Merge pull request #72 from iclementine/example_readme

add README for transformer_tts, waveflow and wavenet
2020-12-30 15:56:46 +08:00 · 2020-12-30 15:56:46 +08:00 · 737b09d03c
parent f9b39b97dd f5027a5e6f
commit 737b09d03c
3 changed files with 144 additions and 0 deletions
--- a/examples/transformer_tts/README.md
+++ b/examples/transformer_tts/README.md
@ -0,0 +1,48 @@
 # TransformerTTS with LJSpeech
 ## Dataset
 ### Download the datasaet.
 ```bash
 wget https://data.keithito.com/data/speech/LJSpeech-1.1.tar.bz2
 ```
 ### Extract the dataset.
 ```bash
 tar xjvf LJSpeech-1.1.tar.bz2
 ```
 ### Preprocess the dataset. 
 Assume the path to save the preprocessed dataset is `ljspeech_transformer_tts`. Run the command below to preprocess the dataset.
 ```bash
 python preprocess.py --input=LJSpeech-1.1/  --output=ljspeech_transformer_tts
 ```
 ## Train the model
 The training script requires 4 command line arguments.
 `--data` is the path of the training dataset, `--output` is the path of the output direcctory (we recommend to use a subdirectory in `runs` to manage different experiments.)
 `--device` should be "cpu" or "gpu", `--nprocs` is the number of processes to train the model in parallel.
 ```bash
 python train.py --data=ljspeech_transformer_tts/ --output=runs/test --device="gpu" --nprocs=1
 ```
 If you want distributed training, set a larger `--nprocs` (e.g. 4). Note that distributed training with cpu is not supported yet.
 ## Synthesize
 Synthesize waveform. We assume the `--input` is a text file, one sentence per line, and `--output` is a directory to save the synthesized mel spectrogram(log magnitude) in `.npy` format. The mel spectrograms can be used with `Waveflow` to generate waveforms.
 `--checkpoint_path` should be the path of the parameter file (`.pdparams`) to load. Note that the extention name `.pdparmas` is not included here.
 `--device` specifies to device to run synthesis on.
 ```bash
 python synthesize.py --input=sentence.txt --output=mels/ --checkpoint_path='step-310000' --device="gpu" --verbose
 ```
--- a/examples/waveflow/README.md
+++ b/examples/waveflow/README.md
@ -0,0 +1,48 @@
 # WaveFlow with LJSpeech
 ## Dataset
 ### Download the datasaet.
 ```bash
 wget https://data.keithito.com/data/speech/LJSpeech-1.1.tar.bz2
 ```
 ### Extract the dataset.
 ```bash
 tar xjvf LJSpeech-1.1.tar.bz2
 ```
 ### Preprocess the dataset. 
 Assume the path to save the preprocessed dataset is `ljspeech_waveflow`. Run the command below to preprocess the dataset.
 ```bash
 python preprocess.py --input=LJSpeech-1.1/  --output=ljspeech_waveflow
 ```
 ## Train the model
 The training script requires 4 command line arguments.
 `--data` is the path of the training dataset, `--output` is the path of the output directory (we recommend to use a subdirectory in `runs` to manage different experiments.)
 `--device` should be "cpu" or "gpu", `--nprocs` is the number of processes to train the model in parallel.
 ```bash
 python train.py --data=ljspeech_waveflow/ --output=runs/test --device="gpu" --nprocs=1
 ```
 If you want distributed training, set a larger `--nprocs` (e.g. 4). Note that distributed training with cpu is not supported yet.
 ## Synthesize
 Synthesize waveform. We assume the `--input` is a directory containing several mel spectrograms(log magnitude) in `.npy` format. The output would be saved in `--output` directory, containing several `.wav` files, each with the same name as the mel spectrogram does.
 `--checkpoint_path` should be the path of the parameter file (`.pdparams`) to load. Note that the extention name `.pdparmas` is not included here.
 `--device` specifies to device to run synthesis on.
 ```bash
 python synthesize.py --input=mels/ --output=wavs/ --checkpoint_path='step-2000000' --device="gpu" --verbose
 ```
--- a/examples/wavenet/README.md
+++ b/examples/wavenet/README.md
@ -0,0 +1,48 @@
 # WaveNet with LJSpeech
 ## Dataset
 ### Download the datasaet.
 ```bash
 wget https://data.keithito.com/data/speech/LJSpeech-1.1.tar.bz2
 ```
 ### Extract the dataset.
 ```bash
 tar xjvf LJSpeech-1.1.tar.bz2
 ```
 ### Preprocess the dataset. 
 Assume the path to save the preprocessed dataset is `ljspeech_wavenet`. Run the command below to preprocess the dataset.
 ```bash
 python preprocess.py --input=LJSpeech-1.1/  --output=ljspeech_wavenet
 ```
 ## Train the model
 The training script requires 4 command line arguments.
 `--data` is the path of the training dataset, `--output` is the path of the output directory (we recommend to use a subdirectory in `runs` to manage different experiments.)
 `--device` should be "cpu" or "gpu", `--nprocs` is the number of processes to train the model in parallel.
 ```bash
 python train.py --data=ljspeech_wavenet/ --output=runs/test --device="gpu" --nprocs=1
 ```
 If you want distributed training, set a larger `--nprocs` (e.g. 4). Note that distributed training with cpu is not supported yet.
 ## Synthesize
 Synthesize waveform. We assume the `--input` is a directory containing several mel spectrograms(normalized into range[0, 1)) in `.npy` format. The output would be saved in `--output` directory, containing several `.wav` files, each with the same name as the mel spectrogram does.
 `--checkpoint_path` should be the path of the parameter file (`.pdparams`) to load. Note that the extention name `.pdparmas` is not included here.
 `--device` specifies to device to run synthesis on. Due to the autoregressiveness of wavenet, using cpu may be faster.
 ```bash
 python synthesize.py --input=mels/ --output=wavs/ --checkpoint_path='step-2450000' --device="cpu" --verbose
 ```