Skip to content
New issue

Have a question about this project? Sign up for a free GitHub account to open an issue and contact its maintainers and the community.

By clicking “Sign up for GitHub”, you agree to our terms of service and privacy statement. We’ll occasionally send you account related emails.

Already on GitHub? Sign in to your account

convert English models? #14

Closed
BarfingLemurs opened this issue Aug 19, 2023 · 2 comments
Closed

convert English models? #14

BarfingLemurs opened this issue Aug 19, 2023 · 2 comments

Comments

@BarfingLemurs
Copy link

@weirdseed
Does any english models work with this?
I receive the same error message as #6, using the windows executable or the python script. I also get the same error message on linux with the executable. What am I doing wrong?

I tried two (and more) models:

I do see weirdseed/Vits-Android-ncnn#13 "English support is added", how does it work?

Do you have a working model example I can copy the config for?

@weirdseed
Copy link
Owner

This project only supports the original vits models or moegoe models.
Here is an config example of vctk model.

  {
    "train": {
      "log_interval": 200,
      "eval_interval": 1000,
      "seed": 1234,
      "epochs": 10000,
      "learning_rate": 2e-4,
      "betas": [0.8, 0.99],
      "eps": 1e-9,
      "batch_size": 64,
      "fp16_run": true,
      "lr_decay": 0.999875,
      "segment_size": 8192,
      "init_lr_ratio": 1,
      "warmup_epochs": 0,
      "c_mel": 45,
      "c_kl": 1.0
    },
    "data": {
      "training_files":"filelists/vctk_audio_sid_text_train_filelist.txt.cleaned",
      "validation_files":"filelists/vctk_audio_sid_text_val_filelist.txt.cleaned",
      "text_cleaners":["english_cleaners2"],
      "max_wav_value": 32768.0,
      "sampling_rate": 22050,
      "filter_length": 1024,
      "hop_length": 256,
      "win_length": 1024,
      "n_mel_channels": 80,
      "mel_fmin": 0.0,
      "mel_fmax": null,
      "add_blank": true,
      "n_speakers": 109,
      "cleaned_text": true
    },
    "model": {
      "inter_channels": 192,
      "hidden_channels": 192,
      "filter_channels": 768,
      "n_heads": 2,
      "n_layers": 6,
      "kernel_size": 3,
      "p_dropout": 0.1,
      "resblock": "1",
      "resblock_kernel_sizes": [3,7,11],
      "resblock_dilation_sizes": [[1,3,5], [1,3,5], [1,3,5]],
      "upsample_rates": [8,8,2,2],
      "upsample_initial_channel": 512,
      "upsample_kernel_sizes": [16,16,4,4],
      "n_layers_q": 3,
      "use_spectral_norm": false,
      "gin_channels": 256
    },
    "speakers": [0, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, 100, 101, 102, 103, 104, 105, 106, 107, 108],
    "symbols": ["_", ";", ":", ",", ".", "!", "?", "¡", "¿", "—", "…", "\"", "«", "»", "“", "”", " ", "A", "B", "C", "D", "E", "F", "G", "H", "I", "J", "K", "L", "M", "N", "O", "P", "Q", "R", "S", "T", "U", "V", "W", "X", "Y", "Z", "a", "b", "c", "d", "e", "f", "g", "h", "i", "j", "k", "l", "m", "n", "o", "p", "q", "r", "s", "t", "u", "v", "w", "x", "y", "z", "ɑ", "ɐ", "ɒ", "æ", "ɓ", "ʙ", "β", "ɔ", "ɕ", "ç", "ɗ", "ɖ", "ð", "ʤ", "ə", "ɘ", "ɚ", "ɛ", "ɜ", "ɝ", "ɞ", "ɟ", "ʄ", "ɡ", "ɠ", "ɢ", "ʛ", "ɦ", "ɧ", "ħ", "ɥ", "ʜ", "ɨ", "ɪ", "ʝ", "ɭ", "ɬ", "ɫ", "ɮ", "ʟ", "ɱ", "ɯ", "ɰ", "ŋ", "ɳ", "ɲ", "ɴ", "ø", "ɵ", "ɸ", "θ", "œ", "ɶ", "ʘ", "ɹ", "ɺ", "ɾ", "ɻ", "ʀ", "ʁ", "ɽ", "ʂ", "ʃ", "ʈ", "ʧ", "ʉ", "ʊ", "ʋ", "ⱱ", "ʌ", "ɣ", "ɤ", "ʍ", "χ", "ʎ", "ʏ", "ʑ", "ʐ", "ʒ", "ʔ", "ʡ", "ʕ", "ʢ", "ǀ", "ǁ", "ǂ", "ǃ", "ˈ", "ˌ", "ː", "ˑ", "ʼ", "ʴ", "ʰ", "ʱ", "ʲ", "ʷ", "ˠ", "ˤ", "˞", "↓", "↑", "→", "↗", "↘", "'", "̩", "'", "ᵻ"]
  }

@BarfingLemurs
Copy link
Author

BarfingLemurs commented Aug 19, 2023

Thanks for the help! I could convert the moegoe 365 model successfully.
but the original vits models don't work for me with that config.
I found the issue, the exe is too old for converting the english. need to use python script.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment
Labels
None yet
Projects
None yet
Development

No branches or pull requests

2 participants