# Voice cloning: compact engine sources

- Model weights, tokenizer models, BOS tensors, metadata, and model card: KevinAHM/pocket-tts-onnx, revision `58a6d00cf13d239b6748cb0769f35c580a8f606c`.
  https://huggingface.co/KevinAHM/pocket-tts-onnx/tree/58a6d00cf13d239b6748cb0769f35c580a8f606c
  The repository's `onnx/LICENSE` is CC BY 4.0. The mirrored license is preserved in `voice-clone-pocket-LICENSE.txt`.
- Original model author: Kyutai. https://huggingface.co/kyutai/pocket-tts
- Browser inference implementation: KevinAHM/pocket-tts-web, revision `d0c0c79b7712256a32d691c67f20b8ae2e020d00`, `inference-worker.js`, Apache 2.0 (`CODE-LICENSE`). The local engine adapts the upload-reference encoder, voice conditioning, autoregressive flow sampling, and streaming decoder. It adds checked immutable loading, cancellation, explicit tensor/session disposal, typed errors, and original-token chunking. Preset voice loading is omitted.
  https://huggingface.co/spaces/KevinAHM/pocket-tts-web/blob/d0c0c79b7712256a32d691c67f20b8ae2e020d00/inference-worker.js
- SentencePiece browser runtime: the same Space and revision, `sentencepiece.js`. This dependency is mirrored intact with its embedded WASM and Microsoft helper copyright/permission notice. SentencePiece is Copyright Google Inc., licensed under Apache 2.0; the JavaScript wrapper is provided under the Space's Apache 2.0 code license. https://github.com/google/sentencepiece/blob/master/LICENSE
  https://huggingface.co/spaces/KevinAHM/pocket-tts-web/blob/d0c0c79b7712256a32d691c67f20b8ae2e020d00/sentencepiece.js
- ONNX Runtime Web: existing repository vendor, `1.31.0-dev.20260914-8d85527a0`, MIT license already retained under `scripts/licenses/onnxruntime-*` and the vendor's notices. It runs with one CPU WASM thread. https://github.com/microsoft/onnxruntime

The Space's `onnx/ONNX-LICENSE` differs from the model repository's license. No model weight is sourced from that Space; all model requests and mirror identities use the model repository and its explicit CC BY 4.0 license.

Five independent language bundles: English (`english_2026-04`), German, Spanish, Italian, Portuguese. Each contains its own upload-reference encoder. No `voices.bin` or preset speaker embedding is loaded. Input reference audio and text remain in the browser; remote requests fetch only fixed public assets. Download byte totals include the shared SentencePiece runtime and the reused ORT WASM, and do not represent peak working memory.
