oboroge0
hayamimi
早耳 - Real-time multilingual speech-to-text on CPU only. Live subtitles, browser dashboard, speaker labels, translation. No GPU, no cloud.
Documentation snapshot
README 快照
本页保存的是公开项目资料快照,阅读过程不需要连接 GitHub。
hayamimi (早耳)
图片:tests 图片:license 图片:release
Real-time, multilingual speech-to-text on CPU only. Live subtitles, a browser dashboard, speaker labels, and on-the-fly translation — no GPU, no cloud API, under 2GB RAM.
日本語版 README は README.ja.md にあります。
“早耳” (hayamimi) is Japanese for “quick ear” — someone who picks up on things fast. That’s the design goal: partial subtitles appear while you’re still talking, and a finalized line lands roughly 100ms after you stop.
Why
Most CPU-only real-time transcription setups fall back to a single general-purpose model (Whisper) and accept its accuracy ceiling. hayamimi instead routes each utterance to whichever specialist model is best for its language, all running as quantized (INT8) ONNX models via sherpa-onnx — no PyTorch, no CUDA.
On real broadcast Japanese audio (see docs/SCORECARD.md), that routing
gets 5.8% CER, less than half of whisper-large-v3-turbo’s 13.8% on the
same clips, while running at 10-50x realtime on a 6-core desktop CPU.
Features
| Feature | What it does |
|---|---|
| 5-route language catalog | ja/zh/ko/yue/en+24 EU languages each go to a dedicated best-in-class model; everything else (~1600 languages) falls back to Meta’s Omnilingual ASR |
| Partial subtitles | in-progress draft text updates every ~0.5s while you’re still speaking |
| Fast finals | a finalized line typically lands ~100ms after you stop talking (ja; see docs/GOALS.md for other languages) |
| Two-pass refinement | after 2s of silence, recent utterances are batch re-decoded for a higher-accuracy “clean” transcript (ja real-broadcast CER 15.5% -> 12.0%) |
| Speaker labels | --speakers tags each utterance S1/S2/… using CAM++ speaker embeddings (turn-taking, not full diarization) |
| Translation | --translate en,zh,ko translates Japanese lines live (en via FuguMT, zh/ko via M2M-100) |
| Hotwords / user dictionary | --hotwords biases decoding toward proper nouns; --replace does post-hoc find/replace |
| OBS overlay + dashboard | --serve starts a local HTTP server with a browser-source overlay and a live dashboard |
| Memory-bounded | LRU model eviction keeps resident models under a configurable cap (default: <2GB total) |
| CPU-only | every model runs as quantized ONNX via sherpa-onnx; no GPU or PyTorch required |
Demo UI
--serve starts a local server exposing three views:
http://localhost:8765/dashboard— the live dashboard: a partial-text strip for in-progress speech, a finals feed with language badges, speaker chips, and per-line latency, inline translations under each line, and a second column with the refined (two-pass) transcript as it lands.http://localhost:8765/— a minimal OBS browser-source overlay (add this URL as a Browser Source in OBS for stream captions).http://localhost:8765/transcript— plain scrolling transcript history.
图片:dashboard
🎬 Watch the demo video — real 4-language audio (ja/en/ko/zh) transcribed live, replayed frame-accurately from a captured session.
Requirements
Python 3.10+ and ffmpeg on PATH. Developed and tested on Windows 11; macOS/Linux are expected to work (all runtimes are cross-platform) but are not yet CI-tested end to end — reports welcome.
Quickstart
python -m venv .venv
# Windows
.venv\Scripts\pip install -r requirements.txt
.venv\Scripts\python scripts\download_models.py
# macOS / Linux
.venv/bin/pip install -r requirements.txt
.venv/bin/python scripts/download_models.py
# Real-time transcription from your microphone
.venv/Scripts/python scripts/realtime_transcribe.py # Windows
.venv/bin/python scripts/realtime_transcribe.py # macOS/Linux
# With the dashboard + OBS overlay
.venv/Scripts/python scripts/realtime_transcribe.py --serve
# -> open http://localhost:8765/dashboard in a browser
scripts/download_models.py pulls ~3.1GB of pretrained models into
models/ (git-ignored). Pass --minimal for a ~1.1GB ja/en-only install
(ReazonSpeech, whisper-tiny, Silero VAD, Japanese punctuation). See
THIRD_PARTY_NOTICES.md for what each model’s license commits you to.
CLI reference
All flags are on scripts/realtime_transcribe.py:
| Flag | Default | Description |
|---|---|---|
--wav PATH | mic input | simulate streaming from a 16kHz mono WAV file instead of the microphone |
--no-realtime | off | with --wav, don’t sleep between chunks (fast batch processing) |
--threads N | 4 | inference threads per model |
--no-partial | off | disable in-progress draft subtitles |
--min-silence SEC | 0.35 | silence duration that ends an utterance; lower = snappier finals, more splits |
--max-speech SEC | 12.0 | force-finalize an utterance after this many seconds of continuous speech |
--max-resident N | 3 | max non-tier0 models kept resident (LRU eviction); <=0 = unlimited |
--serve [PORT] | off, 8765 | serve the dashboard + OBS overlay at http://localhost:PORT |
--no-refine | off | disable the second-pass re-decode of utterance groups |
--transcript PATH | none | append refined transcript lines to this file |
--hotwords PATH | none | hotword list (one per line) to bias Japanese decoding toward proper nouns |
--replace PATH | none | user dictionary: wrong=right per line, applied to all output |
--lang-switch-guard SEC | 2.0 | treat a new-language detection shorter than this as noise and keep the session language (0 disables) |
--speakers | off | label utterances with speaker ids (S1, S2, …) |
--translate [LANGS] | off, en | translate Japanese lines to these comma-separated languages (en/zh/ko) |
Architecture
┌─────────────┐
mic / wav ───────────▶ │ Silero VAD │ 0.35s end-of-speech + 0.8s preroll
└──────┬──────┘
│ speech segment
▼
┌───────────────────────────┐
│ whisper-tiny spoken-LID │ runs on first ~4s while
│ (+ char-set arbitration) │ the segment is still coming in
└─────────────┬─────────────┘
│ language tag
┌───────────────┼────────────────┬─────────────┬──────────────┐
▼ ▼ ▼ ▼ ▼
┌───────┐ ┌─────────┐ ┌──────────┐ ┌─────────┐ ┌──────────┐
│ ja │ │ zh │ │ ko/yue │ │ en + 24 │ │ ~1600 │
│ Reazon│ │Paraformer│ │SenseVoice│ │EU langs │ │ other │
│Speech │ │ -zh │ │ small │ │Parakeet │ │Omnilingual│
│Zipform│ │ │ │ │ │TDT v3 │ │ ASR │
└───┬───┘ └────┬────┘ └────┬─────┘ └────┬────┘ └────┬─────┘
└───────────────┴────────────────┴─────────────┴─────────────┘
│
partial (every ~0.5s) │ final (~0.1s after end-of-speech)
◀───────────────────────────┴───────────────────────▶
│
┌─────────────────────┼─────────────────────┐
▼ ▼ ▼
┌────────────────┐ ┌──────────────────┐ ┌────────────────┐
│ ja punctuation │ │ speaker labeling │ │ translation │
│ (BERT restore) │ │ (CAM++, --speakers)│ │ (FuguMT/M2M-100)│
└────────────────┘ └──────────────────┘ └────────────────┘
│
2s silence: batch re-decode recent utterances (two-pass refine)
│
▼
dashboard / OBS overlay / transcript file
Models are lazy-loaded on first use; an LRU cache evicts the
least-recently-used non-Japanese models (--max-resident) so memory stays
bounded no matter how many languages a session wanders through.
Measured performance
End-to-end (LID -> routing -> decode -> ja punctuation), real speech, no
preroll/two-pass (single clips). en uses WER, all others use CER (yue
t2s-normalized). Full methodology in docs/SCORECARD.md.
| Language | Clips | LID accuracy | Route | Mean error | Mean RTF |
|---|---|---|---|---|---|
| ja | 15 | 15/15 | ReazonSpeech | 7.5% | 0.071 |
| en | 15 | 15/15 | Parakeet v3 | 2.3% | 0.109 |
| zh | 12 | 12/12 | Paraformer-zh | 5.3% | 0.102 |
| ko | 12 | 12/12 | SenseVoice | 8.1% | 0.062 |
| yue | 12 | 12/12 | SenseVoice | 6.1% | 0.061 |
RTF (real-time factor) well under 0.2 across every route means each route
runs 9-16x faster than realtime on CPU alone — see docs/GOALS.md for the
full target table and docs/BENCHMARKS.md for the complete iteration log
(30+ measured changes, latency/memory/accuracy tradeoffs and why each one was
made or rejected).
Headline numbers from that log:
- Japanese CER 5.8% (beam search) on real broadcast audio, vs. 13.8% for
whisper-large-v3-turboon the same clips — less than half the error rate. - ~100ms mean final latency (ja, punctuated); ~236ms mean / 552ms max across a 5-language soak test with every feature enabled.
- <2GB RAM with
--max-resident 3(1.35GB at--max-resident 2).
Limitations (honest list)
- Code-switching mid-sentence is not supported. The router picks one language per utterance; a sentence that mixes Japanese and English within itself will have the minority-language portion mangled or dropped. Utterance-level switching (e.g. an interpreter alternating full sentences) works well; word-level switching within one sentence does not.
- Very short utterances after a jingle/sting/BGM burst can misroute.
The language-switch guard (
--lang-switch-guard) mitigates this but a session’s very first utterance (before any session language is established) and confidently-wrong LID+decode combinations (where the garbled text happens to match the wrong language’s character set) are known blind spots — seedocs/BENCHMARKS.md’s iteration #29 for a quantified before/after. - Two overlapping speakers are not separated.
--speakersdoes turn-taking speaker labeling (one embedding per finalized VAD segment, nearest-centroid assignment), not true diarization — simultaneous speech gets one label. - Translation quality has a real ceiling, not just a tuning one.
FuguMT (ja->en) and M2M-100 (ja->zh/ko) are small models; repetition loops
are suppressed but not eliminated, and numeric values are not reliably
preserved in ja->zh/ko translation (see
docs/TRANSLATE.mdanddocs/TRANSLATE_M2M.mdfor measured failure cases before you rely on this for anything numeric or financial). - The end-to-end mic pipeline has not been independently verified beyond
this project’s own testing — see
docs/GOALS.md’s remaining-work section. File an issue if your results differ from the numbers above.
License
Source code is MIT (LICENSE, copyright oboroge0). No model weights are
committed to this repository — scripts/download_models.py fetches them
from their original publishers at install time, and each carries its own
license (THIRD_PARTY_NOTICES.md has the full table).
One model is not permissive: the ja->en translation model
(mojicast-fugumt-ja-en-ct2, used by --translate en) is
CC BY-SA 4.0 (share-alike). If you redistribute that model’s weights,
you must keep attribution and license any redistribution under CC BY-SA 4.0
too. This does not affect hayamimi’s own code license, and does not affect
--translate zh,ko (M2M-100, MIT).
Credits
hayamimi exists on top of, and would not exist without:
- k2-fsa/sherpa-onnx — the ONNX Runtime inference engine every model here runs through.
- ReazonSpeech (Reazon Human Interaction Lab) — the Japanese ASR model that anchors this project’s accuracy claim.
- NVIDIA NeMo / Parakeet — English + 24 European languages.
- Meta AI Omnilingual ASR — the ~1600-language fallback that makes “multilingual” not a lie.
- FunASR / SenseVoice (Alibaba DAMO Academy) — Chinese, Korean, and Cantonese ASR.
- Mojicast (ishiki-emo) — design inspiration for the live-captioning pipeline, and the source of the converted punctuation/translation model artifacts this project uses. Mojicast is itself a full offline real-time captioning app worth checking out.
- Silero VAD — voice activity detection.
- 3D-Speaker (Alibaba DAMO
Academy) — the CAM++ speaker embedding model behind
--speakers. - Kiwi — Korean morphological tokenizer, used to fix SenseVoice’s token-spaced Korean output.
Contributing
See CONTRIBUTING.md.
Official distribution
获取与安装
暂未发现可确认的官方软件包地址
当前 README 快照没有出现 npm、PyPI、Crates.io、pub.dev 等官方包页链接。本站不会根据仓库名称猜测下载地址。
本站不托管项目文件;需要安装时,请以项目维护者发布的官方文档为准。
Before installing
使用前核验
本站保存公开资料用于阅读,不代表安全审计或功能背书。安装前请核对许可证、依赖来源和发布签名,不要直接运行来源不明的二进制文件或高权限脚本。