<?xml version="1.0" encoding="UTF-8"?>
<feed xmlns="http://www.w3.org/2005/Atom">
    <title>speech-recognition</title>
    <link rel="self" type="application/atom+xml" href="https://links.biapy.com/guest/tags/90/feed"/>
    <updated>2026-08-25T23:06:46+00:00</updated>
    <id>https://links.biapy.com/guest/tags/90/feed</id>
            <entry>
            <id>https://links.biapy.com/links/13675</id>
            <title type="text"><![CDATA[Voxtype]]></title>
            <link rel="alternate" href="https://voxtype.io/" />
            <link rel="via" type="application/atom+xml" href="https://links.biapy.com/links/13675"/>
            <author>
                <name><![CDATA[Biapy]]></name>
            </author>
            <summary type="text">
                <![CDATA[Voice-to-text with push-to-talk for Wayland compositors.

Voice-to-text for Linux. 9-11× realtime on your CPU. Local by default.

Hold a hotkey (default: ScrollLock) while speaking, release to transcribe and output the text at your cursor position. Voxtype runs Cohere Transcribe (#1 on the Open ASR Leaderboard) faster than realtime on a plain Zen 4 CPU. Parakeet, Whisper, and five more engines if you want them. No cloud, no subscription, no telemetry.

- [Voxtype @ GitHub](https://github.com/peteonrails/voxtype).

Related contents:

- [88: Talking to my Computer @ Linux Matters](https://linuxmatters.sh/88/).]]>
            </summary>
            <updated>2026-08-19T06:49:18+00:00</updated>
        </entry>
            <entry>
            <id>https://links.biapy.com/links/11998</id>
            <title type="text"><![CDATA[shuo 说]]></title>
            <link rel="alternate" href="https://github.com/NickTikhonov/shuo" />
            <link rel="via" type="application/atom+xml" href="https://links.biapy.com/links/11998"/>
            <author>
                <name><![CDATA[Biapy]]></name>
            </author>
            <summary type="text">
                <![CDATA[sub-500ms latency phone agent orchestration.
A voice agent framework in ~600 lines of Python.

Related contents:

- [How I built a sub-500ms latency voice agent from scratch @ Nick Tikhonov](https://www.ntik.me/posts/voice-agent).]]>
            </summary>
            <updated>2026-03-03T13:11:31+00:00</updated>
        </entry>
            <entry>
            <id>https://links.biapy.com/links/11997</id>
            <title type="text"><![CDATA[Silero VAD]]></title>
            <link rel="alternate" href="https://github.com/snakers4/silero-vad" />
            <link rel="via" type="application/atom+xml" href="https://links.biapy.com/links/11997"/>
            <author>
                <name><![CDATA[Biapy]]></name>
            </author>
            <summary type="text">
                <![CDATA[pre-trained enterprise-grade Voice Activity Detector.

Related contents:

- [How I built a sub-500ms latency voice agent from scratch @ Nick Tikhonov](https://www.ntik.me/posts/voice-agent).]]>
            </summary>
            <updated>2026-03-03T13:08:44+00:00</updated>
        </entry>
            <entry>
            <id>https://links.biapy.com/links/11950</id>
            <title type="text"><![CDATA[Moonshine Voice]]></title>
            <link rel="alternate" href="https://github.com/moonshine-ai/moonshine" />
            <link rel="via" type="application/atom+xml" href="https://links.biapy.com/links/11950"/>
            <author>
                <name><![CDATA[Biapy]]></name>
            </author>
            <summary type="text">
                <![CDATA[Fast and accurate automatic speech recognition (ASR) for edge devices.

Moonshine Voice is an open source AI toolkit for developers building real-time voice applications.]]>
            </summary>
            <updated>2026-02-27T12:40:37+00:00</updated>
        </entry>
            <entry>
            <id>https://links.biapy.com/links/10480</id>
            <title type="text"><![CDATA[Handy]]></title>
            <link rel="alternate" href="https://handy.computer/" />
            <link rel="via" type="application/atom+xml" href="https://links.biapy.com/links/10480"/>
            <author>
                <name><![CDATA[Biapy]]></name>
            </author>
            <summary type="text">
                <![CDATA[speak into any text field.

A free, open source, and extensible speech-to-text application that works completely offline. 

Handy is a cross-platform desktop application built with Tauri (Rust + React/TypeScript) that provides simple, privacy-focused speech transcription. Press a shortcut, speak, and have your words appear in any text field—all without sending your voice to the cloud.

- [Handy @ GitHub](https://github.com/cjpais/Handy).

Related contents:

- [Handy - Un outil de reconnaissance vocale incroyable (et open source) @ Korben :fr:](https://korben.info/handy-computer-speech-to-text-accessibility-open-s.html).]]>
            </summary>
            <updated>2025-11-03T08:51:14+00:00</updated>
        </entry>
            <entry>
            <id>https://links.biapy.com/links/10139</id>
            <title type="text"><![CDATA[Monologue]]></title>
            <link rel="alternate" href="https://www.monologue.to/" />
            <link rel="via" type="application/atom+xml" href="https://links.biapy.com/links/10139"/>
            <author>
                <name><![CDATA[Biapy]]></name>
            </author>
            <summary type="text">
                <![CDATA[Speech to text. talk to the computer.

Related contents:

- [I Started Talking to My Computer Instead of Typing. It Changed How I Think. @ Working Overtime&amp;#039;s Every](https://every.to/working-overtime/i-didn-t-know-typing-held-me-back-until-i-started-thinking-out-loud).]]>
            </summary>
            <updated>2025-09-12T05:47:57+00:00</updated>
        </entry>
            <entry>
            <id>https://links.biapy.com/links/25</id>
            <title type="text"><![CDATA[Whispering]]></title>
            <link rel="alternate" href="https://github.com/epicenter-so/epicenter/tree/main/apps/whispering" />
            <link rel="via" type="application/atom+xml" href="https://links.biapy.com/links/25"/>
            <author>
                <name><![CDATA[Biapy]]></name>
            </author>
            <summary type="text">
                <![CDATA[Whispering is an open-source speech-to-text application. Press a keyboard shortcut, speak, and your words will transcribe, transform, then copy and paste at the cursor.]]>
            </summary>
            <updated>2026-08-13T10:00:27+00:00</updated>
        </entry>
            <entry>
            <id>https://links.biapy.com/links/78</id>
            <title type="text"><![CDATA[Kyutai STT]]></title>
            <link rel="alternate" href="https://kyutai.org/next/stt" />
            <link rel="via" type="application/atom+xml" href="https://links.biapy.com/links/78"/>
            <author>
                <name><![CDATA[Biapy]]></name>
            </author>
            <summary type="text">
                <![CDATA[Kyutai&amp;#039;s Speech-To-Text and Text-To-Speech models based on the Delayed Streams Modeling framework. 

Kyutai STT is a streaming speech-to-text model architecture, providing an unmatched trade-off between latency and accuracy, perfect for interactive applications. Its support for batching allows for processing hundreds of concurrent conversations on a single GPU.

- [Kyutai STT &amp;amp; TTS @ GitHub](https://github.com/kyutai-labs/delayed-streams-modeling).

Related contents:

- [Transcribe speech 100x faster and 100x cheaper with open models @ Modal](https://modal.com/blog/fast-cheap-batch-transcription).]]>
            </summary>
            <updated>2026-02-11T07:15:28+00:00</updated>
        </entry>
            <entry>
            <id>https://links.biapy.com/links/79</id>
            <title type="text"><![CDATA[nvidia/canary-1b-flash @ Hugging Face]]></title>
            <link rel="alternate" href="https://huggingface.co/nvidia/canary-1b-flash" />
            <link rel="via" type="application/atom+xml" href="https://links.biapy.com/links/79"/>
            <author>
                <name><![CDATA[Biapy]]></name>
            </author>
            <summary type="text">
                <![CDATA[canary-1b-flash supports automatic speech-to-text recognition (ASR) in four languages (English, German, French, Spanish) and translation from English to German/French/Spanish and from German/French/Spanish to English with or without punctuation and capitalization (PnC).

Related contents:

- [Transcribe speech 100x faster and 100x cheaper with open models @ Modal](https://modal.com/blog/fast-cheap-batch-transcription).]]>
            </summary>
            <updated>2025-09-18T05:53:10+00:00</updated>
        </entry>
            <entry>
            <id>https://links.biapy.com/links/80</id>
            <title type="text"><![CDATA[nvidia/parakeet-tdt-0.6b-v2 @ Hugging Face]]></title>
            <link rel="alternate" href="https://huggingface.co/nvidia/parakeet-tdt-0.6b-v2" />
            <link rel="via" type="application/atom+xml" href="https://links.biapy.com/links/80"/>
            <author>
                <name><![CDATA[Biapy]]></name>
            </author>
            <summary type="text">
                <![CDATA[parakeet-tdt-0.6b-v2 is a 600-million-parameter automatic speech recognition (ASR) model designed for high-quality English transcription, featuring support for punctuation, capitalization, and accurate timestamp prediction.

Related contents:

- [Transcribe speech 100x faster and 100x cheaper with open models @ Modal](https://modal.com/blog/fast-cheap-batch-transcription).]]>
            </summary>
            <updated>2025-09-18T05:53:12+00:00</updated>
        </entry>
            <entry>
            <id>https://links.biapy.com/links/1398</id>
            <title type="text"><![CDATA[superwhisper]]></title>
            <link rel="alternate" href="https://superwhisper.com/" />
            <link rel="via" type="application/atom+xml" href="https://links.biapy.com/links/1398"/>
            <author>
                <name><![CDATA[Biapy]]></name>
            </author>
            <summary type="text">
                <![CDATA[Al powered voice to text.

Write 3x faster, without lifting a finger. 

Related contents:

- [Vibe Coding and the Future of Software Engineering @ Alex P](https://alexp.pl/2025/02/19/vibe-coding.html).]]>
            </summary>
            <updated>2026-04-07T10:16:23+00:00</updated>
        </entry>
            <entry>
            <id>https://links.biapy.com/links/1511</id>
            <title type="text"><![CDATA[Neon Core]]></title>
            <link rel="alternate" href="https://github.com/NeonGeckoCom/NeonCore" />
            <link rel="via" type="application/atom+xml" href="https://links.biapy.com/links/1511"/>
            <author>
                <name><![CDATA[Biapy]]></name>
            </author>
            <summary type="text">
                <![CDATA[Neon Core extends Mycroft core with more modular code, extended multi-user support, and more.

Neon AI is an open source voice assistant.]]>
            </summary>
            <updated>2025-08-28T20:08:22+00:00</updated>
        </entry>
            <entry>
            <id>https://links.biapy.com/links/1512</id>
            <title type="text"><![CDATA[Open Voice OS]]></title>
            <link rel="alternate" href="https://www.openvoiceos.org/" />
            <link rel="via" type="application/atom+xml" href="https://links.biapy.com/links/1512"/>
            <author>
                <name><![CDATA[Biapy]]></name>
            </author>
            <summary type="text">
                <![CDATA[🌟 OpenVoiceOS is an open-source platform for smart speakers and other voice-centric devices.

OpenVoiceOS is a community-driven, open-source voice AI platform for creating custom voice-controlled ​interfaces across devices with NLP, a customizable UI, and a focus on privacy and security.

- [Open Voice OS core @ GitHub](https://github.com/OpenVoiceOS/ovos-core).]]>
            </summary>
            <updated>2025-08-28T20:08:22+00:00</updated>
        </entry>
            <entry>
            <id>https://links.biapy.com/links/2274</id>
            <title type="text"><![CDATA[EmoBox]]></title>
            <link rel="alternate" href="https://emo-box.github.io/" />
            <link rel="via" type="application/atom+xml" href="https://links.biapy.com/links/2274"/>
            <author>
                <name><![CDATA[Biapy]]></name>
            </author>
            <summary type="text">
                <![CDATA[Multilingual Multi-corpus Speech Emotion Recognition Toolkit and Benchmark.

EmoBox, a groundbreaking multilingual multi-corpus speech emotion recognition (SER) toolkit designed to streamline research in this field. EmoBox is accompanied by a meticulously curated benchmark tailored for both intra-corpus and cross-corpus evaluation settings. 

- [EmoBox @ GitHub](https://github.com/emo-box/emobox).]]>
            </summary>
            <updated>2025-08-28T22:16:39+00:00</updated>
        </entry>
            <entry>
            <id>https://links.biapy.com/links/2975</id>
            <title type="text"><![CDATA[Hertz-dev]]></title>
            <link rel="alternate" href="https://github.com/Standard-Intelligence/hertz-dev" />
            <link rel="via" type="application/atom+xml" href="https://links.biapy.com/links/2975"/>
            <author>
                <name><![CDATA[Biapy]]></name>
            </author>
            <summary type="text">
                <![CDATA[Hertz-dev is an open-source, first-of-its-kind base model for full-duplex conversational audio.]]>
            </summary>
            <updated>2025-08-29T00:12:18+00:00</updated>
        </entry>
            <entry>
            <id>https://links.biapy.com/links/3204</id>
            <title type="text"><![CDATA[🍓 Ichigo]]></title>
            <link rel="alternate" href="https://github.com/homebrewltd/ichigo" />
            <link rel="via" type="application/atom+xml" href="https://links.biapy.com/links/3204"/>
            <author>
                <name><![CDATA[Biapy]]></name>
            </author>
            <summary type="text">
                <![CDATA[Llama3.1 learns to Listen. Local real-time voice AI (Formerly llama3-s).

🍓 Ichigo is an open, ongoing research experiment to extend a text-based LLM to have native &amp;quot;listening&amp;quot; ability. Think of it as an open data, open weight, on device Siri.]]>
            </summary>
            <updated>2025-08-29T00:50:41+00:00</updated>
        </entry>
            <entry>
            <id>https://links.biapy.com/links/3880</id>
            <title type="text"><![CDATA[say]]></title>
            <link rel="alternate" href="https://github.com/8ta4/say" />
            <link rel="via" type="application/atom+xml" href="https://links.biapy.com/links/3880"/>
            <author>
                <name><![CDATA[Biapy]]></name>
            </author>
            <summary type="text">
                <![CDATA[say is always on, recording and transcribing your voice 24/7. Whenever inspiration strikes, just say it.]]>
            </summary>
            <updated>2025-08-29T02:43:39+00:00</updated>
        </entry>
            <entry>
            <id>https://links.biapy.com/links/4690</id>
            <title type="text"><![CDATA[Distil-Whisper]]></title>
            <link rel="alternate" href="https://github.com/huggingface/distil-whisper" />
            <link rel="via" type="application/atom+xml" href="https://links.biapy.com/links/4690"/>
            <author>
                <name><![CDATA[Biapy]]></name>
            </author>
            <summary type="text">
                <![CDATA[Distilled variant of Whisper for speech recognition. 6x faster, 50% smaller, within 1% word error rate.

- [Distil-Whisper @ Hugging Face](https://huggingface.co/collections/distil-whisper/distil-whisper-models-65411987e6727569748d2eb6).
- [Distil-Whisper – Pour faire de la reconnaissance vocale rapide @ Korben :fr:](https://korben.info/distil-whisper-revolution-reconnaissance-vocale-automatique.html).]]>
            </summary>
            <updated>2025-08-29T04:58:56+00:00</updated>
        </entry>
            <entry>
            <id>https://links.biapy.com/links/5280</id>
            <title type="text"><![CDATA[AI Transcriptions by Riverside]]></title>
            <link rel="alternate" href="https://riverside.fm/transcription" />
            <link rel="via" type="application/atom+xml" href="https://links.biapy.com/links/5280"/>
            <author>
                <name><![CDATA[Biapy]]></name>
            </author>
            <summary type="text">
                <![CDATA[Accurate AI Transcriptions in Minutes.

Web service proposing to transcribe video and/or audio content using AI]]>
            </summary>
            <updated>2025-08-29T06:36:41+00:00</updated>
        </entry>
            <entry>
            <id>https://links.biapy.com/links/5643</id>
            <title type="text"><![CDATA[Whisper]]></title>
            <link rel="alternate" href="https://openai.com/index/whisper/" />
            <link rel="via" type="application/atom+xml" href="https://links.biapy.com/links/5643"/>
            <author>
                <name><![CDATA[Biapy]]></name>
            </author>
            <summary type="text">
                <![CDATA[Whisper is a general-purpose speech recognition model. It is trained on a large dataset of diverse audio and is also a multitasking model that can perform multilingual speech recognition, speech translation, and language identification.

[Whisper @ GitHub](https://github.com/openai/whisper).]]>
            </summary>
            <updated>2025-08-29T07:37:12+00:00</updated>
        </entry>
            <entry>
            <id>https://links.biapy.com/links/6100</id>
            <title type="text"><![CDATA[Amberscript]]></title>
            <link rel="alternate" href="https://www.amberscript.com/en/" />
            <link rel="via" type="application/atom+xml" href="https://links.biapy.com/links/6100"/>
            <author>
                <name><![CDATA[Biapy]]></name>
            </author>
            <summary type="text">
                <![CDATA[Audio &amp;amp; Video Transcription | Speech-to-text.
Smarter subtitling and transcription.
We combine artificial and human intelligence to bring you accurate and fast transcripts, captions, and translated subtitles with ease.]]>
            </summary>
            <updated>2025-08-29T08:53:51+00:00</updated>
        </entry>
            <entry>
            <id>https://links.biapy.com/links/6803</id>
            <title type="text"><![CDATA[Whisper]]></title>
            <link rel="alternate" href="https://github.com/openai/whisper" />
            <link rel="via" type="application/atom+xml" href="https://links.biapy.com/links/6803"/>
            <author>
                <name><![CDATA[Biapy]]></name>
            </author>
            <summary type="text">
                <![CDATA[Whisper is a general-purpose speech recognition model. It is trained on a large dataset of diverse audio and is also a multi-task model that can perform multilingual speech recognition as well as speech translation and language identification.]]>
            </summary>
            <updated>2025-08-29T10:50:56+00:00</updated>
        </entry>
    </feed>
