The Speakr project is designed to help users manage and search through large collections of audio files, such as meeting recordings, interviews, and voice notes. It allows users to either upload existing files or record directly from their browser, capturing both microphone input and system sounds. Once processed, the audio is converted into clear, searchable text. This text is split by speaker, and the system provides summaries and notes that can be used for quick reference. Users can click on specific parts of the transcription to jump to the exact moment in the audio.
One of Speakr’s unique features is "Inquire," which lets users ask questions about the entire collection of recordings. For example, a user could ask, "What did we decide about the price, and who was not in agreement?" The system then searches through the recordings and provides answers, often with links to specific parts of the audio. There's also an optional beta feature called "agent mode," which automatically reads through relevant recordings and provides answers with numbered links to the relevant excerpts.
While Speakr is marketed as a self-hosted tool—emphasizing user privacy—the transcription and processing of audio involve external services. The audio is sent to a third-party engine, such as OpenAI or OpenRouter, to convert it into text. The text is then processed by another model to generate summaries, titles, and chat responses. For users who want to keep everything local, an alternative called Ollama can be used instead of the external services. However, this requires a graphics card and a large Whisper model for effective transcription, as noted in the project’s documentation. Without a graphics card, the system still works but will be slower.
Speakr is built using Docker and supports both SQLite and PostgreSQL for data storage. It requires at least 2 GB of RAM for the application alone and includes features common in multi-user applications, such as user accounts, single sign-on, sharing options, signed APIs, and webhooks. The software is released under the AGPLv3 license, which allows for free use but requires that any modifications be shared with the community. A commercial license is also available for users who prefer not to be bound by the AGPLv3 terms.
Other similar tools include Scriberr, which transcribes audio locally without sending data to external servers, but its development has paused. Another alternative is Meetily, which focuses on individual workstations rather than team servers and offers voice separation as a paid feature.
Speakr Project Offers Self-Hosted Audio Transcription and Search with Privacy Concerns
AI-rewritten from original reportingHow it works
audio-transcriptionself-hostedprivacyai-insightsdockeropen-source
Original sources:
- 🇫🇷Korben



