Josh Knapp 87b3ad94f9 Improve import UX: progress overlay, pyannote fix, debug logging
- Enhanced ProgressOverlay with spinner, better styling, and z-index 9999
- Import button shows "Processing..." with pulse animation while transcribing
- Fix pyannote API: use token= instead of deprecated use_auth_token=
- Read HF_TOKEN from environment for pyannote model download
- Add console logging for click-to-seek debugging
- Add color-scheme: dark for native form controls

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-02-26 17:43:49 -08:00

Voice to Notes

A desktop application that transcribes audio/video recordings with speaker identification, producing editable transcriptions with synchronized audio playback.

Goals

  • Speech-to-Text Transcription — Accurately convert spoken audio from recordings into text
  • Speaker Identification (Diarization) — Detect and distinguish between different speakers in a conversation
  • Speaker Naming — Assign and persist speaker names/IDs across the transcription
  • Synchronized Playback — Click any transcribed text segment to play back the corresponding audio for review and correction
  • Export Formats
    • Closed captioning files (SRT, VTT) for video
    • Plain text documents with speaker labels
  • AI Integration — Connect to AI providers to ask questions about the conversation and generate condensed notes/summaries

Platform Support

Platform Status
Linux Planned (initial target)
Windows Planned (initial target)
macOS Future (pending hardware)

Project Status

Early planning phase — Architecture and technology decisions in progress.

License

MIT

Description
Convert recorded audio to text with speaker identifying and text to audio scrubbing
Readme MIT 1.1 MiB
2026-03-24 02:04:26 +00:00
Languages
Python 36.6%
Svelte 30.3%
Rust 29.6%
TypeScript 2.2%
Shell 0.5%
Other 0.8%