Fix speaker diarization: WAV conversion, pyannote 4.0 compat, telemetry bug

- Convert non-WAV audio to 16kHz mono WAV before diarization (pyannote
  v4.0.4 AudioDecoder returns None duration for FLAC, causing crash)
- Handle pyannote 4.0 DiarizeOutput return type (unwrap .speaker_diarization)
- Disable pyannote telemetry (np.isfinite(None) bug with max_speakers)
- Use huggingface_hub.login() to persist token for all sub-downloads
- Pre-download sub-models (segmentation-3.0, speaker-diarization-community-1)
- Add third required model license link in settings UI
- Improve SpeakerManager hints based on settings state
- Add word-wrap to transcript text

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
This commit is contained in:
2026-02-26 19:46:07 -08:00
co-authored by Claude Opus 4.6
parent a3612c986d
commit 585411f402
6 changed files with 133 additions and 25 deletions
+8 -4
View File
@@ -1,5 +1,6 @@
<script lang="ts">
import { speakers } from '$lib/stores/transcript';
import { settings } from '$lib/stores/settings';
import type { Speaker } from '$lib/types/transcript';
let editingSpeakerId = $state<string | null>(null);
@@ -35,10 +36,13 @@
<h3>Speakers</h3>
{#if $speakers.length === 0}
<p class="empty-hint">No speakers detected</p>
<p class="setup-hint">
Speaker detection requires a HuggingFace token.
Set the <code>HF_TOKEN</code> environment variable and restart.
</p>
{#if $settings.skip_diarization}
<p class="setup-hint">Speaker detection is disabled. Enable it in Settings &gt; Speakers.</p>
{:else if !$settings.hf_token}
<p class="setup-hint">Speaker detection requires a HuggingFace token. Configure it in Settings &gt; Speakers.</p>
{:else}
<p class="setup-hint">Speaker detection ran but found no distinct speakers, or the model may need to be downloaded. Check Settings &gt; Speakers.</p>
{/if}
{:else}
<ul class="speaker-list">
{#each $speakers as speaker (speaker.id)}