Archived
First commit of local Closed Caption App
This commit is contained in:
@@ -0,0 +1,17 @@
|
||||
# Whisper Configuration
|
||||
# For local Whisper, ensure whisper is installed via pip
|
||||
WHISPER_MODEL=base # Options: tiny, base, small, medium, large
|
||||
|
||||
# Remote Whisper API (e.g., OpenAI Whisper API)
|
||||
OPENAI_API_KEY=your_openai_api_key_here
|
||||
|
||||
# ElevenLabs Speech-to-Text API
|
||||
ELEVENLABS_API_KEY=your_elevenlabs_api_key_here
|
||||
|
||||
# Audio Processing Configuration
|
||||
# Audio threshold for silence detection (0.0-1.0, lower = more sensitive)
|
||||
# 0.005 = Very sensitive, 0.01 = Default, 0.05 = Less sensitive
|
||||
AUDIO_THRESHOLD=0.01
|
||||
|
||||
# Server Configuration
|
||||
SERVER_PORT=3000
|
||||
+16
@@ -0,0 +1,16 @@
|
||||
node_modules/
|
||||
dist/
|
||||
.env
|
||||
.DS_Store
|
||||
*.log
|
||||
*.lock
|
||||
package-lock.json
|
||||
yarn.lock
|
||||
build/
|
||||
out/
|
||||
.vscode/
|
||||
.idea/
|
||||
*.wav
|
||||
*.mp3
|
||||
*.webm
|
||||
*.txt
|
||||
@@ -0,0 +1,141 @@
|
||||
# Installing Whisper for Local Transcription
|
||||
|
||||
This guide explains how to install OpenAI's Whisper for local speech-to-text transcription.
|
||||
|
||||
## Prerequisites
|
||||
|
||||
- Python 3.8 or newer
|
||||
- pip (Python package manager)
|
||||
- ffmpeg
|
||||
|
||||
## Installation Steps
|
||||
|
||||
### 1. Install Python Dependencies
|
||||
|
||||
```bash
|
||||
# Install Whisper
|
||||
pip install openai-whisper
|
||||
|
||||
# Or install with specific version
|
||||
pip install openai-whisper==20230918
|
||||
```
|
||||
|
||||
### 2. Install ffmpeg
|
||||
|
||||
#### Ubuntu/Debian:
|
||||
```bash
|
||||
sudo apt update
|
||||
sudo apt install ffmpeg
|
||||
```
|
||||
|
||||
#### macOS (using Homebrew):
|
||||
```bash
|
||||
brew install ffmpeg
|
||||
```
|
||||
|
||||
#### Windows:
|
||||
1. Download ffmpeg from https://ffmpeg.org/download.html
|
||||
2. Extract the archive
|
||||
3. Add the bin folder to your PATH
|
||||
|
||||
### 3. Verify Installation
|
||||
|
||||
```bash
|
||||
# Test Whisper installation
|
||||
whisper --help
|
||||
|
||||
# Test with an audio file
|
||||
whisper audio.mp3 --model base
|
||||
```
|
||||
|
||||
## Available Models
|
||||
|
||||
Whisper offers several models with different speed/accuracy tradeoffs:
|
||||
|
||||
| Model | Parameters | Relative Speed | Required VRAM |
|
||||
|--------|------------|----------------|---------------|
|
||||
| tiny | 39 M | ~32x | ~1 GB |
|
||||
| base | 74 M | ~16x | ~1 GB |
|
||||
| small | 244 M | ~6x | ~2 GB |
|
||||
| medium | 769 M | ~2x | ~5 GB |
|
||||
| large | 1550 M | 1x | ~10 GB |
|
||||
|
||||
## Configuration
|
||||
|
||||
Set your preferred model in the `.env` file:
|
||||
|
||||
```env
|
||||
WHISPER_MODEL=base # Options: tiny, base, small, medium, large
|
||||
```
|
||||
|
||||
## GPU Acceleration Setup
|
||||
|
||||
### NVIDIA GPUs (CUDA)
|
||||
```bash
|
||||
# Install PyTorch with CUDA support
|
||||
pip install torch torchvision torchaudio --index-url https://download.pytorch.org/whl/cu118
|
||||
|
||||
# Verify CUDA is available
|
||||
python -c "import torch; print(f'CUDA available: {torch.cuda.is_available()}')"
|
||||
```
|
||||
|
||||
### AMD GPUs (ROCm) - Linux Only
|
||||
```bash
|
||||
# Install ROCm (Ubuntu/Debian)
|
||||
wget -q -O - https://repo.radeon.com/rocm/rocm.gpg.key | sudo apt-key add -
|
||||
echo 'deb [arch=amd64] https://repo.radeon.com/rocm/apt/debian/ ubuntu main' | sudo tee /etc/apt/sources.list.d/rocm.list
|
||||
sudo apt update
|
||||
sudo apt install rocm-dkms
|
||||
|
||||
# Install PyTorch with ROCm support
|
||||
pip install torch torchvision torchaudio --index-url https://download.pytorch.org/whl/rocm5.4.2
|
||||
|
||||
# Add user to render/video groups
|
||||
sudo usermod -a -G render,video $USER
|
||||
# Logout and login for group changes to take effect
|
||||
```
|
||||
|
||||
### Intel GPUs (OpenVINO)
|
||||
```bash
|
||||
# Install OpenVINO toolkit
|
||||
pip install openvino-dev[pytorch]
|
||||
|
||||
# For Arc GPUs, you may need the compute runtime
|
||||
# Ubuntu/Debian:
|
||||
wget -qO - https://repositories.intel.com/graphics/intel-graphics.key | sudo apt-key add -
|
||||
sudo apt-add-repository 'deb [arch=amd64] https://repositories.intel.com/graphics/ubuntu focal main'
|
||||
sudo apt update
|
||||
sudo apt install intel-opencl-icd intel-level-zero-gpu level-zero
|
||||
|
||||
# Note: You may need a Whisper fork with OpenVINO support
|
||||
pip install git+https://github.com/openvinotoolkit/whisper.git
|
||||
```
|
||||
|
||||
## Troubleshooting
|
||||
|
||||
### Issue: "whisper: command not found"
|
||||
- Make sure Python's Scripts directory is in your PATH
|
||||
- On Windows: `C:\Users\[Username]\AppData\Local\Programs\Python\Python3X\Scripts`
|
||||
- On macOS/Linux: `~/.local/bin`
|
||||
|
||||
### Issue: GPU not detected
|
||||
- Use the "Detect GPU" button in the app's settings page
|
||||
- For NVIDIA: Ensure `nvidia-smi` command works
|
||||
- For AMD: Ensure `rocm-smi` command works
|
||||
- For Intel: Check `lspci | grep -i intel` shows your GPU
|
||||
|
||||
### Issue: CUDA out of memory
|
||||
- Try using a smaller Whisper model (tiny or base)
|
||||
- Reduce batch size or use CPU for very long audio
|
||||
|
||||
### Issue: Audio Format Not Supported
|
||||
Whisper works best with WAV, MP3, and M4A files. The app automatically converts audio to a compatible format.
|
||||
|
||||
## Alternative: Docker Installation
|
||||
|
||||
You can also run Whisper in a Docker container:
|
||||
|
||||
```bash
|
||||
docker pull openai/whisper
|
||||
docker run -it -v $(pwd):/app openai/whisper audio.mp3 --model base
|
||||
```
|
||||
@@ -0,0 +1,172 @@
|
||||
# Live Transcription App
|
||||
|
||||
A desktop application for live speech-to-text transcription with OBS compatibility. Supports multiple transcription backends including local/remote Whisper and ElevenLabs Scribe, with GPU acceleration for NVIDIA, AMD, and Intel GPUs.
|
||||
|
||||
## Features
|
||||
|
||||
- **Multiple Transcription Backends:**
|
||||
- Whisper (Local) - Run OpenAI's Whisper model locally
|
||||
- Whisper (Remote) - Use OpenAI's Whisper API
|
||||
- ElevenLabs (Scribe) - Use ElevenLabs' Scribe v1 speech-to-text API
|
||||
|
||||
- **OBS Integration:**
|
||||
- Transparent/frameless window mode
|
||||
- Always-on-top option
|
||||
- Customizable text display
|
||||
|
||||
- **Customization Options:**
|
||||
- Text duration control (1-60 seconds)
|
||||
- Text and background colors
|
||||
- Font selection (system fonts + Google Fonts)
|
||||
- Adjustable font size (16-72px)
|
||||
|
||||
- **Advanced Features:**
|
||||
- GPU acceleration (NVIDIA CUDA, AMD ROCm, Intel OpenVINO)
|
||||
- Ambient noise calibration
|
||||
- Persistent settings storage
|
||||
- Real-time audio level monitoring
|
||||
- Queue management for optimal performance
|
||||
- Configurable transcription speed (1-5 second chunks)
|
||||
|
||||
## Installation
|
||||
|
||||
1. Clone the repository:
|
||||
```bash
|
||||
git clone <repository-url>
|
||||
cd closed-caption-app
|
||||
```
|
||||
|
||||
2. Install dependencies:
|
||||
```bash
|
||||
npm install
|
||||
```
|
||||
|
||||
3. Install Whisper (for local transcription):
|
||||
- See [INSTALL_WHISPER.md](INSTALL_WHISPER.md) for detailed instructions
|
||||
|
||||
4. Configure environment:
|
||||
```bash
|
||||
cp .env.example .env
|
||||
# Edit .env with your API keys
|
||||
```
|
||||
|
||||
## Configuration
|
||||
|
||||
### Settings Page
|
||||
|
||||
Access the settings page by clicking the ⚙️ Settings button in the main interface. The settings page allows you to:
|
||||
|
||||
- Configure API keys for OpenAI and ElevenLabs
|
||||
- Select Whisper model size (tiny to large)
|
||||
- Enable GPU acceleration (auto-detects NVIDIA/AMD/Intel)
|
||||
- Calibrate ambient noise levels
|
||||
- Customize display settings (fonts, colors, duration)
|
||||
- Test API connections
|
||||
- Save settings persistently
|
||||
|
||||
### Environment Variables (Optional)
|
||||
|
||||
You can also edit `.env` file for initial configuration:
|
||||
|
||||
```env
|
||||
# For local Whisper
|
||||
WHISPER_MODEL=base
|
||||
|
||||
# For remote Whisper (OpenAI API)
|
||||
OPENAI_API_KEY=your_api_key_here
|
||||
|
||||
# For ElevenLabs speech-to-text
|
||||
ELEVENLABS_API_KEY=your_api_key_here
|
||||
|
||||
# Audio threshold for silence detection
|
||||
AUDIO_THRESHOLD=0.01
|
||||
```
|
||||
|
||||
Note: Settings configured through the UI will override environment variables.
|
||||
|
||||
## Usage
|
||||
|
||||
1. Start the application:
|
||||
```bash
|
||||
npm start
|
||||
```
|
||||
|
||||
2. Click ⚙️ Settings to configure:
|
||||
- API keys for remote services
|
||||
- GPU acceleration (if available)
|
||||
- Audio threshold and ambient noise calibration
|
||||
- Display preferences
|
||||
|
||||
3. Select your preferred transcription backend
|
||||
4. Configure display settings (color, font, duration)
|
||||
5. Click "Start Recording" to begin transcription
|
||||
6. Use "Open Display Window" for OBS capture
|
||||
|
||||
### Ambient Noise Calibration
|
||||
|
||||
1. Go to Settings → Audio Processing
|
||||
2. Click "Start Calibration"
|
||||
3. Stay quiet for 10 seconds
|
||||
4. The app will learn your room's ambient noise level
|
||||
5. Threshold will be automatically adjusted
|
||||
|
||||
### OBS Setup
|
||||
|
||||
1. Add a "Window Capture" source in OBS
|
||||
2. Select "Live Transcription" window
|
||||
3. Enable "Allow Transparency" in the source properties
|
||||
4. Position and scale as needed
|
||||
|
||||
### Keyboard Shortcuts
|
||||
|
||||
- `ESC` - Exit display mode (when in fullscreen)
|
||||
|
||||
## Development
|
||||
|
||||
Run in development mode:
|
||||
```bash
|
||||
npm run dev
|
||||
```
|
||||
|
||||
## Build
|
||||
|
||||
Build for distribution:
|
||||
```bash
|
||||
npm run build
|
||||
```
|
||||
|
||||
## System Requirements
|
||||
|
||||
- Node.js 16+
|
||||
- Python 3.8+ (for local Whisper)
|
||||
- ffmpeg (for audio processing)
|
||||
- Microphone access
|
||||
|
||||
## Documentation
|
||||
|
||||
- [Settings Guide](SETTINGS.md) - Detailed guide for all settings and configuration options
|
||||
- [Whisper Installation](INSTALL_WHISPER.md) - Complete guide for installing Whisper with GPU support
|
||||
|
||||
## Troubleshooting
|
||||
|
||||
### No transcription appearing
|
||||
- Check microphone permissions
|
||||
- Verify selected backend is configured (use Settings page)
|
||||
- Check audio level meter is responding to sound
|
||||
- Verify audio threshold isn't too high
|
||||
|
||||
### Whisper local not working
|
||||
- Ensure Whisper is installed: `whisper --help`
|
||||
- Check Python and ffmpeg are in PATH
|
||||
- Try a smaller model (tiny/base) for testing
|
||||
- Check GPU detection in Settings if using acceleration
|
||||
|
||||
### Window not transparent in OBS
|
||||
- Enable "Allow Transparency" in OBS source settings
|
||||
- Use the separate display window (Open Display Window button)
|
||||
|
||||
### High CPU/GPU usage
|
||||
- Use Settings to select smaller Whisper model
|
||||
- Increase transcription chunk size (3-5 seconds)
|
||||
- Calibrate ambient noise to reduce unnecessary processing
|
||||
- Check queue status for backlog
|
||||
+152
@@ -0,0 +1,152 @@
|
||||
# Settings Guide
|
||||
|
||||
This guide explains all the settings available in the Live Transcription App.
|
||||
|
||||
## Accessing Settings
|
||||
|
||||
Click the **⚙️ Settings** button in the top-right corner of the main window to open the settings page.
|
||||
|
||||
## Settings Overview
|
||||
|
||||
### API Configuration
|
||||
|
||||
#### OpenAI API Key
|
||||
- Required for using Whisper Remote transcription
|
||||
- Get your key from: https://platform.openai.com/api-keys
|
||||
- Stored securely and never shown in plain text after saving
|
||||
|
||||
#### ElevenLabs API Key
|
||||
- Required for using ElevenLabs Scribe transcription
|
||||
- Get your key from: https://elevenlabs.io/api
|
||||
- Stored securely and never shown in plain text after saving
|
||||
|
||||
### Audio Processing
|
||||
|
||||
#### Whisper Model
|
||||
Select the Whisper model size based on your needs:
|
||||
- **Tiny**: Fastest, least accurate (39M parameters)
|
||||
- **Base**: Good balance (74M parameters) - Default
|
||||
- **Small**: More accurate (244M parameters)
|
||||
- **Medium**: High accuracy (769M parameters)
|
||||
- **Large**: Highest accuracy, slowest (1550M parameters)
|
||||
|
||||
#### GPU Acceleration
|
||||
Choose your GPU acceleration method:
|
||||
- **Disabled (CPU only)**: Uses CPU for processing
|
||||
- **NVIDIA CUDA**: For NVIDIA GPUs (RTX, GTX, etc.)
|
||||
- **AMD ROCm**: For AMD GPUs (Linux only)
|
||||
- **Intel OpenVINO**: For Intel Arc and integrated GPUs
|
||||
|
||||
Use the **Detect GPU** button to automatically detect and select the best option.
|
||||
|
||||
#### Audio Threshold
|
||||
- Controls sensitivity to sound (0.001 to 0.1)
|
||||
- Lower values = more sensitive (picks up quiet sounds)
|
||||
- Higher values = less sensitive (only loud sounds)
|
||||
- Default: 0.01
|
||||
|
||||
#### Transcription Speed
|
||||
Controls how often audio chunks are processed:
|
||||
- **Very Fast (1s)**: Near real-time, may cut off words
|
||||
- **Fast (2s)**: Good responsiveness - Default
|
||||
- **Balanced (3s)**: Better for complete phrases
|
||||
- **Quality (5s)**: Best accuracy for long sentences
|
||||
|
||||
#### Ambient Noise Calibration
|
||||
1. Click **Start Calibration**
|
||||
2. Stay quiet for 10 seconds
|
||||
3. The app measures your room's background noise
|
||||
4. Automatically sets optimal threshold
|
||||
5. Shows current ambient level when calibrated
|
||||
|
||||
### Display Settings
|
||||
|
||||
#### Default Backend
|
||||
Choose which transcription service to use by default:
|
||||
- Whisper (Local)
|
||||
- Whisper (Remote)
|
||||
- ElevenLabs (Scribe)
|
||||
|
||||
#### Text Duration
|
||||
How long transcribed text stays visible (1-60 seconds)
|
||||
|
||||
#### Text Color
|
||||
Color of the transcribed text
|
||||
|
||||
#### Background Color
|
||||
Background color of the transcription area
|
||||
|
||||
#### Font Family
|
||||
Choose from system fonts or Google Fonts
|
||||
|
||||
#### Font Size
|
||||
Adjust text size from 16px to 72px
|
||||
|
||||
### Server Settings
|
||||
|
||||
#### Server Port
|
||||
- Default: 3000
|
||||
- Change if port conflicts with other applications
|
||||
- Requires app restart after changing
|
||||
|
||||
## Settings Storage
|
||||
|
||||
Settings are stored in:
|
||||
- **Windows**: `%USERPROFILE%\.live-transcription\settings.json`
|
||||
- **macOS**: `~/.live-transcription/settings.json`
|
||||
- **Linux**: `~/.live-transcription/settings.json`
|
||||
|
||||
## Testing Features
|
||||
|
||||
### Test API Connections
|
||||
- Tests OpenAI API key validity
|
||||
- Tests ElevenLabs API key validity
|
||||
- Shows success/failure status
|
||||
|
||||
### Detect GPU
|
||||
- Automatically detects available GPUs
|
||||
- Shows NVIDIA CUDA availability
|
||||
- Shows AMD ROCm availability
|
||||
- Shows Intel GPU availability
|
||||
- Recommends best option
|
||||
|
||||
## Best Practices
|
||||
|
||||
1. **For Best Performance**:
|
||||
- Use GPU acceleration if available
|
||||
- Choose appropriate model size for your hardware
|
||||
- Use "Fast (2s)" transcription speed
|
||||
|
||||
2. **For Best Accuracy**:
|
||||
- Use larger Whisper models (Small/Medium/Large)
|
||||
- Use "Quality (5s)" transcription speed
|
||||
- Calibrate ambient noise in your environment
|
||||
|
||||
3. **For Low-End Hardware**:
|
||||
- Use Tiny or Base models
|
||||
- Disable GPU if causing issues
|
||||
- Increase audio threshold to reduce processing
|
||||
|
||||
## Troubleshooting
|
||||
|
||||
### Settings Not Saving
|
||||
- Check write permissions for settings directory
|
||||
- Try "Reset to Defaults" if settings are corrupted
|
||||
|
||||
### GPU Not Detected
|
||||
- Ensure GPU drivers are installed
|
||||
- For NVIDIA: Install CUDA toolkit
|
||||
- For AMD: Install ROCm (Linux only)
|
||||
- For Intel: Install OpenVINO runtime
|
||||
|
||||
### High CPU Usage
|
||||
- Enable GPU acceleration if available
|
||||
- Use smaller Whisper model
|
||||
- Increase audio threshold
|
||||
- Use longer transcription chunks
|
||||
|
||||
### API Keys Not Working
|
||||
- Use "Test API Connections" to verify
|
||||
- Check for typos or extra spaces
|
||||
- Ensure API keys have correct permissions
|
||||
- Check API usage limits
|
||||
@@ -0,0 +1,33 @@
|
||||
{
|
||||
"name": "live-transcription-app",
|
||||
"version": "1.0.0",
|
||||
"description": "Live transcription app with multiple backend support",
|
||||
"main": "src/main.js",
|
||||
"scripts": {
|
||||
"start": "electron .",
|
||||
"dev": "electron . --dev",
|
||||
"build": "electron-builder",
|
||||
"server": "node src/server/index.js"
|
||||
},
|
||||
"keywords": ["transcription", "whisper", "elevenlabs", "obs"],
|
||||
"author": "",
|
||||
"license": "MIT",
|
||||
"devDependencies": {
|
||||
"@types/node": "^20.11.0",
|
||||
"electron": "^28.1.0",
|
||||
"electron-builder": "^24.9.1"
|
||||
},
|
||||
"dependencies": {
|
||||
"axios": "^1.6.5",
|
||||
"cors": "^2.8.5",
|
||||
"dotenv": "^16.3.1",
|
||||
"express": "^4.18.2",
|
||||
"form-data": "^4.0.0",
|
||||
"multer": "^2.0.0-rc.4",
|
||||
"node-record-lpcm16": "^1.0.1",
|
||||
"openai": "^4.24.1",
|
||||
"socket.io": "^4.7.4",
|
||||
"socket.io-client": "^4.7.4",
|
||||
"ws": "^8.16.0"
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,181 @@
|
||||
* {
|
||||
margin: 0;
|
||||
padding: 0;
|
||||
box-sizing: border-box;
|
||||
}
|
||||
|
||||
body {
|
||||
font-family: Arial, sans-serif;
|
||||
background: #000;
|
||||
color: #fff;
|
||||
overflow: hidden;
|
||||
}
|
||||
|
||||
#app {
|
||||
height: 100vh;
|
||||
display: flex;
|
||||
flex-direction: column;
|
||||
}
|
||||
|
||||
.controls {
|
||||
background: #1a1a1a;
|
||||
padding: 20px;
|
||||
border-bottom: 1px solid #333;
|
||||
}
|
||||
|
||||
.header {
|
||||
display: flex;
|
||||
justify-content: space-between;
|
||||
align-items: center;
|
||||
margin-bottom: 20px;
|
||||
}
|
||||
|
||||
.header h3 {
|
||||
color: #fff;
|
||||
font-size: 18px;
|
||||
margin: 0;
|
||||
}
|
||||
|
||||
.header a {
|
||||
text-decoration: none;
|
||||
}
|
||||
|
||||
.header .btn {
|
||||
margin: 0;
|
||||
}
|
||||
|
||||
.control-panel {
|
||||
display: grid;
|
||||
grid-template-columns: repeat(auto-fit, minmax(250px, 1fr));
|
||||
gap: 15px;
|
||||
-webkit-app-region: no-drag;
|
||||
}
|
||||
|
||||
.control-group {
|
||||
display: flex;
|
||||
flex-direction: column;
|
||||
gap: 5px;
|
||||
}
|
||||
|
||||
.control-group label {
|
||||
font-size: 14px;
|
||||
color: #ccc;
|
||||
}
|
||||
|
||||
.control-group input,
|
||||
.control-group select {
|
||||
padding: 8px;
|
||||
border: 1px solid #333;
|
||||
background: #2a2a2a;
|
||||
color: #fff;
|
||||
border-radius: 4px;
|
||||
}
|
||||
|
||||
.control-group input[type="color"] {
|
||||
height: 40px;
|
||||
cursor: pointer;
|
||||
}
|
||||
|
||||
.control-group input[type="range"] {
|
||||
width: 100%;
|
||||
}
|
||||
|
||||
#font-size-display {
|
||||
color: #ccc;
|
||||
font-size: 14px;
|
||||
}
|
||||
|
||||
.btn {
|
||||
padding: 10px 20px;
|
||||
border: none;
|
||||
border-radius: 4px;
|
||||
cursor: pointer;
|
||||
font-size: 14px;
|
||||
transition: opacity 0.2s;
|
||||
}
|
||||
|
||||
.btn:disabled {
|
||||
opacity: 0.5;
|
||||
cursor: not-allowed;
|
||||
}
|
||||
|
||||
.btn-primary {
|
||||
background: #4CAF50;
|
||||
color: white;
|
||||
}
|
||||
|
||||
.btn-primary:hover:not(:disabled) {
|
||||
background: #45a049;
|
||||
}
|
||||
|
||||
.btn-secondary {
|
||||
background: #f44336;
|
||||
color: white;
|
||||
}
|
||||
|
||||
.btn-secondary:hover:not(:disabled) {
|
||||
background: #da190b;
|
||||
}
|
||||
|
||||
.btn-small {
|
||||
padding: 6px 12px;
|
||||
font-size: 12px;
|
||||
}
|
||||
|
||||
.transcription-display {
|
||||
flex: 1;
|
||||
display: flex;
|
||||
align-items: center;
|
||||
justify-content: center;
|
||||
padding: 40px;
|
||||
background: #000;
|
||||
position: relative;
|
||||
}
|
||||
|
||||
#transcription-text {
|
||||
text-align: center;
|
||||
font-size: 32px;
|
||||
color: #fff;
|
||||
line-height: 1.5;
|
||||
max-width: 90%;
|
||||
word-wrap: break-word;
|
||||
transition: opacity 0.3s ease;
|
||||
}
|
||||
|
||||
.fade-out {
|
||||
opacity: 0 !important;
|
||||
}
|
||||
|
||||
.audio-level-container {
|
||||
display: flex;
|
||||
align-items: center;
|
||||
gap: 10px;
|
||||
}
|
||||
|
||||
.audio-level-bar {
|
||||
flex: 1;
|
||||
height: 20px;
|
||||
background: #333;
|
||||
border-radius: 10px;
|
||||
overflow: hidden;
|
||||
position: relative;
|
||||
}
|
||||
|
||||
.audio-level-fill {
|
||||
height: 100%;
|
||||
background: linear-gradient(90deg, #4CAF50 0%, #FFC107 50%, #f44336 100%);
|
||||
width: 0%;
|
||||
transition: width 0.1s ease;
|
||||
}
|
||||
|
||||
#audio-level-text {
|
||||
color: #ccc;
|
||||
font-size: 12px;
|
||||
min-width: 30px;
|
||||
}
|
||||
|
||||
#queue-status {
|
||||
color: #ccc;
|
||||
font-size: 12px;
|
||||
font-family: monospace;
|
||||
}
|
||||
@@ -0,0 +1,56 @@
|
||||
<!DOCTYPE html>
|
||||
<html lang="en">
|
||||
<head>
|
||||
<meta charset="UTF-8">
|
||||
<meta name="viewport" content="width=device-width, initial-scale=1.0">
|
||||
<title>Live Transcription Display</title>
|
||||
<link rel="preconnect" href="https://fonts.googleapis.com">
|
||||
<link rel="preconnect" href="https://fonts.gstatic.com" crossorigin>
|
||||
<link id="google-font-link" rel="stylesheet" href="">
|
||||
<style>
|
||||
* {
|
||||
margin: 0;
|
||||
padding: 0;
|
||||
box-sizing: border-box;
|
||||
}
|
||||
|
||||
body {
|
||||
background: transparent;
|
||||
overflow: hidden;
|
||||
-webkit-app-region: drag;
|
||||
user-select: none;
|
||||
}
|
||||
|
||||
#display-container {
|
||||
width: 100vw;
|
||||
height: 100vh;
|
||||
display: flex;
|
||||
align-items: center;
|
||||
justify-content: center;
|
||||
padding: 20px;
|
||||
background: #000;
|
||||
}
|
||||
|
||||
#transcription-text {
|
||||
text-align: center;
|
||||
font-size: 32px;
|
||||
color: #fff;
|
||||
line-height: 1.5;
|
||||
word-wrap: break-word;
|
||||
transition: opacity 0.3s ease;
|
||||
font-family: Arial, sans-serif;
|
||||
}
|
||||
|
||||
.fade-out {
|
||||
opacity: 0 !important;
|
||||
}
|
||||
</style>
|
||||
</head>
|
||||
<body>
|
||||
<div id="display-container">
|
||||
<div id="transcription-text"></div>
|
||||
</div>
|
||||
|
||||
<script src="js/display.js"></script>
|
||||
</body>
|
||||
</html>
|
||||
@@ -0,0 +1,118 @@
|
||||
<!DOCTYPE html>
|
||||
<html lang="en">
|
||||
<head>
|
||||
<meta charset="UTF-8">
|
||||
<meta name="viewport" content="width=device-width, initial-scale=1.0">
|
||||
<meta http-equiv="Content-Security-Policy" content="default-src 'self' http://localhost:3000 https://fonts.googleapis.com https://fonts.gstatic.com; script-src 'self' 'unsafe-inline' http://localhost:3000; style-src 'self' 'unsafe-inline' https://fonts.googleapis.com; font-src 'self' https://fonts.gstatic.com; connect-src 'self' http://localhost:3000 ws://localhost:3000">
|
||||
<title>Live Transcription</title>
|
||||
<link rel="stylesheet" href="css/styles.css">
|
||||
<link rel="preconnect" href="https://fonts.googleapis.com">
|
||||
<link rel="preconnect" href="https://fonts.gstatic.com" crossorigin>
|
||||
<link id="google-font-link" rel="stylesheet" href="">
|
||||
</head>
|
||||
<body>
|
||||
<div id="app">
|
||||
<div id="controls" class="controls">
|
||||
<div class="header">
|
||||
<h3>Live Transcription</h3>
|
||||
<button id="settings-btn" class="btn btn-small">⚙️ Settings</button>
|
||||
</div>
|
||||
|
||||
<div class="control-panel">
|
||||
<div class="control-group">
|
||||
<label>Backend:</label>
|
||||
<select id="backend-select">
|
||||
<option value="whisper-local">Whisper (Local)</option>
|
||||
<option value="whisper-remote">Whisper (Remote)</option>
|
||||
<option value="elevenlabs">ElevenLabs (Scribe)</option>
|
||||
</select>
|
||||
</div>
|
||||
|
||||
<div class="control-group">
|
||||
<label>Text Duration (seconds):</label>
|
||||
<input type="number" id="duration-input" min="1" max="60" value="5">
|
||||
</div>
|
||||
|
||||
<div class="control-group">
|
||||
<label>Transcription Speed:</label>
|
||||
<select id="chunk-size-select">
|
||||
<option value="1000">Very Fast (1s chunks)</option>
|
||||
<option value="2000" selected>Fast (2s chunks)</option>
|
||||
<option value="3000">Balanced (3s chunks)</option>
|
||||
<option value="5000">Quality (5s chunks)</option>
|
||||
</select>
|
||||
</div>
|
||||
|
||||
<div class="control-group">
|
||||
<label>Text Color:</label>
|
||||
<input type="color" id="color-input" value="#ffffff">
|
||||
</div>
|
||||
|
||||
<div class="control-group">
|
||||
<label>Background Color:</label>
|
||||
<input type="color" id="bg-color-input" value="#000000">
|
||||
</div>
|
||||
|
||||
<div class="control-group">
|
||||
<label>Font Family:</label>
|
||||
<select id="font-select">
|
||||
<optgroup label="System Fonts">
|
||||
<option value="Arial, sans-serif">Arial</option>
|
||||
<option value="'Times New Roman', serif">Times New Roman</option>
|
||||
<option value="'Courier New', monospace">Courier New</option>
|
||||
<option value="Georgia, serif">Georgia</option>
|
||||
<option value="Verdana, sans-serif">Verdana</option>
|
||||
<option value="'Trebuchet MS', sans-serif">Trebuchet MS</option>
|
||||
</optgroup>
|
||||
<optgroup label="Google Fonts">
|
||||
<option value="'Roboto', sans-serif" data-google="Roboto">Roboto</option>
|
||||
<option value="'Open Sans', sans-serif" data-google="Open+Sans">Open Sans</option>
|
||||
<option value="'Lato', sans-serif" data-google="Lato">Lato</option>
|
||||
<option value="'Montserrat', sans-serif" data-google="Montserrat">Montserrat</option>
|
||||
<option value="'Poppins', sans-serif" data-google="Poppins">Poppins</option>
|
||||
<option value="'Raleway', sans-serif" data-google="Raleway">Raleway</option>
|
||||
</optgroup>
|
||||
</select>
|
||||
</div>
|
||||
|
||||
<div class="control-group">
|
||||
<label>Font Size:</label>
|
||||
<input type="range" id="font-size-input" min="16" max="72" value="32">
|
||||
<span id="font-size-display">32px</span>
|
||||
</div>
|
||||
|
||||
<div class="control-group">
|
||||
<button id="start-btn" class="btn btn-primary">Start Recording</button>
|
||||
<button id="stop-btn" class="btn btn-secondary" disabled>Stop Recording</button>
|
||||
</div>
|
||||
|
||||
<div class="control-group">
|
||||
<label>Audio Level:</label>
|
||||
<div class="audio-level-container">
|
||||
<div id="audio-level-bar" class="audio-level-bar">
|
||||
<div id="audio-level-fill" class="audio-level-fill"></div>
|
||||
</div>
|
||||
<span id="audio-level-text">0%</span>
|
||||
</div>
|
||||
</div>
|
||||
|
||||
<div class="control-group">
|
||||
<label>Queue Status:</label>
|
||||
<div id="queue-status">Queue: 0, Running: 0</div>
|
||||
</div>
|
||||
|
||||
<div class="control-group">
|
||||
<button id="toggle-display-btn" class="btn">Open Display Window</button>
|
||||
</div>
|
||||
</div>
|
||||
</div>
|
||||
|
||||
<div id="transcription-display" class="transcription-display">
|
||||
<div id="transcription-text"></div>
|
||||
</div>
|
||||
</div>
|
||||
|
||||
<script src="http://localhost:3000/socket.io/socket.io.js"></script>
|
||||
<script src="js/app.js"></script>
|
||||
</body>
|
||||
</html>
|
||||
@@ -0,0 +1,444 @@
|
||||
let socket;
|
||||
let mediaRecorder;
|
||||
let audioChunks = [];
|
||||
let isRecording = false;
|
||||
let displayWindowOpen = false;
|
||||
let currentTimeout;
|
||||
|
||||
// Store current settings
|
||||
let displaySettings = {
|
||||
textColor: '#ffffff',
|
||||
bgColor: '#000000',
|
||||
fontFamily: 'Arial, sans-serif',
|
||||
fontSize: '32px',
|
||||
duration: 5,
|
||||
googleFont: null
|
||||
};
|
||||
|
||||
// Wait for DOM to be fully loaded
|
||||
document.addEventListener('DOMContentLoaded', () => {
|
||||
// Initialize socket connection
|
||||
try {
|
||||
socket = io('http://localhost:3000');
|
||||
|
||||
socket.on('connect', () => {
|
||||
console.log('Connected to server');
|
||||
});
|
||||
|
||||
socket.on('connect_error', (error) => {
|
||||
console.error('Socket connection error:', error);
|
||||
});
|
||||
} catch (error) {
|
||||
console.error('Failed to initialize socket:', error);
|
||||
}
|
||||
|
||||
const startBtn = document.getElementById('start-btn');
|
||||
const stopBtn = document.getElementById('stop-btn');
|
||||
const backendSelect = document.getElementById('backend-select');
|
||||
const durationInput = document.getElementById('duration-input');
|
||||
const colorInput = document.getElementById('color-input');
|
||||
const bgColorInput = document.getElementById('bg-color-input');
|
||||
const fontSelect = document.getElementById('font-select');
|
||||
const fontSizeInput = document.getElementById('font-size-input');
|
||||
const fontSizeDisplay = document.getElementById('font-size-display');
|
||||
const transcriptionText = document.getElementById('transcription-text');
|
||||
const transcriptionDisplay = document.getElementById('transcription-display');
|
||||
const toggleDisplayBtn = document.getElementById('toggle-display-btn');
|
||||
const googleFontLink = document.getElementById('google-font-link');
|
||||
const audioLevelFill = document.getElementById('audio-level-fill');
|
||||
const audioLevelText = document.getElementById('audio-level-text');
|
||||
const queueStatus = document.getElementById('queue-status');
|
||||
const chunkSizeSelect = document.getElementById('chunk-size-select');
|
||||
const settingsBtn = document.getElementById('settings-btn');
|
||||
|
||||
// Load available backends
|
||||
async function loadBackends() {
|
||||
try {
|
||||
const response = await fetch('http://localhost:3000/api/backends');
|
||||
const data = await response.json();
|
||||
|
||||
backendSelect.innerHTML = '';
|
||||
data.backends.forEach(backend => {
|
||||
const option = document.createElement('option');
|
||||
option.value = backend.id;
|
||||
option.textContent = backend.name + (backend.available ? '' : ' (Unavailable)');
|
||||
option.disabled = !backend.available;
|
||||
backendSelect.appendChild(option);
|
||||
});
|
||||
|
||||
backendSelect.value = data.current;
|
||||
} catch (error) {
|
||||
console.error('Failed to load backends:', error);
|
||||
}
|
||||
}
|
||||
|
||||
// Initialize
|
||||
loadBackends();
|
||||
loadAppSettings();
|
||||
|
||||
// Update queue status periodically
|
||||
async function updateQueueStatus() {
|
||||
try {
|
||||
const response = await fetch('http://localhost:3000/api/queue-status');
|
||||
const status = await response.json();
|
||||
queueStatus.textContent = `Queue: ${status.queueLength}, Running: ${status.runningJobs}`;
|
||||
} catch (error) {
|
||||
// Silently fail - not critical
|
||||
}
|
||||
}
|
||||
|
||||
setInterval(updateQueueStatus, 2000); // Update every 2 seconds
|
||||
|
||||
// Save app setting
|
||||
async function saveAppSetting(key, value) {
|
||||
try {
|
||||
await fetch('http://localhost:3000/api/settings', {
|
||||
method: 'POST',
|
||||
headers: { 'Content-Type': 'application/json' },
|
||||
body: JSON.stringify({ [key]: value })
|
||||
});
|
||||
} catch (error) {
|
||||
console.warn('Could not save app setting:', key, error);
|
||||
}
|
||||
}
|
||||
|
||||
// Load app settings
|
||||
async function loadAppSettings() {
|
||||
try {
|
||||
const response = await fetch('http://localhost:3000/api/settings');
|
||||
const settings = await response.json();
|
||||
|
||||
// Apply settings to UI elements
|
||||
if (settings.selectedBackend) {
|
||||
backendSelect.value = settings.selectedBackend;
|
||||
}
|
||||
if (settings.textDuration) {
|
||||
durationInput.value = settings.textDuration;
|
||||
}
|
||||
if (settings.textColor) {
|
||||
colorInput.value = settings.textColor;
|
||||
}
|
||||
if (settings.backgroundColor) {
|
||||
bgColorInput.value = settings.backgroundColor;
|
||||
}
|
||||
if (settings.fontFamily) {
|
||||
fontSelect.value = settings.fontFamily;
|
||||
}
|
||||
if (settings.fontSize) {
|
||||
fontSizeInput.value = settings.fontSize;
|
||||
fontSizeDisplay.textContent = settings.fontSize + 'px';
|
||||
}
|
||||
if (settings.chunkSize) {
|
||||
chunkSizeSelect.value = settings.chunkSize;
|
||||
}
|
||||
|
||||
// Update display settings
|
||||
updateDisplaySettings();
|
||||
|
||||
console.log('App settings loaded');
|
||||
} catch (error) {
|
||||
console.warn('Could not load app settings, using defaults:', error);
|
||||
}
|
||||
}
|
||||
|
||||
// Update display settings and notify display window
|
||||
function updateDisplaySettings() {
|
||||
displaySettings = {
|
||||
textColor: colorInput.value,
|
||||
bgColor: bgColorInput.value,
|
||||
fontFamily: fontSelect.value,
|
||||
fontSize: fontSizeInput.value + 'px',
|
||||
duration: parseInt(durationInput.value),
|
||||
googleFont: fontSelect.options[fontSelect.selectedIndex].getAttribute('data-google')
|
||||
};
|
||||
|
||||
// Apply to local display
|
||||
transcriptionText.style.color = displaySettings.textColor;
|
||||
transcriptionText.style.fontSize = displaySettings.fontSize;
|
||||
transcriptionText.style.fontFamily = displaySettings.fontFamily;
|
||||
transcriptionDisplay.style.backgroundColor = displaySettings.bgColor;
|
||||
|
||||
// Update Google Font link
|
||||
if (displaySettings.googleFont) {
|
||||
googleFontLink.href = `https://fonts.googleapis.com/css2?family=${displaySettings.googleFont}&display=swap`;
|
||||
}
|
||||
|
||||
// Send to display window
|
||||
if (window.electronAPI) {
|
||||
window.electronAPI.updateDisplaySettings(displaySettings);
|
||||
}
|
||||
}
|
||||
|
||||
// Font selection handling with auto-save
|
||||
fontSelect.addEventListener('change', () => {
|
||||
updateDisplaySettings();
|
||||
saveAppSetting('fontFamily', fontSelect.value);
|
||||
const selectedOption = fontSelect.options[fontSelect.selectedIndex];
|
||||
const googleFont = selectedOption.getAttribute('data-google');
|
||||
saveAppSetting('googleFont', googleFont);
|
||||
});
|
||||
|
||||
colorInput.addEventListener('change', () => {
|
||||
updateDisplaySettings();
|
||||
saveAppSetting('textColor', colorInput.value);
|
||||
});
|
||||
|
||||
bgColorInput.addEventListener('change', () => {
|
||||
updateDisplaySettings();
|
||||
saveAppSetting('backgroundColor', bgColorInput.value);
|
||||
});
|
||||
|
||||
durationInput.addEventListener('change', () => {
|
||||
updateDisplaySettings();
|
||||
saveAppSetting('textDuration', parseInt(durationInput.value));
|
||||
});
|
||||
|
||||
fontSizeInput.addEventListener('input', (e) => {
|
||||
fontSizeDisplay.textContent = e.target.value + 'px';
|
||||
updateDisplaySettings();
|
||||
saveAppSetting('fontSize', parseInt(e.target.value));
|
||||
});
|
||||
|
||||
chunkSizeSelect.addEventListener('change', () => {
|
||||
saveAppSetting('chunkSize', parseInt(chunkSizeSelect.value));
|
||||
});
|
||||
|
||||
// Backend selection
|
||||
backendSelect.addEventListener('change', async (e) => {
|
||||
try {
|
||||
await fetch('http://localhost:3000/api/backend', {
|
||||
method: 'POST',
|
||||
headers: { 'Content-Type': 'application/json' },
|
||||
body: JSON.stringify({ backend: e.target.value })
|
||||
});
|
||||
|
||||
// Save to settings
|
||||
saveAppSetting('selectedBackend', e.target.value);
|
||||
} catch (error) {
|
||||
console.error('Failed to switch backend:', error);
|
||||
}
|
||||
});
|
||||
|
||||
// Display window toggle
|
||||
toggleDisplayBtn.addEventListener('click', async () => {
|
||||
if (window.electronAPI) {
|
||||
displayWindowOpen = await window.electronAPI.toggleDisplayWindow();
|
||||
toggleDisplayBtn.textContent = displayWindowOpen ? 'Close Display Window' : 'Open Display Window';
|
||||
|
||||
// Send current settings to the new window
|
||||
if (displayWindowOpen) {
|
||||
setTimeout(() => {
|
||||
window.electronAPI.updateDisplaySettings(displaySettings);
|
||||
}, 500);
|
||||
}
|
||||
}
|
||||
});
|
||||
|
||||
// Listen for display window close event
|
||||
if (window.electronAPI) {
|
||||
window.electronAPI.onDisplayWindowClosed(() => {
|
||||
displayWindowOpen = false;
|
||||
toggleDisplayBtn.textContent = 'Open Display Window';
|
||||
});
|
||||
}
|
||||
|
||||
// Socket event listeners
|
||||
if (socket) {
|
||||
socket.on('transcription', (data) => {
|
||||
displayTranscription(data.text);
|
||||
|
||||
// Also send to display window via IPC
|
||||
if (window.electronAPI && displayWindowOpen) {
|
||||
window.electronAPI.sendTranscription(data);
|
||||
}
|
||||
});
|
||||
}
|
||||
|
||||
function displayTranscription(text) {
|
||||
transcriptionText.textContent = text;
|
||||
transcriptionText.classList.remove('fade-out');
|
||||
|
||||
if (currentTimeout) {
|
||||
clearTimeout(currentTimeout);
|
||||
}
|
||||
|
||||
const duration = displaySettings.duration * 1000;
|
||||
currentTimeout = setTimeout(() => {
|
||||
transcriptionText.classList.add('fade-out');
|
||||
setTimeout(() => {
|
||||
transcriptionText.textContent = '';
|
||||
}, 300);
|
||||
}, duration);
|
||||
}
|
||||
|
||||
// Audio level monitoring
|
||||
let audioContext;
|
||||
let analyser;
|
||||
let microphone;
|
||||
let dataArray;
|
||||
let audioLevelInterval;
|
||||
|
||||
function startAudioLevelMonitoring(stream) {
|
||||
audioContext = new (window.AudioContext || window.webkitAudioContext)();
|
||||
analyser = audioContext.createAnalyser();
|
||||
microphone = audioContext.createMediaStreamSource(stream);
|
||||
|
||||
// Configure analyser for better responsiveness
|
||||
analyser.fftSize = 512;
|
||||
analyser.smoothingTimeConstant = 0.3;
|
||||
analyser.minDecibels = -90;
|
||||
analyser.maxDecibels = -10;
|
||||
|
||||
const bufferLength = analyser.frequencyBinCount;
|
||||
dataArray = new Uint8Array(bufferLength);
|
||||
|
||||
microphone.connect(analyser);
|
||||
|
||||
function updateAudioLevel() {
|
||||
if (!isRecording) return;
|
||||
|
||||
analyser.getByteTimeDomainData(dataArray);
|
||||
|
||||
// Calculate RMS of time domain data (more accurate for level detection)
|
||||
let sum = 0;
|
||||
for (let i = 0; i < bufferLength; i++) {
|
||||
const sample = (dataArray[i] - 128) / 128; // Convert to -1 to 1 range
|
||||
sum += sample * sample;
|
||||
}
|
||||
const rms = Math.sqrt(sum / bufferLength);
|
||||
|
||||
// Apply some smoothing and scaling
|
||||
const level = Math.min(rms * 10, 1); // Scale up and cap at 1
|
||||
|
||||
// Update UI
|
||||
const percentage = Math.round(level * 100);
|
||||
audioLevelFill.style.width = `${percentage}%`;
|
||||
audioLevelText.textContent = `${percentage}%`;
|
||||
|
||||
// Store current level for threshold checking
|
||||
window.currentAudioLevel = level;
|
||||
}
|
||||
|
||||
// Update at 60fps for smooth animation
|
||||
audioLevelInterval = setInterval(updateAudioLevel, 16);
|
||||
}
|
||||
|
||||
// Audio recording
|
||||
async function startRecording() {
|
||||
try {
|
||||
const stream = await navigator.mediaDevices.getUserMedia({ audio: true });
|
||||
|
||||
// Start audio level monitoring
|
||||
startAudioLevelMonitoring(stream);
|
||||
|
||||
const options = { mimeType: 'audio/webm' };
|
||||
if (!MediaRecorder.isTypeSupported(options.mimeType)) {
|
||||
options.mimeType = 'audio/ogg';
|
||||
if (!MediaRecorder.isTypeSupported(options.mimeType)) {
|
||||
options.mimeType = '';
|
||||
}
|
||||
}
|
||||
|
||||
mediaRecorder = new MediaRecorder(stream, options);
|
||||
|
||||
mediaRecorder.ondataavailable = (event) => {
|
||||
if (event.data.size > 0) {
|
||||
audioChunks.push(event.data);
|
||||
}
|
||||
};
|
||||
|
||||
mediaRecorder.onstop = async () => {
|
||||
const audioBlob = new Blob(audioChunks, { type: 'audio/webm' });
|
||||
audioChunks = [];
|
||||
await sendAudioForTranscription(audioBlob);
|
||||
};
|
||||
|
||||
mediaRecorder.start();
|
||||
isRecording = true;
|
||||
|
||||
// Auto-stop and restart recording based on selected chunk size
|
||||
const chunkSize = parseInt(chunkSizeSelect.value);
|
||||
const recordingInterval = setInterval(() => {
|
||||
if (isRecording && mediaRecorder.state === 'recording') {
|
||||
mediaRecorder.stop();
|
||||
setTimeout(() => {
|
||||
if (isRecording) {
|
||||
audioChunks = [];
|
||||
mediaRecorder.start();
|
||||
}
|
||||
}, 100);
|
||||
} else if (!isRecording) {
|
||||
clearInterval(recordingInterval);
|
||||
}
|
||||
}, chunkSize);
|
||||
|
||||
startBtn.disabled = true;
|
||||
stopBtn.disabled = false;
|
||||
} catch (error) {
|
||||
console.error('Error starting recording:', error);
|
||||
alert('Failed to start recording. Please check microphone permissions.');
|
||||
}
|
||||
}
|
||||
|
||||
function stopRecording() {
|
||||
if (mediaRecorder && mediaRecorder.state !== 'inactive') {
|
||||
mediaRecorder.stop();
|
||||
mediaRecorder.stream.getTracks().forEach(track => track.stop());
|
||||
}
|
||||
isRecording = false;
|
||||
startBtn.disabled = false;
|
||||
stopBtn.disabled = true;
|
||||
|
||||
// Stop audio level monitoring
|
||||
if (audioLevelInterval) {
|
||||
clearInterval(audioLevelInterval);
|
||||
audioLevelInterval = null;
|
||||
}
|
||||
if (audioContext) {
|
||||
audioContext.close();
|
||||
audioContext = null;
|
||||
}
|
||||
|
||||
// Reset audio level display
|
||||
audioLevelFill.style.width = '0%';
|
||||
audioLevelText.textContent = '0%';
|
||||
}
|
||||
|
||||
async function sendAudioForTranscription(audioBlob) {
|
||||
const formData = new FormData();
|
||||
formData.append('audio', audioBlob, 'audio.webm');
|
||||
|
||||
try {
|
||||
const response = await fetch('http://localhost:3000/api/transcribe', {
|
||||
method: 'POST',
|
||||
body: formData
|
||||
});
|
||||
|
||||
if (!response.ok) {
|
||||
const error = await response.json();
|
||||
console.error('Transcription error:', error);
|
||||
}
|
||||
} catch (error) {
|
||||
console.error('Failed to send audio:', error);
|
||||
}
|
||||
}
|
||||
|
||||
// Event listeners
|
||||
startBtn.addEventListener('click', () => {
|
||||
console.log('Start button clicked');
|
||||
startRecording();
|
||||
});
|
||||
|
||||
stopBtn.addEventListener('click', () => {
|
||||
console.log('Stop button clicked');
|
||||
stopRecording();
|
||||
});
|
||||
|
||||
// Initial settings application
|
||||
updateDisplaySettings();
|
||||
|
||||
// Settings button
|
||||
settingsBtn.addEventListener('click', () => {
|
||||
window.location.href = 'settings.html';
|
||||
});
|
||||
});
|
||||
@@ -0,0 +1,56 @@
|
||||
let currentTimeout;
|
||||
let displaySettings = {
|
||||
textColor: '#ffffff',
|
||||
bgColor: '#000000',
|
||||
fontFamily: 'Arial, sans-serif',
|
||||
fontSize: '32px',
|
||||
duration: 5
|
||||
};
|
||||
|
||||
const transcriptionText = document.getElementById('transcription-text');
|
||||
const displayContainer = document.getElementById('display-container');
|
||||
const googleFontLink = document.getElementById('google-font-link');
|
||||
|
||||
// Apply settings
|
||||
function applySettings(settings) {
|
||||
displaySettings = { ...displaySettings, ...settings };
|
||||
|
||||
transcriptionText.style.color = displaySettings.textColor;
|
||||
transcriptionText.style.fontSize = displaySettings.fontSize;
|
||||
transcriptionText.style.fontFamily = displaySettings.fontFamily;
|
||||
displayContainer.style.backgroundColor = displaySettings.bgColor;
|
||||
|
||||
// Handle Google Fonts
|
||||
if (settings.googleFont) {
|
||||
googleFontLink.href = `https://fonts.googleapis.com/css2?family=${settings.googleFont}&display=swap`;
|
||||
}
|
||||
}
|
||||
|
||||
// Display transcription
|
||||
function displayTranscription(text) {
|
||||
transcriptionText.textContent = text;
|
||||
transcriptionText.classList.remove('fade-out');
|
||||
|
||||
if (currentTimeout) {
|
||||
clearTimeout(currentTimeout);
|
||||
}
|
||||
|
||||
const duration = displaySettings.duration * 1000;
|
||||
currentTimeout = setTimeout(() => {
|
||||
transcriptionText.classList.add('fade-out');
|
||||
setTimeout(() => {
|
||||
transcriptionText.textContent = '';
|
||||
}, 300);
|
||||
}, duration);
|
||||
}
|
||||
|
||||
// Listen for IPC messages
|
||||
if (window.electronAPI) {
|
||||
window.electronAPI.onTranscription((data) => {
|
||||
displayTranscription(data.text);
|
||||
});
|
||||
|
||||
window.electronAPI.onUpdateSettings((settings) => {
|
||||
applySettings(settings);
|
||||
});
|
||||
}
|
||||
@@ -0,0 +1,353 @@
|
||||
let isCalibrating = false;
|
||||
let calibrationData = [];
|
||||
let calibrationInterval;
|
||||
let audioContext;
|
||||
let analyser;
|
||||
let microphone;
|
||||
|
||||
document.addEventListener('DOMContentLoaded', () => {
|
||||
// Get all form elements
|
||||
const elements = {
|
||||
openaiKey: document.getElementById('openai-key'),
|
||||
elevenlabsKey: document.getElementById('elevenlabs-key'),
|
||||
whisperModel: document.getElementById('whisper-model'),
|
||||
gpuAcceleration: document.getElementById('gpu-acceleration'),
|
||||
audioThreshold: document.getElementById('audio-threshold'),
|
||||
thresholdValue: document.getElementById('threshold-value'),
|
||||
chunkSize: document.getElementById('chunk-size'),
|
||||
defaultBackend: document.getElementById('default-backend'),
|
||||
textDuration: document.getElementById('text-duration'),
|
||||
textColor: document.getElementById('text-color'),
|
||||
bgColor: document.getElementById('bg-color'),
|
||||
fontFamily: document.getElementById('font-family'),
|
||||
fontSize: document.getElementById('font-size'),
|
||||
fontSizeValue: document.getElementById('font-size-value'),
|
||||
serverPort: document.getElementById('server-port'),
|
||||
|
||||
// Calibration elements
|
||||
startCalibration: document.getElementById('start-calibration'),
|
||||
stopCalibration: document.getElementById('stop-calibration'),
|
||||
calibrationProgress: document.getElementById('calibration-progress'),
|
||||
ambientLevel: document.getElementById('ambient-level'),
|
||||
calibrationStatus: document.getElementById('calibration-status'),
|
||||
|
||||
// Action buttons
|
||||
saveSettings: document.getElementById('save-settings'),
|
||||
resetSettings: document.getElementById('reset-settings'),
|
||||
testConnection: document.getElementById('test-connection'),
|
||||
detectGpu: document.getElementById('detect-gpu'),
|
||||
statusMessage: document.getElementById('status-message')
|
||||
};
|
||||
|
||||
// Update threshold value display
|
||||
elements.audioThreshold.addEventListener('input', (e) => {
|
||||
elements.thresholdValue.textContent = parseFloat(e.target.value).toFixed(3);
|
||||
});
|
||||
|
||||
// Update font size value display
|
||||
elements.fontSize.addEventListener('input', (e) => {
|
||||
elements.fontSizeValue.textContent = e.target.value + 'px';
|
||||
});
|
||||
|
||||
// Load settings on page load
|
||||
loadSettings();
|
||||
|
||||
// Save settings
|
||||
elements.saveSettings.addEventListener('click', saveSettings);
|
||||
|
||||
// Reset settings
|
||||
elements.resetSettings.addEventListener('click', resetSettings);
|
||||
|
||||
// Test API connections
|
||||
elements.testConnection.addEventListener('click', testApiConnections);
|
||||
|
||||
// Detect GPU
|
||||
elements.detectGpu.addEventListener('click', detectGpu);
|
||||
|
||||
// Calibration controls
|
||||
elements.startCalibration.addEventListener('click', startCalibration);
|
||||
elements.stopCalibration.addEventListener('click', stopCalibration);
|
||||
|
||||
async function loadSettings() {
|
||||
try {
|
||||
const response = await fetch('http://localhost:3000/api/settings');
|
||||
const settings = await response.json();
|
||||
|
||||
// Populate form fields
|
||||
elements.openaiKey.value = settings.openaiApiKey || '';
|
||||
elements.elevenlabsKey.value = settings.elevenlabsApiKey || '';
|
||||
elements.whisperModel.value = settings.whisperModel || 'base';
|
||||
elements.gpuAcceleration.value = settings.gpuAcceleration || 'none';
|
||||
elements.audioThreshold.value = settings.audioThreshold || 0.01;
|
||||
elements.thresholdValue.textContent = (settings.audioThreshold || 0.01).toFixed(3);
|
||||
elements.chunkSize.value = settings.chunkSize || 2000;
|
||||
elements.defaultBackend.value = settings.selectedBackend || 'whisper-local';
|
||||
elements.textDuration.value = settings.textDuration || 5;
|
||||
elements.textColor.value = settings.textColor || '#ffffff';
|
||||
elements.bgColor.value = settings.backgroundColor || '#000000';
|
||||
elements.fontFamily.value = settings.fontFamily || 'Arial, sans-serif';
|
||||
elements.fontSize.value = settings.fontSize || 32;
|
||||
elements.fontSizeValue.textContent = (settings.fontSize || 32) + 'px';
|
||||
elements.serverPort.value = settings.serverPort || 3000;
|
||||
|
||||
// Update ambient noise level display
|
||||
if (settings.ambientNoiseLevel !== null) {
|
||||
elements.ambientLevel.textContent = `Current: ${settings.ambientNoiseLevel.toFixed(4)}`;
|
||||
}
|
||||
|
||||
showStatus('Settings loaded successfully', 'success');
|
||||
} catch (error) {
|
||||
console.error('Error loading settings:', error);
|
||||
showStatus('Error loading settings', 'error');
|
||||
}
|
||||
}
|
||||
|
||||
async function saveSettings() {
|
||||
try {
|
||||
const settings = {
|
||||
openaiApiKey: elements.openaiKey.value.trim(),
|
||||
elevenlabsApiKey: elements.elevenlabsKey.value.trim(),
|
||||
whisperModel: elements.whisperModel.value,
|
||||
gpuAcceleration: elements.gpuAcceleration.value,
|
||||
audioThreshold: parseFloat(elements.audioThreshold.value),
|
||||
chunkSize: parseInt(elements.chunkSize.value),
|
||||
selectedBackend: elements.defaultBackend.value,
|
||||
textDuration: parseInt(elements.textDuration.value),
|
||||
textColor: elements.textColor.value,
|
||||
backgroundColor: elements.bgColor.value,
|
||||
fontFamily: elements.fontFamily.value,
|
||||
fontSize: parseInt(elements.fontSize.value),
|
||||
googleFont: elements.fontFamily.options[elements.fontFamily.selectedIndex].getAttribute('data-google'),
|
||||
serverPort: parseInt(elements.serverPort.value)
|
||||
};
|
||||
|
||||
const response = await fetch('http://localhost:3000/api/settings', {
|
||||
method: 'POST',
|
||||
headers: { 'Content-Type': 'application/json' },
|
||||
body: JSON.stringify(settings)
|
||||
});
|
||||
|
||||
if (response.ok) {
|
||||
showStatus('Settings saved successfully! Some changes may require a restart.', 'success');
|
||||
} else {
|
||||
throw new Error('Failed to save settings');
|
||||
}
|
||||
} catch (error) {
|
||||
console.error('Error saving settings:', error);
|
||||
showStatus('Error saving settings', 'error');
|
||||
}
|
||||
}
|
||||
|
||||
async function resetSettings() {
|
||||
if (!confirm('Are you sure you want to reset all settings to defaults? This cannot be undone.')) {
|
||||
return;
|
||||
}
|
||||
|
||||
try {
|
||||
const response = await fetch('http://localhost:3000/api/settings/reset', {
|
||||
method: 'POST'
|
||||
});
|
||||
|
||||
if (response.ok) {
|
||||
await loadSettings();
|
||||
showStatus('Settings reset to defaults', 'success');
|
||||
} else {
|
||||
throw new Error('Failed to reset settings');
|
||||
}
|
||||
} catch (error) {
|
||||
console.error('Error resetting settings:', error);
|
||||
showStatus('Error resetting settings', 'error');
|
||||
}
|
||||
}
|
||||
|
||||
async function detectGpu() {
|
||||
showStatus('Detecting GPU...', 'info');
|
||||
|
||||
try {
|
||||
const response = await fetch('http://localhost:3000/api/detect-gpu', {
|
||||
method: 'POST'
|
||||
});
|
||||
|
||||
const result = await response.json();
|
||||
|
||||
let message = 'GPU Detection Results:\n';
|
||||
message += `NVIDIA CUDA: ${result.cuda ? 'Available' : 'Not found'}\n`;
|
||||
message += `AMD ROCm: ${result.rocm ? 'Available' : 'Not found'}\n`;
|
||||
message += `Intel OpenVINO: ${result.openvino ? 'Available' : 'Not found'}\n`;
|
||||
|
||||
if (result.recommended) {
|
||||
message += `\nRecommended: ${result.recommended}`;
|
||||
elements.gpuAcceleration.value = result.recommended;
|
||||
}
|
||||
|
||||
showStatus(message.replace(/\n/g, '<br>'), 'success');
|
||||
} catch (error) {
|
||||
console.error('Error detecting GPU:', error);
|
||||
showStatus('Error detecting GPU', 'error');
|
||||
}
|
||||
}
|
||||
|
||||
async function testApiConnections() {
|
||||
showStatus('Testing API connections...', 'info');
|
||||
|
||||
try {
|
||||
const response = await fetch('http://localhost:3000/api/test-connections', {
|
||||
method: 'POST',
|
||||
headers: { 'Content-Type': 'application/json' },
|
||||
body: JSON.stringify({
|
||||
openaiApiKey: elements.openaiKey.value.trim(),
|
||||
elevenlabsApiKey: elements.elevenlabsKey.value.trim()
|
||||
})
|
||||
});
|
||||
|
||||
const results = await response.json();
|
||||
|
||||
let message = 'API Connection Test Results:\n';
|
||||
if (results.openai) {
|
||||
message += `✓ OpenAI: ${results.openai.success ? 'Connected' : 'Failed - ' + results.openai.error}\n`;
|
||||
}
|
||||
if (results.elevenlabs) {
|
||||
message += `✓ ElevenLabs: ${results.elevenlabs.success ? 'Connected' : 'Failed - ' + results.elevenlabs.error}\n`;
|
||||
}
|
||||
|
||||
showStatus(message.replace(/\n/g, '<br>'), results.openai?.success || results.elevenlabs?.success ? 'success' : 'error');
|
||||
} catch (error) {
|
||||
console.error('Error testing connections:', error);
|
||||
showStatus('Error testing API connections', 'error');
|
||||
}
|
||||
}
|
||||
|
||||
async function startCalibration() {
|
||||
try {
|
||||
const stream = await navigator.mediaDevices.getUserMedia({ audio: true });
|
||||
|
||||
audioContext = new (window.AudioContext || window.webkitAudioContext)();
|
||||
analyser = audioContext.createAnalyser();
|
||||
microphone = audioContext.createMediaStreamSource(stream);
|
||||
|
||||
analyser.fftSize = 512;
|
||||
analyser.smoothingTimeConstant = 0.3;
|
||||
|
||||
const bufferLength = analyser.frequencyBinCount;
|
||||
const dataArray = new Uint8Array(bufferLength);
|
||||
|
||||
microphone.connect(analyser);
|
||||
|
||||
isCalibrating = true;
|
||||
calibrationData = [];
|
||||
elements.startCalibration.disabled = true;
|
||||
elements.stopCalibration.disabled = false;
|
||||
elements.calibrationStatus.textContent = 'Calibrating... Stay quiet for 10 seconds.';
|
||||
|
||||
let progress = 0;
|
||||
const calibrationDuration = 10000; // 10 seconds
|
||||
const updateInterval = 100; // Update every 100ms
|
||||
const totalUpdates = calibrationDuration / updateInterval;
|
||||
|
||||
calibrationInterval = setInterval(() => {
|
||||
if (!isCalibrating) return;
|
||||
|
||||
analyser.getByteTimeDomainData(dataArray);
|
||||
|
||||
// Calculate RMS
|
||||
let sum = 0;
|
||||
for (let i = 0; i < bufferLength; i++) {
|
||||
const sample = (dataArray[i] - 128) / 128;
|
||||
sum += sample * sample;
|
||||
}
|
||||
const rms = Math.sqrt(sum / bufferLength);
|
||||
calibrationData.push(rms);
|
||||
|
||||
progress++;
|
||||
const progressPercent = (progress / totalUpdates) * 100;
|
||||
elements.calibrationProgress.style.width = `${progressPercent}%`;
|
||||
|
||||
if (progress >= totalUpdates) {
|
||||
finishCalibration();
|
||||
}
|
||||
}, updateInterval);
|
||||
|
||||
} catch (error) {
|
||||
console.error('Error starting calibration:', error);
|
||||
showStatus('Error accessing microphone for calibration', 'error');
|
||||
stopCalibration();
|
||||
}
|
||||
}
|
||||
|
||||
function stopCalibration() {
|
||||
isCalibrating = false;
|
||||
|
||||
if (calibrationInterval) {
|
||||
clearInterval(calibrationInterval);
|
||||
calibrationInterval = null;
|
||||
}
|
||||
|
||||
if (audioContext) {
|
||||
audioContext.close();
|
||||
audioContext = null;
|
||||
}
|
||||
|
||||
elements.startCalibration.disabled = false;
|
||||
elements.stopCalibration.disabled = true;
|
||||
elements.calibrationProgress.style.width = '0%';
|
||||
elements.calibrationStatus.textContent = 'Calibration stopped.';
|
||||
}
|
||||
|
||||
async function finishCalibration() {
|
||||
if (calibrationData.length === 0) {
|
||||
stopCalibration();
|
||||
return;
|
||||
}
|
||||
|
||||
// Calculate average ambient noise level
|
||||
const avgLevel = calibrationData.reduce((a, b) => a + b, 0) / calibrationData.length;
|
||||
|
||||
// Set threshold slightly above ambient level
|
||||
const suggestedThreshold = Math.max(avgLevel * 3, 0.005);
|
||||
|
||||
try {
|
||||
const response = await fetch('http://localhost:3000/api/settings', {
|
||||
method: 'POST',
|
||||
headers: { 'Content-Type': 'application/json' },
|
||||
body: JSON.stringify({
|
||||
ambientNoiseLevel: avgLevel,
|
||||
audioThreshold: suggestedThreshold
|
||||
})
|
||||
});
|
||||
|
||||
if (response.ok) {
|
||||
elements.ambientLevel.textContent = `Current: ${avgLevel.toFixed(4)}`;
|
||||
elements.audioThreshold.value = suggestedThreshold;
|
||||
elements.thresholdValue.textContent = suggestedThreshold.toFixed(3);
|
||||
elements.calibrationStatus.textContent = `Calibration complete! New threshold: ${suggestedThreshold.toFixed(3)}`;
|
||||
showStatus('Ambient noise calibration completed successfully', 'success');
|
||||
}
|
||||
} catch (error) {
|
||||
console.error('Error saving calibration:', error);
|
||||
showStatus('Error saving calibration results', 'error');
|
||||
}
|
||||
|
||||
stopCalibration();
|
||||
}
|
||||
|
||||
function showStatus(message, type) {
|
||||
elements.statusMessage.textContent = message;
|
||||
elements.statusMessage.innerHTML = message; // Allow HTML for test results
|
||||
elements.statusMessage.className = `status-${type}`;
|
||||
elements.statusMessage.style.display = 'block';
|
||||
|
||||
// Auto-hide after 5 seconds
|
||||
setTimeout(() => {
|
||||
elements.statusMessage.style.display = 'none';
|
||||
}, 5000);
|
||||
}
|
||||
});
|
||||
|
||||
// Add status message styles
|
||||
const style = document.createElement('style');
|
||||
style.textContent = `
|
||||
.status-success { background: #d4edda; color: #155724; border: 1px solid #c3e6cb; }
|
||||
.status-error { background: #f8d7da; color: #721c24; border: 1px solid #f5c6cb; }
|
||||
.status-info { background: #d1ecf1; color: #0c5460; border: 1px solid #bee5eb; }
|
||||
`;
|
||||
document.head.appendChild(style);
|
||||
@@ -0,0 +1,319 @@
|
||||
<!DOCTYPE html>
|
||||
<html lang="en">
|
||||
<head>
|
||||
<meta charset="UTF-8">
|
||||
<meta name="viewport" content="width=device-width, initial-scale=1.0">
|
||||
<meta http-equiv="Content-Security-Policy" content="default-src 'self' http://localhost:3000 https://fonts.googleapis.com https://fonts.gstatic.com; script-src 'self' 'unsafe-inline' http://localhost:3000; style-src 'self' 'unsafe-inline' https://fonts.googleapis.com; font-src 'self' https://fonts.gstatic.com; connect-src 'self' http://localhost:3000 ws://localhost:3000">
|
||||
<title>Settings - Live Transcription</title>
|
||||
<link rel="stylesheet" href="css/styles.css">
|
||||
<style>
|
||||
body {
|
||||
margin: 0;
|
||||
padding: 0;
|
||||
height: 100vh;
|
||||
overflow-y: auto;
|
||||
background: #0a0a0a;
|
||||
}
|
||||
|
||||
.settings-container {
|
||||
max-width: 800px;
|
||||
margin: 0 auto;
|
||||
padding: 20px;
|
||||
min-height: 100vh;
|
||||
}
|
||||
|
||||
.settings-section {
|
||||
background: #1a1a1a;
|
||||
margin-bottom: 20px;
|
||||
padding: 20px;
|
||||
border-radius: 8px;
|
||||
border: 1px solid #333;
|
||||
}
|
||||
|
||||
.settings-section h3 {
|
||||
color: #fff;
|
||||
margin-bottom: 15px;
|
||||
border-bottom: 1px solid #333;
|
||||
padding-bottom: 10px;
|
||||
}
|
||||
|
||||
.settings-grid {
|
||||
display: grid;
|
||||
grid-template-columns: repeat(auto-fit, minmax(300px, 1fr));
|
||||
gap: 15px;
|
||||
}
|
||||
|
||||
.setting-item {
|
||||
display: flex;
|
||||
flex-direction: column;
|
||||
gap: 5px;
|
||||
}
|
||||
|
||||
.setting-item label {
|
||||
color: #ccc;
|
||||
font-size: 14px;
|
||||
font-weight: 500;
|
||||
}
|
||||
|
||||
.setting-item input,
|
||||
.setting-item select {
|
||||
padding: 8px;
|
||||
border: 1px solid #333;
|
||||
background: #2a2a2a;
|
||||
color: #fff;
|
||||
border-radius: 4px;
|
||||
}
|
||||
|
||||
.setting-item input[type="password"] {
|
||||
font-family: monospace;
|
||||
}
|
||||
|
||||
.calibration-section {
|
||||
background: #2a2a1a;
|
||||
padding: 15px;
|
||||
border-radius: 4px;
|
||||
border: 1px solid #444;
|
||||
}
|
||||
|
||||
.calibration-controls {
|
||||
display: flex;
|
||||
gap: 10px;
|
||||
align-items: center;
|
||||
margin-top: 10px;
|
||||
}
|
||||
|
||||
.calibration-status {
|
||||
color: #ccc;
|
||||
font-size: 12px;
|
||||
margin-top: 5px;
|
||||
}
|
||||
|
||||
.progress-bar {
|
||||
width: 100%;
|
||||
height: 20px;
|
||||
background: #333;
|
||||
border-radius: 10px;
|
||||
overflow: hidden;
|
||||
margin: 10px 0;
|
||||
}
|
||||
|
||||
.progress-fill {
|
||||
height: 100%;
|
||||
background: linear-gradient(90deg, #4CAF50 0%, #FFC107 50%, #f44336 100%);
|
||||
width: 0%;
|
||||
transition: width 0.1s ease;
|
||||
}
|
||||
|
||||
.button-group {
|
||||
display: flex;
|
||||
gap: 10px;
|
||||
margin-top: 20px;
|
||||
}
|
||||
|
||||
.btn-danger {
|
||||
background: #dc3545;
|
||||
color: white;
|
||||
}
|
||||
|
||||
.btn-danger:hover {
|
||||
background: #c82333;
|
||||
}
|
||||
|
||||
.btn-success {
|
||||
background: #28a745;
|
||||
color: white;
|
||||
}
|
||||
|
||||
.btn-success:hover {
|
||||
background: #218838;
|
||||
}
|
||||
|
||||
.back-link {
|
||||
color: #4CAF50;
|
||||
text-decoration: none;
|
||||
margin-bottom: 20px;
|
||||
display: inline-block;
|
||||
}
|
||||
|
||||
.back-link:hover {
|
||||
text-decoration: underline;
|
||||
}
|
||||
|
||||
.help-text {
|
||||
color: #888;
|
||||
font-size: 12px;
|
||||
margin-top: 5px;
|
||||
}
|
||||
</style>
|
||||
</head>
|
||||
<body>
|
||||
<div class="settings-container">
|
||||
<button onclick="window.location.href='index.html'" class="btn btn-small" style="margin-bottom: 20px;">← Back to Main</button>
|
||||
|
||||
<h1 style="color: #fff; margin-bottom: 30px;">Settings</h1>
|
||||
|
||||
<div class="settings-section">
|
||||
<h3>API Configuration</h3>
|
||||
<div class="settings-grid">
|
||||
<div class="setting-item">
|
||||
<label for="openai-key">OpenAI API Key</label>
|
||||
<input type="password" id="openai-key" placeholder="Enter your OpenAI API key">
|
||||
<div class="help-text">Required for remote Whisper transcription</div>
|
||||
</div>
|
||||
|
||||
<div class="setting-item">
|
||||
<label for="elevenlabs-key">ElevenLabs API Key</label>
|
||||
<input type="password" id="elevenlabs-key" placeholder="Enter your ElevenLabs API key">
|
||||
<div class="help-text">Required for ElevenLabs Scribe transcription</div>
|
||||
</div>
|
||||
</div>
|
||||
</div>
|
||||
|
||||
<div class="settings-section">
|
||||
<h3>Audio Processing</h3>
|
||||
<div class="settings-grid">
|
||||
<div class="setting-item">
|
||||
<label for="whisper-model">Whisper Model</label>
|
||||
<select id="whisper-model">
|
||||
<option value="tiny">Tiny (fastest, least accurate)</option>
|
||||
<option value="base">Base (good balance)</option>
|
||||
<option value="small">Small (more accurate)</option>
|
||||
<option value="medium">Medium (high accuracy)</option>
|
||||
<option value="large">Large (highest accuracy, slowest)</option>
|
||||
</select>
|
||||
<div class="help-text">Larger models are more accurate but slower</div>
|
||||
</div>
|
||||
|
||||
<div class="setting-item">
|
||||
<label for="gpu-acceleration">GPU Acceleration</label>
|
||||
<select id="gpu-acceleration">
|
||||
<option value="none">Disabled (CPU only)</option>
|
||||
<option value="cuda">NVIDIA CUDA</option>
|
||||
<option value="rocm">AMD ROCm</option>
|
||||
<option value="openvino">Intel OpenVINO</option>
|
||||
</select>
|
||||
<div class="help-text">
|
||||
CUDA: NVIDIA GPUs<br>
|
||||
ROCm: AMD GPUs (Linux only)<br>
|
||||
OpenVINO: Intel Arc/Integrated GPUs
|
||||
</div>
|
||||
</div>
|
||||
|
||||
<div class="setting-item">
|
||||
<label for="audio-threshold">Audio Threshold</label>
|
||||
<input type="range" id="audio-threshold" min="0.001" max="0.1" step="0.001" value="0.01">
|
||||
<span id="threshold-value">0.01</span>
|
||||
<div class="help-text">Lower = more sensitive to quiet sounds</div>
|
||||
</div>
|
||||
|
||||
<div class="setting-item">
|
||||
<label for="chunk-size">Transcription Speed</label>
|
||||
<select id="chunk-size">
|
||||
<option value="1000">Very Fast (1s chunks)</option>
|
||||
<option value="2000">Fast (2s chunks)</option>
|
||||
<option value="3000">Balanced (3s chunks)</option>
|
||||
<option value="5000">Quality (5s chunks)</option>
|
||||
</select>
|
||||
<div class="help-text">Faster = more responsive but may cut off words</div>
|
||||
</div>
|
||||
</div>
|
||||
|
||||
<div class="calibration-section">
|
||||
<h4 style="color: #fff; margin-bottom: 10px;">Ambient Noise Calibration</h4>
|
||||
<p style="color: #ccc; font-size: 14px;">Calibrate the app to learn your ambient noise level for better threshold detection.</p>
|
||||
|
||||
<div class="progress-bar">
|
||||
<div id="calibration-progress" class="progress-fill"></div>
|
||||
</div>
|
||||
|
||||
<div class="calibration-controls">
|
||||
<button id="start-calibration" class="btn btn-primary">Start Calibration</button>
|
||||
<button id="stop-calibration" class="btn btn-secondary" disabled>Stop</button>
|
||||
<span id="ambient-level">Current: Not calibrated</span>
|
||||
</div>
|
||||
|
||||
<div id="calibration-status" class="calibration-status">
|
||||
Click "Start Calibration" and stay quiet for 10 seconds to measure ambient noise.
|
||||
</div>
|
||||
</div>
|
||||
</div>
|
||||
|
||||
<div class="settings-section">
|
||||
<h3>Display Settings</h3>
|
||||
<div class="settings-grid">
|
||||
<div class="setting-item">
|
||||
<label for="default-backend">Default Backend</label>
|
||||
<select id="default-backend">
|
||||
<option value="whisper-local">Whisper (Local)</option>
|
||||
<option value="whisper-remote">Whisper (Remote)</option>
|
||||
<option value="elevenlabs">ElevenLabs (Scribe)</option>
|
||||
</select>
|
||||
</div>
|
||||
|
||||
<div class="setting-item">
|
||||
<label for="text-duration">Text Duration (seconds)</label>
|
||||
<input type="number" id="text-duration" min="1" max="60" value="5">
|
||||
</div>
|
||||
|
||||
<div class="setting-item">
|
||||
<label for="text-color">Text Color</label>
|
||||
<input type="color" id="text-color" value="#ffffff">
|
||||
</div>
|
||||
|
||||
<div class="setting-item">
|
||||
<label for="bg-color">Background Color</label>
|
||||
<input type="color" id="bg-color" value="#000000">
|
||||
</div>
|
||||
|
||||
<div class="setting-item">
|
||||
<label for="font-family">Font Family</label>
|
||||
<select id="font-family">
|
||||
<optgroup label="System Fonts">
|
||||
<option value="Arial, sans-serif">Arial</option>
|
||||
<option value="'Times New Roman', serif">Times New Roman</option>
|
||||
<option value="'Courier New', monospace">Courier New</option>
|
||||
<option value="Georgia, serif">Georgia</option>
|
||||
<option value="Verdana, sans-serif">Verdana</option>
|
||||
</optgroup>
|
||||
<optgroup label="Google Fonts">
|
||||
<option value="'Roboto', sans-serif" data-google="Roboto">Roboto</option>
|
||||
<option value="'Open Sans', sans-serif" data-google="Open+Sans">Open Sans</option>
|
||||
<option value="'Lato', sans-serif" data-google="Lato">Lato</option>
|
||||
<option value="'Montserrat', sans-serif" data-google="Montserrat">Montserrat</option>
|
||||
<option value="'Poppins', sans-serif" data-google="Poppins">Poppins</option>
|
||||
</optgroup>
|
||||
</select>
|
||||
</div>
|
||||
|
||||
<div class="setting-item">
|
||||
<label for="font-size">Font Size</label>
|
||||
<input type="range" id="font-size" min="16" max="72" value="32">
|
||||
<span id="font-size-value">32px</span>
|
||||
</div>
|
||||
</div>
|
||||
</div>
|
||||
|
||||
<div class="settings-section">
|
||||
<h3>Server Settings</h3>
|
||||
<div class="settings-grid">
|
||||
<div class="setting-item">
|
||||
<label for="server-port">Server Port</label>
|
||||
<input type="number" id="server-port" min="1000" max="65535" value="3000">
|
||||
<div class="help-text">Restart required after changing</div>
|
||||
</div>
|
||||
</div>
|
||||
</div>
|
||||
|
||||
<div class="button-group">
|
||||
<button id="save-settings" class="btn btn-success">Save Settings</button>
|
||||
<button id="reset-settings" class="btn btn-danger">Reset to Defaults</button>
|
||||
<button id="test-connection" class="btn btn-primary">Test API Connections</button>
|
||||
<button id="detect-gpu" class="btn btn-primary">Detect GPU</button>
|
||||
</div>
|
||||
|
||||
<div id="status-message" style="margin-top: 20px; padding: 10px; border-radius: 4px; display: none;"></div>
|
||||
</div>
|
||||
|
||||
<script src="js/settings.js"></script>
|
||||
</body>
|
||||
</html>
|
||||
+130
@@ -0,0 +1,130 @@
|
||||
const { app, BrowserWindow, ipcMain } = require('electron');
|
||||
const path = require('path');
|
||||
const { spawn } = require('child_process');
|
||||
|
||||
let mainWindow;
|
||||
let displayWindow;
|
||||
let serverProcess;
|
||||
|
||||
function createWindow() {
|
||||
mainWindow = new BrowserWindow({
|
||||
width: 800,
|
||||
height: 600,
|
||||
webPreferences: {
|
||||
nodeIntegration: false,
|
||||
contextIsolation: true,
|
||||
preload: path.join(__dirname, 'preload.js')
|
||||
}
|
||||
});
|
||||
|
||||
mainWindow.loadFile(path.join(__dirname, '../public/index.html'));
|
||||
|
||||
if (process.argv.includes('--dev')) {
|
||||
mainWindow.webContents.openDevTools();
|
||||
}
|
||||
|
||||
mainWindow.on('closed', () => {
|
||||
mainWindow = null;
|
||||
if (displayWindow) {
|
||||
displayWindow.close();
|
||||
}
|
||||
});
|
||||
}
|
||||
|
||||
function createDisplayWindow() {
|
||||
displayWindow = new BrowserWindow({
|
||||
width: 600,
|
||||
height: 200,
|
||||
transparent: true,
|
||||
frame: false,
|
||||
alwaysOnTop: true,
|
||||
skipTaskbar: true,
|
||||
webPreferences: {
|
||||
nodeIntegration: false,
|
||||
contextIsolation: true,
|
||||
preload: path.join(__dirname, 'preload.js')
|
||||
}
|
||||
});
|
||||
|
||||
displayWindow.loadFile(path.join(__dirname, '../public/display.html'));
|
||||
|
||||
displayWindow.on('closed', () => {
|
||||
displayWindow = null;
|
||||
if (mainWindow) {
|
||||
mainWindow.webContents.send('display-window-closed');
|
||||
}
|
||||
});
|
||||
|
||||
return displayWindow;
|
||||
}
|
||||
|
||||
function startServer() {
|
||||
return new Promise((resolve) => {
|
||||
serverProcess = spawn('node', [path.join(__dirname, 'server/index.js')], {
|
||||
cwd: path.join(__dirname, '..'),
|
||||
env: { ...process.env }
|
||||
});
|
||||
|
||||
serverProcess.stdout.on('data', (data) => {
|
||||
console.log(`Server: ${data}`);
|
||||
if (data.toString().includes('Server running on port')) {
|
||||
resolve();
|
||||
}
|
||||
});
|
||||
|
||||
serverProcess.stderr.on('data', (data) => {
|
||||
console.error(`Server Error: ${data}`);
|
||||
});
|
||||
|
||||
// Fallback resolve after 2 seconds
|
||||
setTimeout(resolve, 2000);
|
||||
});
|
||||
}
|
||||
|
||||
app.whenReady().then(async () => {
|
||||
await startServer();
|
||||
createWindow();
|
||||
});
|
||||
|
||||
app.on('window-all-closed', () => {
|
||||
if (serverProcess) {
|
||||
serverProcess.kill();
|
||||
}
|
||||
if (process.platform !== 'darwin') {
|
||||
app.quit();
|
||||
}
|
||||
});
|
||||
|
||||
app.on('activate', () => {
|
||||
if (BrowserWindow.getAllWindows().length === 0) {
|
||||
createWindow();
|
||||
}
|
||||
});
|
||||
|
||||
ipcMain.handle('toggle-display-window', () => {
|
||||
if (displayWindow) {
|
||||
displayWindow.close();
|
||||
return false;
|
||||
} else {
|
||||
createDisplayWindow();
|
||||
return true;
|
||||
}
|
||||
});
|
||||
|
||||
ipcMain.handle('update-display-settings', (event, settings) => {
|
||||
if (displayWindow) {
|
||||
displayWindow.webContents.send('update-settings', settings);
|
||||
}
|
||||
if (mainWindow) {
|
||||
mainWindow.webContents.send('update-settings', settings);
|
||||
}
|
||||
});
|
||||
|
||||
ipcMain.handle('send-transcription', (event, data) => {
|
||||
if (displayWindow) {
|
||||
displayWindow.webContents.send('transcription', data);
|
||||
}
|
||||
if (mainWindow) {
|
||||
mainWindow.webContents.send('transcription', data);
|
||||
}
|
||||
});
|
||||
@@ -0,0 +1,10 @@
|
||||
const { contextBridge, ipcRenderer } = require('electron');
|
||||
|
||||
contextBridge.exposeInMainWorld('electronAPI', {
|
||||
toggleDisplayWindow: () => ipcRenderer.invoke('toggle-display-window'),
|
||||
updateDisplaySettings: (settings) => ipcRenderer.invoke('update-display-settings', settings),
|
||||
sendTranscription: (data) => ipcRenderer.invoke('send-transcription', data),
|
||||
onTranscription: (callback) => ipcRenderer.on('transcription', (event, data) => callback(data)),
|
||||
onUpdateSettings: (callback) => ipcRenderer.on('update-settings', (event, settings) => callback(settings)),
|
||||
onDisplayWindowClosed: (callback) => ipcRenderer.on('display-window-closed', () => callback())
|
||||
});
|
||||
@@ -0,0 +1,57 @@
|
||||
const axios = require('axios');
|
||||
const FormData = require('form-data');
|
||||
const settingsManager = require('../utils/settingsManager');
|
||||
|
||||
async function transcribe(audioBuffer) {
|
||||
// Check for API key in settings first, then env
|
||||
const apiKey = settingsManager.getSetting('elevenlabsApiKey') || process.env.ELEVENLABS_API_KEY;
|
||||
|
||||
if (!apiKey) {
|
||||
throw new Error('ElevenLabs API key not configured');
|
||||
}
|
||||
|
||||
const formData = new FormData();
|
||||
formData.append('model_id', 'scribe_v1');
|
||||
formData.append('file', audioBuffer, {
|
||||
filename: 'audio.webm',
|
||||
contentType: 'audio/webm'
|
||||
});
|
||||
|
||||
try {
|
||||
const response = await axios.post(
|
||||
'https://api.elevenlabs.io/v1/speech-to-text',
|
||||
formData,
|
||||
{
|
||||
headers: {
|
||||
...formData.getHeaders(),
|
||||
'xi-api-key': apiKey
|
||||
},
|
||||
maxContentLength: Infinity,
|
||||
maxBodyLength: Infinity
|
||||
}
|
||||
);
|
||||
|
||||
// Handle the response
|
||||
if (response.data && typeof response.data.text === 'string') {
|
||||
// Return the text even if it's empty (no speech detected)
|
||||
return response.data.text;
|
||||
} else if (response.data && response.data.transcription) {
|
||||
return response.data.transcription;
|
||||
} else if (typeof response.data === 'string') {
|
||||
return response.data;
|
||||
} else {
|
||||
console.log('ElevenLabs unexpected response:', JSON.stringify(response.data, null, 2));
|
||||
throw new Error('Unexpected response format from ElevenLabs API');
|
||||
}
|
||||
} catch (error) {
|
||||
if (error.response) {
|
||||
console.error('ElevenLabs API error response:', error.response.data);
|
||||
throw new Error(`ElevenLabs API error: ${error.response.data.detail?.message || error.response.data.error || error.response.statusText}`);
|
||||
}
|
||||
throw error;
|
||||
}
|
||||
}
|
||||
|
||||
module.exports = {
|
||||
transcribe
|
||||
};
|
||||
@@ -0,0 +1,192 @@
|
||||
const { spawn } = require('child_process');
|
||||
const fs = require('fs').promises;
|
||||
const path = require('path');
|
||||
const os = require('os');
|
||||
const OpenAI = require('openai');
|
||||
const transcriptionQueue = require('../utils/transcriptionQueue');
|
||||
const { getAudioLevel } = require('../utils/audioUtils');
|
||||
const settingsManager = require('../utils/settingsManager');
|
||||
|
||||
const openai = process.env.OPENAI_API_KEY ? new OpenAI({
|
||||
apiKey: process.env.OPENAI_API_KEY
|
||||
}) : null;
|
||||
|
||||
// Get audio threshold dynamically from settings
|
||||
function getAudioThreshold() {
|
||||
return settingsManager.getSetting('audioThreshold') || parseFloat(process.env.AUDIO_THRESHOLD) || 0.01;
|
||||
}
|
||||
|
||||
async function transcribeLocal(audioBuffer) {
|
||||
// Check audio level before processing
|
||||
const audioLevel = getAudioLevel(audioBuffer);
|
||||
const audioThreshold = getAudioThreshold();
|
||||
|
||||
if (audioLevel < audioThreshold) {
|
||||
console.log(`Audio level ${audioLevel.toFixed(4)} below threshold ${audioThreshold}, skipping transcription`);
|
||||
return ''; // Return empty string for silence
|
||||
}
|
||||
|
||||
console.log(`Audio level ${audioLevel.toFixed(4)} above threshold, queuing for transcription`);
|
||||
|
||||
// Use queue to prevent multiple concurrent Whisper processes
|
||||
return transcriptionQueue.enqueue(async () => {
|
||||
const tempDir = os.tmpdir();
|
||||
const tempFile = path.join(tempDir, `audio_${Date.now()}.wav`);
|
||||
|
||||
try {
|
||||
await fs.writeFile(tempFile, audioBuffer);
|
||||
|
||||
return new Promise((resolve, reject) => {
|
||||
// Build Whisper command with current settings (not cached)
|
||||
const currentModel = settingsManager.getSetting('whisperModel') || process.env.WHISPER_MODEL || 'base';
|
||||
const whisperArgs = [
|
||||
tempFile,
|
||||
'--model', currentModel,
|
||||
'--language', 'en',
|
||||
'--task', 'transcribe',
|
||||
'--output_format', 'txt',
|
||||
'--output_dir', tempDir
|
||||
];
|
||||
|
||||
// Get current GPU settings (not cached)
|
||||
const gpuAcceleration = settingsManager.getSetting('gpuAcceleration') || 'none';
|
||||
|
||||
switch (gpuAcceleration) {
|
||||
case 'cuda':
|
||||
// NVIDIA CUDA
|
||||
whisperArgs.push('--device', 'cuda');
|
||||
whisperArgs.push('--fp16', 'True');
|
||||
break;
|
||||
|
||||
case 'rocm':
|
||||
// AMD ROCm (requires PyTorch built with ROCm)
|
||||
whisperArgs.push('--device', 'cuda'); // ROCm uses CUDA interface
|
||||
whisperArgs.push('--fp16', 'True');
|
||||
// Set environment variable for ROCm
|
||||
process.env.HSA_OVERRIDE_GFX_VERSION = '10.3.0'; // May need adjustment
|
||||
break;
|
||||
|
||||
case 'openvino':
|
||||
// Intel OpenVINO
|
||||
whisperArgs.push('--device', 'cpu'); // OpenVINO uses CPU device
|
||||
// Note: Requires whisper-openvino fork or custom implementation
|
||||
whisperArgs.push('--fp16', 'False');
|
||||
whisperArgs.push('--threads', '4');
|
||||
break;
|
||||
|
||||
default:
|
||||
// CPU optimizations
|
||||
whisperArgs.push('--device', 'cpu');
|
||||
whisperArgs.push('--fp16', 'False');
|
||||
whisperArgs.push('--threads', '2');
|
||||
}
|
||||
|
||||
console.log(`Starting Whisper with model: ${currentModel}, GPU: ${gpuAcceleration}`);
|
||||
console.log('Whisper args:', whisperArgs.join(' '));
|
||||
const whisperProcess = spawn('whisper', whisperArgs);
|
||||
|
||||
let output = '';
|
||||
let error = '';
|
||||
|
||||
whisperProcess.stdout.on('data', (data) => {
|
||||
output += data.toString();
|
||||
});
|
||||
|
||||
whisperProcess.stderr.on('data', (data) => {
|
||||
error += data.toString();
|
||||
});
|
||||
|
||||
whisperProcess.on('close', async (code) => {
|
||||
if (code !== 0) {
|
||||
reject(new Error(`Whisper process exited with code ${code}: ${error}`));
|
||||
return;
|
||||
}
|
||||
|
||||
try {
|
||||
const outputFile = tempFile.replace('.wav', '.txt');
|
||||
const transcription = await fs.readFile(outputFile, 'utf8');
|
||||
|
||||
// Clean up files
|
||||
await fs.unlink(tempFile).catch(() => {});
|
||||
await fs.unlink(outputFile).catch(() => {});
|
||||
|
||||
resolve(transcription.trim());
|
||||
} catch (err) {
|
||||
reject(err);
|
||||
}
|
||||
});
|
||||
|
||||
// Kill process if it takes too long (30 seconds timeout)
|
||||
const timeout = setTimeout(() => {
|
||||
whisperProcess.kill('SIGKILL');
|
||||
reject(new Error('Whisper process timeout'));
|
||||
}, 30000);
|
||||
|
||||
whisperProcess.on('close', () => {
|
||||
clearTimeout(timeout);
|
||||
});
|
||||
});
|
||||
} catch (error) {
|
||||
await fs.unlink(tempFile).catch(() => {});
|
||||
throw error;
|
||||
}
|
||||
}, `whisper_${Date.now()}`);
|
||||
}
|
||||
|
||||
async function transcribeRemote(audioBuffer) {
|
||||
// Check for API key in settings first, then env
|
||||
const apiKey = settingsManager.getSetting('openaiApiKey') || process.env.OPENAI_API_KEY;
|
||||
|
||||
if (!apiKey) {
|
||||
throw new Error('OpenAI API key not configured');
|
||||
}
|
||||
|
||||
// Create OpenAI client with current API key
|
||||
const currentOpenAI = new OpenAI({ apiKey });
|
||||
|
||||
// Check audio level before processing
|
||||
const audioLevel = getAudioLevel(audioBuffer);
|
||||
const audioThreshold = getAudioThreshold();
|
||||
|
||||
if (audioLevel < audioThreshold) {
|
||||
console.log(`Audio level ${audioLevel.toFixed(4)} below threshold ${audioThreshold}, skipping remote transcription`);
|
||||
return ''; // Return empty string for silence
|
||||
}
|
||||
|
||||
console.log(`Audio level ${audioLevel.toFixed(4)} above threshold, sending to OpenAI`);
|
||||
|
||||
// Use queue for remote requests too to prevent overwhelming the API
|
||||
return transcriptionQueue.enqueue(async () => {
|
||||
const tempDir = os.tmpdir();
|
||||
const tempFile = path.join(tempDir, `audio_${Date.now()}.wav`);
|
||||
|
||||
try {
|
||||
await fs.writeFile(tempFile, audioBuffer);
|
||||
|
||||
const fileBuffer = await fs.readFile(tempFile);
|
||||
const transcription = await currentOpenAI.audio.transcriptions.create({
|
||||
file: new File([fileBuffer], 'audio.wav', { type: 'audio/wav' }),
|
||||
model: "whisper-1",
|
||||
language: "en"
|
||||
});
|
||||
|
||||
await fs.unlink(tempFile).catch(() => {});
|
||||
|
||||
return transcription.text;
|
||||
} catch (error) {
|
||||
await fs.unlink(tempFile).catch(() => {});
|
||||
throw error;
|
||||
}
|
||||
}, `openai_${Date.now()}`);
|
||||
}
|
||||
|
||||
// Clean up old queued jobs periodically
|
||||
setInterval(() => {
|
||||
transcriptionQueue.clearOldJobs(30000); // Clear jobs older than 30 seconds
|
||||
}, 10000);
|
||||
|
||||
module.exports = {
|
||||
transcribeLocal,
|
||||
transcribeRemote,
|
||||
getQueueStatus: () => transcriptionQueue.getStatus()
|
||||
};
|
||||
@@ -0,0 +1,271 @@
|
||||
require('dotenv').config();
|
||||
const express = require('express');
|
||||
const http = require('http');
|
||||
const socketIo = require('socket.io');
|
||||
const cors = require('cors');
|
||||
const path = require('path');
|
||||
const multer = require('multer');
|
||||
|
||||
const whisperBackend = require('./backends/whisper');
|
||||
const elevenLabsBackend = require('./backends/elevenlabs');
|
||||
const settingsManager = require('./utils/settingsManager');
|
||||
|
||||
const app = express();
|
||||
const server = http.createServer(app);
|
||||
const io = socketIo(server, {
|
||||
cors: {
|
||||
origin: "*",
|
||||
methods: ["GET", "POST"]
|
||||
}
|
||||
});
|
||||
|
||||
const upload = multer({ storage: multer.memoryStorage() });
|
||||
|
||||
app.use(cors());
|
||||
app.use(express.json());
|
||||
app.use(express.static(path.join(__dirname, '../../public')));
|
||||
|
||||
const activeBackend = {
|
||||
current: 'whisper-local',
|
||||
instance: null
|
||||
};
|
||||
|
||||
// Initialize settings
|
||||
settingsManager.loadSettings().then(() => {
|
||||
settingsManager.applyToEnv();
|
||||
activeBackend.current = settingsManager.getSetting('selectedBackend') || 'whisper-local';
|
||||
console.log('Settings loaded, active backend:', activeBackend.current);
|
||||
});
|
||||
|
||||
app.post('/api/backend', (req, res) => {
|
||||
const { backend } = req.body;
|
||||
activeBackend.current = backend;
|
||||
console.log(`Switched to backend: ${backend}`);
|
||||
res.json({ success: true, backend });
|
||||
});
|
||||
|
||||
app.post('/api/transcribe', upload.single('audio'), async (req, res) => {
|
||||
try {
|
||||
if (!req.file) {
|
||||
return res.status(400).json({ error: 'No audio file provided' });
|
||||
}
|
||||
|
||||
let transcription = '';
|
||||
|
||||
switch (activeBackend.current) {
|
||||
case 'whisper-local':
|
||||
transcription = await whisperBackend.transcribeLocal(req.file.buffer);
|
||||
break;
|
||||
case 'whisper-remote':
|
||||
transcription = await whisperBackend.transcribeRemote(req.file.buffer);
|
||||
break;
|
||||
case 'elevenlabs':
|
||||
transcription = await elevenLabsBackend.transcribe(req.file.buffer);
|
||||
break;
|
||||
default:
|
||||
throw new Error(`Unknown backend: ${activeBackend.current}`);
|
||||
}
|
||||
|
||||
// Only emit non-empty transcriptions
|
||||
if (transcription && transcription.trim()) {
|
||||
io.emit('transcription', { text: transcription, timestamp: Date.now() });
|
||||
}
|
||||
res.json({ transcription });
|
||||
} catch (error) {
|
||||
console.error('Transcription error:', error);
|
||||
res.status(500).json({ error: error.message });
|
||||
}
|
||||
});
|
||||
|
||||
app.get('/api/backends', (req, res) => {
|
||||
res.json({
|
||||
backends: [
|
||||
{ id: 'whisper-local', name: 'Whisper (Local)', available: true },
|
||||
{ id: 'whisper-remote', name: 'Whisper (OpenAI API)', available: !!process.env.OPENAI_API_KEY },
|
||||
{ id: 'elevenlabs', name: 'ElevenLabs (Scribe)', available: !!process.env.ELEVENLABS_API_KEY }
|
||||
],
|
||||
current: activeBackend.current
|
||||
});
|
||||
});
|
||||
|
||||
app.get('/api/queue-status', (req, res) => {
|
||||
try {
|
||||
const status = whisperBackend.getQueueStatus();
|
||||
res.json(status);
|
||||
} catch (error) {
|
||||
res.status(500).json({ error: 'Failed to get queue status' });
|
||||
}
|
||||
});
|
||||
|
||||
// Settings endpoints
|
||||
app.get('/api/settings', (req, res) => {
|
||||
try {
|
||||
const settings = settingsManager.getSettings();
|
||||
// Don't send API keys to client for security
|
||||
const clientSettings = { ...settings };
|
||||
if (clientSettings.openaiApiKey) {
|
||||
clientSettings.openaiApiKey = '***hidden***';
|
||||
}
|
||||
if (clientSettings.elevenlabsApiKey) {
|
||||
clientSettings.elevenlabsApiKey = '***hidden***';
|
||||
}
|
||||
res.json(clientSettings);
|
||||
} catch (error) {
|
||||
res.status(500).json({ error: 'Failed to load settings' });
|
||||
}
|
||||
});
|
||||
|
||||
app.post('/api/settings', async (req, res) => {
|
||||
try {
|
||||
const updates = req.body;
|
||||
|
||||
// Only update API keys if they're not the hidden placeholder
|
||||
if (updates.openaiApiKey === '***hidden***') {
|
||||
delete updates.openaiApiKey;
|
||||
}
|
||||
if (updates.elevenlabsApiKey === '***hidden***') {
|
||||
delete updates.elevenlabsApiKey;
|
||||
}
|
||||
|
||||
const success = await settingsManager.updateSettings(updates);
|
||||
if (success) {
|
||||
// Reload settings to ensure they're current
|
||||
await settingsManager.loadSettings();
|
||||
settingsManager.applyToEnv();
|
||||
|
||||
// Update active backend if changed
|
||||
if (updates.selectedBackend) {
|
||||
activeBackend.current = updates.selectedBackend;
|
||||
}
|
||||
|
||||
// Log GPU changes for debugging
|
||||
if (updates.gpuAcceleration !== undefined) {
|
||||
console.log(`GPU acceleration changed to: ${updates.gpuAcceleration}`);
|
||||
}
|
||||
if (updates.whisperModel !== undefined) {
|
||||
console.log(`Whisper model changed to: ${updates.whisperModel}`);
|
||||
}
|
||||
|
||||
res.json({ success: true });
|
||||
} else {
|
||||
res.status(500).json({ error: 'Failed to save settings' });
|
||||
}
|
||||
} catch (error) {
|
||||
console.error('Error updating settings:', error);
|
||||
res.status(500).json({ error: 'Failed to save settings' });
|
||||
}
|
||||
});
|
||||
|
||||
app.post('/api/settings/reset', async (req, res) => {
|
||||
try {
|
||||
const success = await settingsManager.resetToDefaults();
|
||||
if (success) {
|
||||
settingsManager.applyToEnv();
|
||||
activeBackend.current = 'whisper-local';
|
||||
res.json({ success: true });
|
||||
} else {
|
||||
res.status(500).json({ error: 'Failed to reset settings' });
|
||||
}
|
||||
} catch (error) {
|
||||
console.error('Error resetting settings:', error);
|
||||
res.status(500).json({ error: 'Failed to reset settings' });
|
||||
}
|
||||
});
|
||||
|
||||
app.post('/api/detect-gpu', async (req, res) => {
|
||||
const { exec } = require('child_process');
|
||||
const util = require('util');
|
||||
const execPromise = util.promisify(exec);
|
||||
|
||||
const result = {
|
||||
cuda: false,
|
||||
rocm: false,
|
||||
openvino: false,
|
||||
recommended: 'none'
|
||||
};
|
||||
|
||||
try {
|
||||
// Check for NVIDIA GPU (CUDA)
|
||||
try {
|
||||
await execPromise('nvidia-smi');
|
||||
result.cuda = true;
|
||||
result.recommended = 'cuda';
|
||||
} catch (e) {
|
||||
// NVIDIA not available
|
||||
}
|
||||
|
||||
// Check for AMD GPU (ROCm)
|
||||
try {
|
||||
await execPromise('rocm-smi');
|
||||
result.rocm = true;
|
||||
if (!result.cuda) {
|
||||
result.recommended = 'rocm';
|
||||
}
|
||||
} catch (e) {
|
||||
// ROCm not available
|
||||
}
|
||||
|
||||
// Check for Intel GPU/OpenVINO
|
||||
try {
|
||||
const { stdout } = await execPromise('lspci | grep -i intel | grep -i vga || echo ""');
|
||||
if (stdout.includes('Intel')) {
|
||||
result.openvino = true;
|
||||
if (!result.cuda && !result.rocm) {
|
||||
result.recommended = 'openvino';
|
||||
}
|
||||
}
|
||||
} catch (e) {
|
||||
// Intel GPU check failed
|
||||
}
|
||||
|
||||
res.json(result);
|
||||
} catch (error) {
|
||||
console.error('GPU detection error:', error);
|
||||
res.status(500).json({ error: 'Failed to detect GPU' });
|
||||
}
|
||||
});
|
||||
|
||||
app.post('/api/test-connections', async (req, res) => {
|
||||
const { openaiApiKey, elevenlabsApiKey } = req.body;
|
||||
const results = {};
|
||||
|
||||
// Test OpenAI connection
|
||||
if (openaiApiKey && openaiApiKey !== '***hidden***') {
|
||||
try {
|
||||
const OpenAI = require('openai');
|
||||
const openai = new OpenAI({ apiKey: openaiApiKey });
|
||||
await openai.models.list();
|
||||
results.openai = { success: true };
|
||||
} catch (error) {
|
||||
results.openai = { success: false, error: error.message };
|
||||
}
|
||||
}
|
||||
|
||||
// Test ElevenLabs connection
|
||||
if (elevenlabsApiKey && elevenlabsApiKey !== '***hidden***') {
|
||||
try {
|
||||
const axios = require('axios');
|
||||
await axios.get('https://api.elevenlabs.io/v1/user', {
|
||||
headers: { 'xi-api-key': elevenlabsApiKey }
|
||||
});
|
||||
results.elevenlabs = { success: true };
|
||||
} catch (error) {
|
||||
results.elevenlabs = { success: false, error: error.response?.data?.detail?.message || error.message };
|
||||
}
|
||||
}
|
||||
|
||||
res.json(results);
|
||||
});
|
||||
|
||||
io.on('connection', (socket) => {
|
||||
console.log('Client connected');
|
||||
|
||||
socket.on('disconnect', () => {
|
||||
console.log('Client disconnected');
|
||||
});
|
||||
});
|
||||
|
||||
const PORT = process.env.SERVER_PORT || 3000;
|
||||
server.listen(PORT, () => {
|
||||
console.log(`Server running on port ${PORT}`);
|
||||
});
|
||||
@@ -0,0 +1,85 @@
|
||||
// Audio utility functions for level detection and processing
|
||||
|
||||
/**
|
||||
* Calculate the RMS (Root Mean Square) amplitude of an audio buffer
|
||||
* @param {Float32Array} audioData - Audio samples
|
||||
* @returns {number} RMS value between 0 and 1
|
||||
*/
|
||||
function calculateRMS(audioData) {
|
||||
let sum = 0;
|
||||
for (let i = 0; i < audioData.length; i++) {
|
||||
sum += audioData[i] * audioData[i];
|
||||
}
|
||||
return Math.sqrt(sum / audioData.length);
|
||||
}
|
||||
|
||||
/**
|
||||
* Check if audio contains speech above threshold
|
||||
* @param {Buffer} audioBuffer - Raw audio buffer
|
||||
* @param {number} threshold - Minimum RMS threshold (0-1)
|
||||
* @returns {Promise<boolean>} True if audio is above threshold
|
||||
*/
|
||||
async function hasAudioAboveThreshold(audioBuffer, threshold = 0.01) {
|
||||
try {
|
||||
// For WebM/audio files, we'd need to decode them properly
|
||||
// For now, we'll do a simple buffer analysis
|
||||
// This is a simplified approach - in production you'd want proper audio decoding
|
||||
|
||||
// Convert buffer to Float32Array (simplified approach)
|
||||
const samples = new Float32Array(audioBuffer.length / 4);
|
||||
for (let i = 0; i < samples.length; i++) {
|
||||
samples[i] = audioBuffer.readFloatLE(i * 4) || 0;
|
||||
}
|
||||
|
||||
const rms = calculateRMS(samples);
|
||||
console.log(`Audio RMS level: ${rms.toFixed(4)}, threshold: ${threshold}`);
|
||||
|
||||
return rms > threshold;
|
||||
} catch (error) {
|
||||
console.warn('Error analyzing audio level, proceeding with transcription:', error);
|
||||
return true; // If we can't analyze, proceed with transcription
|
||||
}
|
||||
}
|
||||
|
||||
/**
|
||||
* Simple audio level detection for basic threshold checking
|
||||
* @param {Buffer} audioBuffer - Raw audio buffer
|
||||
* @returns {number} Estimated audio level (0-1)
|
||||
*/
|
||||
function getAudioLevel(audioBuffer) {
|
||||
try {
|
||||
// For WebM files, we can't easily parse the audio data, so we'll use a simple heuristic
|
||||
// based on the buffer size and non-zero bytes
|
||||
|
||||
let nonZeroBytes = 0;
|
||||
let sumAmplitude = 0;
|
||||
|
||||
// Sample every 100th byte to get a rough estimate
|
||||
for (let i = 0; i < audioBuffer.length; i += 100) {
|
||||
const byte = audioBuffer[i];
|
||||
if (byte !== 0) {
|
||||
nonZeroBytes++;
|
||||
sumAmplitude += Math.abs(byte - 128) / 128; // Normalize assuming 8-bit audio
|
||||
}
|
||||
}
|
||||
|
||||
if (nonZeroBytes === 0) return 0;
|
||||
|
||||
const avgAmplitude = sumAmplitude / nonZeroBytes;
|
||||
const activityRatio = nonZeroBytes / (audioBuffer.length / 100);
|
||||
|
||||
// Combine amplitude and activity for better detection
|
||||
const level = Math.min((avgAmplitude * activityRatio) * 2, 1);
|
||||
|
||||
return level;
|
||||
} catch (error) {
|
||||
console.warn('Error calculating audio level:', error);
|
||||
return 0.1; // Return low level if we can't calculate, but not zero
|
||||
}
|
||||
}
|
||||
|
||||
module.exports = {
|
||||
calculateRMS,
|
||||
hasAudioAboveThreshold,
|
||||
getAudioLevel
|
||||
};
|
||||
@@ -0,0 +1,128 @@
|
||||
const fs = require('fs').promises;
|
||||
const path = require('path');
|
||||
const os = require('os');
|
||||
|
||||
// Settings file location
|
||||
const SETTINGS_DIR = path.join(os.homedir(), '.live-transcription');
|
||||
const SETTINGS_FILE = path.join(SETTINGS_DIR, 'settings.json');
|
||||
|
||||
// Default settings
|
||||
const DEFAULT_SETTINGS = {
|
||||
// API Keys
|
||||
openaiApiKey: '',
|
||||
elevenlabsApiKey: '',
|
||||
|
||||
// Audio Processing
|
||||
audioThreshold: 0.01,
|
||||
chunkSize: 2000,
|
||||
whisperModel: 'base',
|
||||
gpuAcceleration: 'none', // none, cuda, rocm, openvino
|
||||
ambientNoiseLevel: null,
|
||||
|
||||
// UI Settings
|
||||
selectedBackend: 'whisper-local',
|
||||
textDuration: 5,
|
||||
textColor: '#ffffff',
|
||||
backgroundColor: '#000000',
|
||||
fontFamily: 'Arial, sans-serif',
|
||||
fontSize: 32,
|
||||
googleFont: null,
|
||||
|
||||
// Server Settings
|
||||
serverPort: 3000
|
||||
};
|
||||
|
||||
class SettingsManager {
|
||||
constructor() {
|
||||
this.settings = { ...DEFAULT_SETTINGS };
|
||||
this.loaded = false;
|
||||
}
|
||||
|
||||
async ensureSettingsDir() {
|
||||
try {
|
||||
await fs.mkdir(SETTINGS_DIR, { recursive: true });
|
||||
} catch (error) {
|
||||
console.warn('Could not create settings directory:', error);
|
||||
}
|
||||
}
|
||||
|
||||
async loadSettings() {
|
||||
try {
|
||||
await this.ensureSettingsDir();
|
||||
const data = await fs.readFile(SETTINGS_FILE, 'utf8');
|
||||
this.settings = { ...DEFAULT_SETTINGS, ...JSON.parse(data) };
|
||||
this.loaded = true;
|
||||
console.log('Settings loaded from:', SETTINGS_FILE);
|
||||
} catch (error) {
|
||||
console.log('Using default settings (no settings file found)');
|
||||
this.settings = { ...DEFAULT_SETTINGS };
|
||||
this.loaded = true;
|
||||
}
|
||||
return this.settings;
|
||||
}
|
||||
|
||||
async saveSettings(newSettings = null) {
|
||||
try {
|
||||
await this.ensureSettingsDir();
|
||||
if (newSettings) {
|
||||
this.settings = { ...this.settings, ...newSettings };
|
||||
}
|
||||
await fs.writeFile(SETTINGS_FILE, JSON.stringify(this.settings, null, 2));
|
||||
console.log('Settings saved to:', SETTINGS_FILE);
|
||||
return true;
|
||||
} catch (error) {
|
||||
console.error('Error saving settings:', error);
|
||||
return false;
|
||||
}
|
||||
}
|
||||
|
||||
getSetting(key) {
|
||||
return this.settings[key];
|
||||
}
|
||||
|
||||
getSettings() {
|
||||
return { ...this.settings };
|
||||
}
|
||||
|
||||
async updateSetting(key, value) {
|
||||
this.settings[key] = value;
|
||||
return await this.saveSettings();
|
||||
}
|
||||
|
||||
async updateSettings(updates) {
|
||||
this.settings = { ...this.settings, ...updates };
|
||||
return await this.saveSettings();
|
||||
}
|
||||
|
||||
// Get environment variables with settings fallback
|
||||
getEnvWithFallback(envKey, settingsKey) {
|
||||
return process.env[envKey] || this.settings[settingsKey] || DEFAULT_SETTINGS[settingsKey];
|
||||
}
|
||||
|
||||
// Apply settings to process.env (for compatibility with existing code)
|
||||
applyToEnv() {
|
||||
if (this.settings.openaiApiKey) {
|
||||
process.env.OPENAI_API_KEY = this.settings.openaiApiKey;
|
||||
}
|
||||
if (this.settings.elevenlabsApiKey) {
|
||||
process.env.ELEVENLABS_API_KEY = this.settings.elevenlabsApiKey;
|
||||
}
|
||||
if (this.settings.audioThreshold !== null) {
|
||||
process.env.AUDIO_THRESHOLD = this.settings.audioThreshold.toString();
|
||||
}
|
||||
if (this.settings.whisperModel) {
|
||||
process.env.WHISPER_MODEL = this.settings.whisperModel;
|
||||
}
|
||||
if (this.settings.serverPort) {
|
||||
process.env.SERVER_PORT = this.settings.serverPort.toString();
|
||||
}
|
||||
}
|
||||
|
||||
async resetToDefaults() {
|
||||
this.settings = { ...DEFAULT_SETTINGS };
|
||||
return await this.saveSettings();
|
||||
}
|
||||
}
|
||||
|
||||
// Export singleton instance
|
||||
module.exports = new SettingsManager();
|
||||
@@ -0,0 +1,92 @@
|
||||
// Queue system to manage transcription requests and prevent resource overload
|
||||
|
||||
class TranscriptionQueue {
|
||||
constructor(maxConcurrent = 1) {
|
||||
this.queue = [];
|
||||
this.running = [];
|
||||
this.maxConcurrent = maxConcurrent;
|
||||
}
|
||||
|
||||
/**
|
||||
* Add a transcription job to the queue
|
||||
* @param {Function} transcriptionFn - Function that returns a Promise
|
||||
* @param {string} jobId - Unique identifier for the job
|
||||
* @returns {Promise} Promise that resolves with transcription result
|
||||
*/
|
||||
async enqueue(transcriptionFn, jobId = Date.now().toString()) {
|
||||
return new Promise((resolve, reject) => {
|
||||
const job = {
|
||||
id: jobId,
|
||||
fn: transcriptionFn,
|
||||
resolve,
|
||||
reject,
|
||||
timestamp: Date.now()
|
||||
};
|
||||
|
||||
this.queue.push(job);
|
||||
this.processQueue();
|
||||
});
|
||||
}
|
||||
|
||||
/**
|
||||
* Process jobs in the queue
|
||||
*/
|
||||
async processQueue() {
|
||||
if (this.running.length >= this.maxConcurrent || this.queue.length === 0) {
|
||||
return;
|
||||
}
|
||||
|
||||
const job = this.queue.shift();
|
||||
this.running.push(job);
|
||||
|
||||
console.log(`Starting transcription job ${job.id}, queue length: ${this.queue.length}, running: ${this.running.length}`);
|
||||
|
||||
try {
|
||||
const result = await job.fn();
|
||||
job.resolve(result);
|
||||
} catch (error) {
|
||||
job.reject(error);
|
||||
} finally {
|
||||
// Remove from running jobs
|
||||
const index = this.running.findIndex(j => j.id === job.id);
|
||||
if (index > -1) {
|
||||
this.running.splice(index, 1);
|
||||
}
|
||||
|
||||
console.log(`Completed transcription job ${job.id}, queue length: ${this.queue.length}, running: ${this.running.length}`);
|
||||
|
||||
// Process next job
|
||||
this.processQueue();
|
||||
}
|
||||
}
|
||||
|
||||
/**
|
||||
* Clear old jobs from queue (older than maxAge milliseconds)
|
||||
* @param {number} maxAge - Maximum age in milliseconds
|
||||
*/
|
||||
clearOldJobs(maxAge = 30000) {
|
||||
const now = Date.now();
|
||||
const originalLength = this.queue.length;
|
||||
|
||||
this.queue = this.queue.filter(job => (now - job.timestamp) < maxAge);
|
||||
|
||||
if (originalLength !== this.queue.length) {
|
||||
console.log(`Cleared ${originalLength - this.queue.length} old jobs from transcription queue`);
|
||||
}
|
||||
}
|
||||
|
||||
/**
|
||||
* Get queue status
|
||||
* @returns {Object} Queue status information
|
||||
*/
|
||||
getStatus() {
|
||||
return {
|
||||
queueLength: this.queue.length,
|
||||
runningJobs: this.running.length,
|
||||
maxConcurrent: this.maxConcurrent
|
||||
};
|
||||
}
|
||||
}
|
||||
|
||||
// Export singleton instance
|
||||
module.exports = new TranscriptionQueue(1); // Only allow 1 concurrent Whisper process
|
||||
Reference in New Issue
Block a user