Archived
3.8 KiB
3.8 KiB
Installing Whisper for Local Transcription
This guide explains how to install OpenAI's Whisper for local speech-to-text transcription.
Prerequisites
- Python 3.8 or newer
- pip (Python package manager)
- ffmpeg
Installation Steps
1. Install Python Dependencies
# Install Whisper
pip install openai-whisper
# Or install with specific version
pip install openai-whisper==20230918
2. Install ffmpeg
Ubuntu/Debian:
sudo apt update
sudo apt install ffmpeg
macOS (using Homebrew):
brew install ffmpeg
Windows:
- Download ffmpeg from https://ffmpeg.org/download.html
- Extract the archive
- Add the bin folder to your PATH
3. Verify Installation
# Test Whisper installation
whisper --help
# Test with an audio file
whisper audio.mp3 --model base
Available Models
Whisper offers several models with different speed/accuracy tradeoffs:
| Model | Parameters | Relative Speed | Required VRAM |
|---|---|---|---|
| tiny | 39 M | ~32x | ~1 GB |
| base | 74 M | ~16x | ~1 GB |
| small | 244 M | ~6x | ~2 GB |
| medium | 769 M | ~2x | ~5 GB |
| large | 1550 M | 1x | ~10 GB |
Configuration
Set your preferred model in the .env file:
WHISPER_MODEL=base # Options: tiny, base, small, medium, large
GPU Acceleration Setup
NVIDIA GPUs (CUDA)
# Install PyTorch with CUDA support
pip install torch torchvision torchaudio --index-url https://download.pytorch.org/whl/cu118
# Verify CUDA is available
python -c "import torch; print(f'CUDA available: {torch.cuda.is_available()}')"
AMD GPUs (ROCm) - Linux Only
# Install ROCm (Ubuntu/Debian)
wget -q -O - https://repo.radeon.com/rocm/rocm.gpg.key | sudo apt-key add -
echo 'deb [arch=amd64] https://repo.radeon.com/rocm/apt/debian/ ubuntu main' | sudo tee /etc/apt/sources.list.d/rocm.list
sudo apt update
sudo apt install rocm-dkms
# Install PyTorch with ROCm support
pip install torch torchvision torchaudio --index-url https://download.pytorch.org/whl/rocm5.4.2
# Add user to render/video groups
sudo usermod -a -G render,video $USER
# Logout and login for group changes to take effect
Intel GPUs (OpenVINO)
# Install OpenVINO toolkit
pip install openvino-dev[pytorch]
# For Arc GPUs, you may need the compute runtime
# Ubuntu/Debian:
wget -qO - https://repositories.intel.com/graphics/intel-graphics.key | sudo apt-key add -
sudo apt-add-repository 'deb [arch=amd64] https://repositories.intel.com/graphics/ubuntu focal main'
sudo apt update
sudo apt install intel-opencl-icd intel-level-zero-gpu level-zero
# Note: You may need a Whisper fork with OpenVINO support
pip install git+https://github.com/openvinotoolkit/whisper.git
Troubleshooting
Issue: "whisper: command not found"
- Make sure Python's Scripts directory is in your PATH
- On Windows:
C:\Users\[Username]\AppData\Local\Programs\Python\Python3X\Scripts - On macOS/Linux:
~/.local/bin
Issue: GPU not detected
- Use the "Detect GPU" button in the app's settings page
- For NVIDIA: Ensure
nvidia-smicommand works - For AMD: Ensure
rocm-smicommand works - For Intel: Check
lspci | grep -i intelshows your GPU
Issue: CUDA out of memory
- Try using a smaller Whisper model (tiny or base)
- Reduce batch size or use CPU for very long audio
Issue: Audio Format Not Supported
Whisper works best with WAV, MP3, and M4A files. The app automatically converts audio to a compatible format.
Alternative: Docker Installation
You can also run Whisper in a Docker container:
docker pull openai/whisper
docker run -it -v $(pwd):/app openai/whisper audio.mp3 --model base