Compare commits
20
Commits
41b0173680
..
main
| Author | SHA1 | Date | |
|---|---|---|---|
|
|
dc95e5ac55 | ||
|
|
4f7b48c03b | ||
|
|
94651b6ec1 | ||
|
|
aca70dbd0b | ||
|
|
394e4ccf24 | ||
|
|
240330cf3b | ||
|
|
408028c36e | ||
|
|
82fc9ea5f9 | ||
|
|
65c981f889 | ||
|
|
1d390685a6 | ||
|
|
4fbbf89afe | ||
|
|
d8e22a9773 | ||
|
|
8144971707 | ||
|
|
45249174ba | ||
|
|
9d6541ce75 | ||
|
|
37b363b317 | ||
|
|
43f40981db | ||
|
|
35839395f4 | ||
|
|
aa9d6e9765 | ||
|
|
ff48937482 |
@@ -1 +1,2 @@
|
||||
scripts/.env
|
||||
v2/.env
|
||||
+8
-4
@@ -1,5 +1,5 @@
|
||||
# Use Python 3 base image
|
||||
FROM python:3.9-slim
|
||||
FROM python:3.11-slim
|
||||
|
||||
# Set working directory
|
||||
WORKDIR /app
|
||||
@@ -10,11 +10,15 @@ RUN apt-get update && apt-get install -y \
|
||||
libopus0 \
|
||||
&& rm -rf /var/lib/apt/lists/*
|
||||
|
||||
# Install required packages
|
||||
RUN pip install --no-cache-dir discord.py python-dotenv openai
|
||||
# Copy requirements file
|
||||
COPY /scripts/requirements.txt .
|
||||
|
||||
# Copy the bot script
|
||||
# Install required packages from requirements.txt
|
||||
RUN pip install --no-cache-dir -r requirements.txt
|
||||
|
||||
# Copy the bot script and system prompt
|
||||
COPY /scripts/discordbot.py .
|
||||
COPY /scripts/system_prompt.txt .
|
||||
|
||||
# Run the bot on container start
|
||||
CMD ["python", "discordbot.py"]
|
||||
@@ -1,23 +1,214 @@
|
||||
# OpenWebUI-Discordbot
|
||||
# LiteLLM Discord Bot
|
||||
|
||||
A Discord bot that interfaces with an OpenWebUI instance to provide AI-powered responses in your Discord server.
|
||||
A Discord bot that interfaces with LiteLLM proxy to provide AI-powered responses in your Discord server. Supports multiple LLM providers through LiteLLM, conversation history management, image analysis, and configurable system prompts.
|
||||
|
||||
## Features
|
||||
|
||||
- 🤖 **LiteLLM Integration**: Use any LLM provider supported by LiteLLM (OpenAI, Anthropic, Google, local models, etc.)
|
||||
- 💬 **Conversation History**: Intelligent message history with token-aware truncation
|
||||
- 🖼️ **Image Support**: Analyze images attached to messages (for vision-capable models)
|
||||
- ⚙️ **Configurable System Prompts**: Customize bot behavior via file-based prompts
|
||||
- 🔄 **Async Architecture**: Efficient async/await design for responsive interactions
|
||||
- 🐳 **Docker Support**: Easy deployment with Docker
|
||||
|
||||
## Prerequisites
|
||||
|
||||
- Docker (for containerized deployment)
|
||||
- Python 3.8 or higher+ (for local development)
|
||||
- A Discord Bot Token ([How to create a Discord Bot Token](https://www.writebots.com/discord-bot-token/))
|
||||
- Access to an OpenWebUI instance
|
||||
- **Python 3.11+** (for local development) or **Docker** (for containerized deployment)
|
||||
- **Discord Bot Token** ([How to create one](https://www.writebots.com/discord-bot-token/))
|
||||
- **LiteLLM Proxy** instance running ([LiteLLM setup guide](https://docs.litellm.ai/docs/proxy/quick_start))
|
||||
|
||||
## Installation
|
||||
## Quick Start
|
||||
|
||||
##### Running locally
|
||||
### Option 1: Running with Docker (Recommended)
|
||||
|
||||
1. Clone the repository
|
||||
2. Copy `.env.sample` to `.env` and configure your environment variables:
|
||||
```env
|
||||
DISCORD_TOKEN=your_discord_bot_token
|
||||
OPENAI_API_KEY=your_openwebui_api_key
|
||||
OPENWEBUI_API_BASE=http://your_openwebui_instance:port/api
|
||||
MODEL_NAME=your_model_name
|
||||
1. Clone the repository:
|
||||
```bash
|
||||
git clone <repository-url>
|
||||
cd OpenWebUI-Discordbot
|
||||
```
|
||||
|
||||
2. Configure environment variables:
|
||||
```bash
|
||||
cd scripts
|
||||
cp .env.sample .env
|
||||
# Edit .env with your actual values
|
||||
```
|
||||
|
||||
3. Build and run with Docker:
|
||||
```bash
|
||||
docker build -t discord-bot .
|
||||
docker run --env-file scripts/.env discord-bot
|
||||
```
|
||||
|
||||
### Option 2: Running Locally
|
||||
|
||||
1. Clone the repository and navigate to scripts directory:
|
||||
```bash
|
||||
git clone <repository-url>
|
||||
cd OpenWebUI-Discordbot/scripts
|
||||
```
|
||||
|
||||
2. Install dependencies:
|
||||
```bash
|
||||
pip install -r requirements.txt
|
||||
```
|
||||
|
||||
3. Copy and configure environment variables:
|
||||
```bash
|
||||
cp .env.sample .env
|
||||
# Edit .env with your configuration
|
||||
```
|
||||
|
||||
4. Run the bot:
|
||||
```bash
|
||||
python discordbot.py
|
||||
```
|
||||
|
||||
## Configuration
|
||||
|
||||
### Environment Variables
|
||||
|
||||
Create a `.env` file in the `scripts/` directory with the following variables:
|
||||
|
||||
```env
|
||||
# Discord Bot Token - Get from https://discord.com/developers/applications
|
||||
DISCORD_TOKEN=your_discord_bot_token
|
||||
|
||||
# LiteLLM API Configuration
|
||||
LITELLM_API_KEY=sk-1234
|
||||
LITELLM_API_BASE=http://localhost:4000
|
||||
|
||||
# Model name (any model supported by your LiteLLM proxy)
|
||||
MODEL_NAME=gpt-4-turbo-preview
|
||||
|
||||
# System Prompt Configuration (optional)
|
||||
SYSTEM_PROMPT_FILE=./system_prompt.txt
|
||||
|
||||
# Maximum tokens to use for conversation history (optional, default: 3000)
|
||||
MAX_HISTORY_TOKENS=3000
|
||||
```
|
||||
|
||||
### System Prompt Customization
|
||||
|
||||
The bot's behavior is controlled by a system prompt file. Edit `scripts/system_prompt.txt` to customize how the bot responds:
|
||||
|
||||
```txt
|
||||
You are a helpful AI assistant integrated into Discord. Users will interact with you by mentioning you or sending direct messages.
|
||||
|
||||
Key behaviors:
|
||||
- Be concise and friendly in your responses
|
||||
- Use Discord markdown formatting when helpful (code blocks, bold, italics, etc.)
|
||||
- When users attach images, analyze them and provide relevant insights
|
||||
...
|
||||
```
|
||||
|
||||
## Setting Up LiteLLM Proxy
|
||||
|
||||
### Quick Setup (Local)
|
||||
|
||||
1. Install LiteLLM:
|
||||
```bash
|
||||
pip install litellm
|
||||
```
|
||||
|
||||
2. Run the proxy:
|
||||
```bash
|
||||
litellm --model gpt-4-turbo-preview --api_key YOUR_OPENAI_KEY
|
||||
# Or for local models:
|
||||
litellm --model ollama/llama3.2-vision
|
||||
```
|
||||
|
||||
### Production Setup (Docker)
|
||||
|
||||
```bash
|
||||
docker run -p 4000:4000 \
|
||||
-e OPENAI_API_KEY=your_key \
|
||||
ghcr.io/berriai/litellm:main-latest
|
||||
```
|
||||
|
||||
For advanced configuration, create a `litellm_config.yaml`:
|
||||
```yaml
|
||||
model_list:
|
||||
- model_name: gpt-4-turbo
|
||||
litellm_params:
|
||||
model: gpt-4-turbo-preview
|
||||
api_key: os.environ/OPENAI_API_KEY
|
||||
- model_name: claude
|
||||
litellm_params:
|
||||
model: claude-3-sonnet-20240229
|
||||
api_key: os.environ/ANTHROPIC_API_KEY
|
||||
```
|
||||
|
||||
Then run:
|
||||
```bash
|
||||
litellm --config litellm_config.yaml
|
||||
```
|
||||
|
||||
See [LiteLLM documentation](https://docs.litellm.ai/) for more details.
|
||||
|
||||
## Usage
|
||||
|
||||
### Triggering the Bot
|
||||
|
||||
The bot responds to:
|
||||
- **@mentions** in any channel where it has read access
|
||||
- **Direct messages (DMs)**
|
||||
|
||||
Example:
|
||||
```
|
||||
User: @BotName what's the weather like?
|
||||
Bot: I don't have access to real-time weather data, but I can help you with other questions!
|
||||
```
|
||||
|
||||
### Image Analysis
|
||||
|
||||
Attach images to your message (requires vision-capable model):
|
||||
```
|
||||
User: @BotName what's in this image? [image.png]
|
||||
Bot: The image shows a beautiful sunset over the ocean with...
|
||||
```
|
||||
|
||||
### Message History
|
||||
|
||||
The bot automatically maintains conversation context:
|
||||
- Retrieves recent relevant messages from the channel
|
||||
- Limits history based on token count (configurable via `MAX_HISTORY_TOKENS`)
|
||||
- Only includes messages where the bot was mentioned or bot's own responses
|
||||
|
||||
## Architecture Overview
|
||||
|
||||
### Key Improvements from OpenWebUI Version
|
||||
|
||||
1. **LiteLLM Integration**: Switched from OpenWebUI to LiteLLM for broader model support
|
||||
2. **Proper Conversation Format**: Messages use correct role attribution (system/user/assistant)
|
||||
3. **Token-Aware History**: Intelligent truncation to stay within model context limits
|
||||
4. **Async Image Downloads**: Uses `aiohttp` instead of synchronous `requests`
|
||||
5. **File-Based System Prompts**: Easy customization without code changes
|
||||
6. **Better Error Handling**: Improved error messages and validation
|
||||
|
||||
### Project Structure
|
||||
|
||||
```
|
||||
OpenWebUI-Discordbot/
|
||||
├── scripts/
|
||||
│ ├── discordbot.py # Main bot code (production)
|
||||
│ ├── system_prompt.txt # System prompt configuration
|
||||
│ ├── requirements.txt # Python dependencies
|
||||
│ └── .env.sample # Environment variable template
|
||||
├── v2/
|
||||
│ └── bot.py # Development/experimental version
|
||||
├── Dockerfile # Docker containerization
|
||||
├── README.md # This file
|
||||
└── claude.md # Development roadmap & upgrade notes
|
||||
```
|
||||
|
||||
## Upgrading from OpenWebUI
|
||||
|
||||
If you're upgrading from the previous OpenWebUI version:
|
||||
|
||||
1. **Update environment variables**: Rename `OPENWEBUI_API_BASE` → `LITELLM_API_BASE`, `OPENAI_API_KEY` → `LITELLM_API_KEY`
|
||||
2. **Set up LiteLLM proxy**: Follow setup instructions above
|
||||
3. **Install new dependencies**: Run `pip install -r requirements.txt`
|
||||
4. **Optional**: Customize `system_prompt.txt` for your use case
|
||||
|
||||
See `claude.md` for detailed upgrade documentation and future roadmap (MCP tools support, etc.).
|
||||
@@ -0,0 +1,649 @@
|
||||
# OpenWebUI Discord Bot - Upgrade Project
|
||||
|
||||
## Project Overview
|
||||
|
||||
This Discord bot currently interfaces with OpenWebUI to provide AI-powered responses. The goal is to upgrade it to:
|
||||
1. **Switch from OpenWebUI to LiteLLM Proxy** as the backend
|
||||
2. **Add MCP (Model Context Protocol) Tool Support**
|
||||
3. **Implement system prompt management within the application**
|
||||
|
||||
## Current Architecture
|
||||
|
||||
### Files Structure
|
||||
- **Main bot**: [v2/bot.py](v2/bot.py) - Current implementation
|
||||
- **Legacy bot**: [scripts/discordbot.py](scripts/discordbot.py) - Older version with slightly different approach
|
||||
- **Dependencies**: [v2/requirements.txt](v2/requirements.txt)
|
||||
- **Config**: [v2/.env.example](v2/.env.example)
|
||||
|
||||
### Current Implementation Details
|
||||
|
||||
#### Bot Features (v2/bot.py)
|
||||
- **Discord Integration**: Uses discord.py with message intents
|
||||
- **Trigger Methods**:
|
||||
- Bot mentions (@bot)
|
||||
- Direct messages (DMs)
|
||||
- **Message History**: Retrieves last 100 messages for context using `get_chat_history()`
|
||||
- **Image Support**: Downloads and encodes images as base64, sends to API
|
||||
- **API Client**: Uses OpenAI Python SDK pointing to OpenWebUI endpoint
|
||||
- **Message Format**: Embeds chat history in user message context
|
||||
|
||||
#### Current Message Flow
|
||||
1. User mentions bot or DMs it
|
||||
2. Bot fetches channel history (last 100 messages)
|
||||
3. Formats history as: `"AuthorName: message content"`
|
||||
4. Sends to OpenWebUI with format:
|
||||
```python
|
||||
{
|
||||
"role": "user",
|
||||
"content": [
|
||||
{"type": "text", "text": "##CONTEXT##\n{history}\n##ENDCONTEXT##\n\n{user_message}"},
|
||||
{"type": "image_url", "image_url": {...}} # if images present
|
||||
]
|
||||
}
|
||||
```
|
||||
5. Returns AI response and replies to user
|
||||
|
||||
#### Current Limitations
|
||||
- **No system prompt**: Context is embedded in user messages
|
||||
- **No tool calling**: Cannot execute functions or use MCPs
|
||||
- **OpenWebUI dependency**: Tightly coupled to OpenWebUI API structure
|
||||
- **Simple history**: Just text concatenation, no proper conversation threading
|
||||
- **Synchronous image download**: Uses `requests.get()` in async context (should use aiohttp)
|
||||
|
||||
## Target Architecture: LiteLLM + MCP Tools
|
||||
|
||||
### Why LiteLLM?
|
||||
|
||||
LiteLLM is a unified proxy that:
|
||||
- **Standardizes API calls** across 100+ LLM providers (OpenAI, Anthropic, Google, etc.)
|
||||
- **Native tool/function calling support** via OpenAI-compatible API
|
||||
- **Built-in MCP support** for Model Context Protocol tools
|
||||
- **Load balancing** and fallback between models
|
||||
- **Cost tracking** and usage analytics
|
||||
- **Streaming support** for real-time responses
|
||||
|
||||
### LiteLLM Tool Calling
|
||||
|
||||
LiteLLM supports the OpenAI tools format:
|
||||
```python
|
||||
response = client.chat.completions.create(
|
||||
model="gpt-4",
|
||||
messages=[...],
|
||||
tools=[{
|
||||
"type": "function",
|
||||
"function": {
|
||||
"name": "get_weather",
|
||||
"description": "Get current weather",
|
||||
"parameters": {...}
|
||||
}
|
||||
}],
|
||||
tool_choice="auto"
|
||||
)
|
||||
```
|
||||
|
||||
### MCP (Model Context Protocol) Overview
|
||||
|
||||
MCP is a standard protocol for:
|
||||
- **Exposing tools** to LLMs (functions they can call)
|
||||
- **Providing resources** (files, APIs, databases)
|
||||
- **Prompts/templates** for consistent interactions
|
||||
- **Sampling** for multi-step agentic behavior
|
||||
|
||||
**MCP Server Examples**:
|
||||
- `filesystem`: Read/write files
|
||||
- `github`: Access repos, create PRs
|
||||
- `postgres`: Query databases
|
||||
- `brave-search`: Web search
|
||||
- `slack`: Send messages, read channels
|
||||
|
||||
## Upgrade Plan
|
||||
|
||||
### Phase 1: Switch to LiteLLM Proxy
|
||||
|
||||
#### Configuration Changes
|
||||
1. Update environment variables:
|
||||
```env
|
||||
DISCORD_TOKEN=your_discord_bot_token
|
||||
LITELLM_API_KEY=your_litellm_api_key
|
||||
LITELLM_API_BASE=http://localhost:4000 # or your LiteLLM proxy URL
|
||||
MODEL_NAME=gpt-4-turbo-preview # or any LiteLLM-supported model
|
||||
SYSTEM_PROMPT=your_default_system_prompt # New!
|
||||
```
|
||||
|
||||
2. Keep using OpenAI SDK (LiteLLM is OpenAI-compatible):
|
||||
```python
|
||||
from openai import OpenAI
|
||||
|
||||
client = OpenAI(
|
||||
api_key=os.getenv('LITELLM_API_KEY'),
|
||||
base_url=os.getenv('LITELLM_API_BASE')
|
||||
)
|
||||
```
|
||||
|
||||
#### Message Format Refactor
|
||||
**Current approach** (embedding context in user message):
|
||||
```python
|
||||
text_content = f"##CONTEXT##\n{context}\n##ENDCONTEXT##\n\n{user_message}"
|
||||
messages = [{"role": "user", "content": text_content}]
|
||||
```
|
||||
|
||||
**New approach** (proper conversation history):
|
||||
```python
|
||||
messages = [
|
||||
{"role": "system", "content": SYSTEM_PROMPT},
|
||||
# ... previous conversation messages with proper roles ...
|
||||
{"role": "user", "content": user_message}
|
||||
]
|
||||
```
|
||||
|
||||
#### Benefits
|
||||
- Better model understanding of conversation structure
|
||||
- Separate system instructions from conversation
|
||||
- Proper role attribution (user vs assistant)
|
||||
- More efficient token usage
|
||||
|
||||
### Phase 2: Add System Prompt Management
|
||||
|
||||
#### Implementation Options
|
||||
|
||||
**Option A: Simple Environment Variable**
|
||||
- Store in `.env` file
|
||||
- Good for: Single, static system prompt
|
||||
- Example: `SYSTEM_PROMPT="You are a helpful Discord assistant..."`
|
||||
|
||||
**Option B: File-Based System Prompt**
|
||||
- Store in separate file (e.g., `system_prompt.txt`)
|
||||
- Good for: Long, complex prompts that need version control
|
||||
- Hot-reload capability
|
||||
|
||||
**Option C: Per-Channel/Per-Guild Prompts**
|
||||
- Store in JSON/database mapping channel_id → system_prompt
|
||||
- Good for: Multi-tenant bot with different personalities per server
|
||||
- Example:
|
||||
```json
|
||||
{
|
||||
"123456789": "You are a coding assistant...",
|
||||
"987654321": "You are a gaming buddy..."
|
||||
}
|
||||
```
|
||||
|
||||
**Option D: User-Configurable Prompts**
|
||||
- Discord slash commands to set/view system prompt
|
||||
- Store in SQLite/JSON
|
||||
- Commands: `/setprompt`, `/viewprompt`, `/resetprompt`
|
||||
|
||||
**Recommended**: Start with Option B (file-based), add Option D later for flexibility.
|
||||
|
||||
#### System Prompt Best Practices
|
||||
1. **Define bot personality**: Tone, style, formality
|
||||
2. **Set boundaries**: What bot should/shouldn't do
|
||||
3. **Provide context**: "You are in a Discord server, users will mention you"
|
||||
4. **Handle images**: "When users attach images, describe them..."
|
||||
5. **Tool usage guidance**: "Use available tools when appropriate"
|
||||
|
||||
Example system prompt:
|
||||
```
|
||||
You are a helpful AI assistant integrated into Discord. Users will interact with you by mentioning you or sending direct messages.
|
||||
|
||||
Key behaviors:
|
||||
- Be concise and friendly
|
||||
- Use Discord markdown formatting when helpful (code blocks, bold, etc.)
|
||||
- When users attach images, analyze them and provide relevant insights
|
||||
- You have access to various tools - use them when they would help answer the user's question
|
||||
- If you're unsure about something, say so
|
||||
- Keep track of conversation context
|
||||
|
||||
You are not a human, and you should not pretend to be one. Be honest about your capabilities and limitations.
|
||||
```
|
||||
|
||||
### Phase 3: Implement MCP Tool Support
|
||||
|
||||
#### LiteLLM MCP Integration
|
||||
|
||||
LiteLLM can connect to MCP servers in two ways:
|
||||
|
||||
**1. Via LiteLLM Proxy Configuration**
|
||||
Configure in `litellm_config.yaml`:
|
||||
```yaml
|
||||
model_list:
|
||||
- model_name: gpt-4-with-tools
|
||||
litellm_params:
|
||||
model: gpt-4-turbo-preview
|
||||
api_key: os.environ/OPENAI_API_KEY
|
||||
|
||||
mcp_servers:
|
||||
filesystem:
|
||||
command: npx
|
||||
args: [-y, @modelcontextprotocol/server-filesystem, /allowed/path]
|
||||
github:
|
||||
command: npx
|
||||
args: [-y, @modelcontextprotocol/server-github]
|
||||
env:
|
||||
GITHUB_TOKEN: ${GITHUB_TOKEN}
|
||||
```
|
||||
|
||||
**2. Via Direct Tool Definitions in Bot**
|
||||
Define tools manually in the bot code:
|
||||
```python
|
||||
tools = [
|
||||
{
|
||||
"type": "function",
|
||||
"function": {
|
||||
"name": "search_web",
|
||||
"description": "Search the web for information",
|
||||
"parameters": {
|
||||
"type": "object",
|
||||
"properties": {
|
||||
"query": {
|
||||
"type": "string",
|
||||
"description": "The search query"
|
||||
}
|
||||
},
|
||||
"required": ["query"]
|
||||
}
|
||||
}
|
||||
}
|
||||
]
|
||||
|
||||
response = client.chat.completions.create(
|
||||
model=MODEL_NAME,
|
||||
messages=messages,
|
||||
tools=tools,
|
||||
tool_choice="auto"
|
||||
)
|
||||
```
|
||||
|
||||
#### Tool Execution Flow
|
||||
|
||||
1. **Send message with tools available**:
|
||||
```python
|
||||
response = client.chat.completions.create(
|
||||
model=MODEL_NAME,
|
||||
messages=messages,
|
||||
tools=available_tools
|
||||
)
|
||||
```
|
||||
|
||||
2. **Check if model wants to use a tool**:
|
||||
```python
|
||||
if response.choices[0].message.tool_calls:
|
||||
for tool_call in response.choices[0].message.tool_calls:
|
||||
function_name = tool_call.function.name
|
||||
arguments = json.loads(tool_call.function.arguments)
|
||||
# Execute the function
|
||||
result = execute_tool(function_name, arguments)
|
||||
```
|
||||
|
||||
3. **Send tool results back to model**:
|
||||
```python
|
||||
messages.append({
|
||||
"role": "assistant",
|
||||
"content": None,
|
||||
"tool_calls": response.choices[0].message.tool_calls
|
||||
})
|
||||
messages.append({
|
||||
"role": "tool",
|
||||
"content": json.dumps(result),
|
||||
"tool_call_id": tool_call.id
|
||||
})
|
||||
|
||||
# Get final response
|
||||
final_response = client.chat.completions.create(
|
||||
model=MODEL_NAME,
|
||||
messages=messages,
|
||||
tools=available_tools
|
||||
)
|
||||
```
|
||||
|
||||
4. **Return final response to user**
|
||||
|
||||
#### Tool Implementation Patterns
|
||||
|
||||
**Pattern 1: Bot-Managed Tools**
|
||||
Implement tools directly in the bot:
|
||||
```python
|
||||
async def search_web(query: str) -> str:
|
||||
"""Execute web search"""
|
||||
# Use requests/aiohttp to call search API
|
||||
pass
|
||||
|
||||
async def get_weather(location: str) -> str:
|
||||
"""Get weather for location"""
|
||||
# Call weather API
|
||||
pass
|
||||
|
||||
AVAILABLE_TOOLS = {
|
||||
"search_web": search_web,
|
||||
"get_weather": get_weather,
|
||||
}
|
||||
|
||||
async def execute_tool(name: str, arguments: dict) -> str:
|
||||
if name in AVAILABLE_TOOLS:
|
||||
return await AVAILABLE_TOOLS[name](**arguments)
|
||||
return "Tool not found"
|
||||
```
|
||||
|
||||
**Pattern 2: MCP Server Proxy**
|
||||
Let LiteLLM proxy handle MCP servers (recommended):
|
||||
- Configure MCP servers in LiteLLM config
|
||||
- LiteLLM automatically exposes them as tools
|
||||
- Bot just passes tool calls through
|
||||
- Simpler bot code, more scalable
|
||||
|
||||
**Pattern 3: Hybrid**
|
||||
- Common tools via LiteLLM proxy MCP
|
||||
- Discord-specific tools in bot (e.g., "get_server_info", "list_channels")
|
||||
|
||||
#### Recommended Starter Tools
|
||||
|
||||
1. **Web Search** (via Brave/Google MCP server)
|
||||
- Let bot search for current information
|
||||
|
||||
2. **File Operations** (via filesystem MCP server - with restrictions!)
|
||||
- Read documentation, configs
|
||||
- Useful in developer-focused servers
|
||||
|
||||
3. **Wikipedia** (via wikipedia MCP server)
|
||||
- Factual information lookup
|
||||
|
||||
4. **Time/Date** (custom function)
|
||||
- Simple, no external dependency
|
||||
|
||||
5. **Discord Server Info** (custom function)
|
||||
- Get channel list, member count, server info
|
||||
- Discord-specific utility
|
||||
|
||||
### Phase 4: Improve Message History Management
|
||||
|
||||
#### Current Issues
|
||||
- Fetches all messages every time (inefficient)
|
||||
- No conversation threading (treats all channel messages as one context)
|
||||
- No token limit awareness
|
||||
- Channel history might contain irrelevant conversations
|
||||
|
||||
#### Improvements
|
||||
|
||||
**1. Per-Conversation Threading**
|
||||
```python
|
||||
# Track conversations by thread or by user
|
||||
conversation_storage = {
|
||||
"channel_id:user_id": [
|
||||
{"role": "user", "content": "..."},
|
||||
{"role": "assistant", "content": "..."},
|
||||
]
|
||||
}
|
||||
```
|
||||
|
||||
**2. Token-Aware History Truncation**
|
||||
```python
|
||||
def trim_history(messages, max_tokens=4000):
|
||||
"""Keep only recent messages that fit in token budget"""
|
||||
# Use tiktoken to count tokens
|
||||
# Remove oldest messages until under limit
|
||||
pass
|
||||
```
|
||||
|
||||
**3. Message Deduplication**
|
||||
Only include messages directly related to bot conversations:
|
||||
- Messages mentioning bot
|
||||
- Bot's responses
|
||||
- Optionally: X messages before each bot mention for context
|
||||
|
||||
**4. Caching & Persistence**
|
||||
- Cache conversation history in memory
|
||||
- Optional: Persist to SQLite/Redis for bot restarts
|
||||
- Clear old conversations after inactivity
|
||||
|
||||
## Implementation Checklist
|
||||
|
||||
### Preparation
|
||||
- [ ] Set up LiteLLM proxy locally or remotely
|
||||
- [ ] Configure LiteLLM with desired model(s)
|
||||
- [ ] Decide on MCP servers to enable
|
||||
- [ ] Design system prompt strategy
|
||||
- [ ] Review token limits for target models
|
||||
|
||||
### Code Changes
|
||||
|
||||
#### File: v2/bot.py
|
||||
- [ ] Update imports (add `json`, improve `aiohttp` usage)
|
||||
- [ ] Change environment variables:
|
||||
- [ ] `OPENWEBUI_API_BASE` → `LITELLM_API_BASE`
|
||||
- [ ] Add `SYSTEM_PROMPT` or `SYSTEM_PROMPT_FILE`
|
||||
- [ ] Update OpenAI client initialization
|
||||
- [ ] Refactor `get_ai_response()`:
|
||||
- [ ] Add system message
|
||||
- [ ] Convert history to proper message format (alternating user/assistant)
|
||||
- [ ] Add tool support parameters
|
||||
- [ ] Implement tool execution loop
|
||||
- [ ] Refactor `get_chat_history()`:
|
||||
- [ ] Return structured messages instead of text concatenation
|
||||
- [ ] Filter for bot-relevant messages
|
||||
- [ ] Add token counting/truncation
|
||||
- [ ] Fix `download_image()` to use aiohttp instead of requests
|
||||
- [ ] Add tool definition functions
|
||||
- [ ] Add tool execution handler
|
||||
- [ ] Add error handling for tool failures
|
||||
|
||||
#### New File: v2/tools.py (optional)
|
||||
- [ ] Define tool schemas
|
||||
- [ ] Implement tool execution functions
|
||||
- [ ] Export tool registry
|
||||
|
||||
#### New File: v2/system_prompt.txt or system_prompts.json
|
||||
- [ ] Write default system prompt
|
||||
- [ ] Optional: Add per-guild prompts
|
||||
|
||||
#### File: v2/requirements.txt
|
||||
- [ ] Keep: `discord.py`, `openai`, `python-dotenv`
|
||||
- [ ] Add: `aiohttp` (if not using requests), `tiktoken` (for token counting)
|
||||
- [ ] Optional: `anthropic` (if using Claude directly), `litellm` (if using SDK directly)
|
||||
|
||||
#### File: v2/.env.example
|
||||
- [ ] Update variable names
|
||||
- [ ] Add system prompt variables
|
||||
- [ ] Document new configuration options
|
||||
|
||||
### Testing
|
||||
- [ ] Test basic message responses (no tools)
|
||||
- [ ] Test with images attached
|
||||
- [ ] Test tool calling with simple tool (e.g., get_time)
|
||||
- [ ] Test tool calling with external MCP server
|
||||
- [ ] Test conversation threading
|
||||
- [ ] Test token limit handling
|
||||
- [ ] Test error scenarios (API down, tool failure, etc.)
|
||||
- [ ] Test in multiple Discord servers/channels
|
||||
|
||||
### Documentation
|
||||
- [ ] Update README.md with new setup instructions
|
||||
- [ ] Document LiteLLM proxy setup
|
||||
- [ ] Document MCP server configuration
|
||||
- [ ] Add example system prompts
|
||||
- [ ] Document available tools
|
||||
- [ ] Add troubleshooting section
|
||||
|
||||
## Technical Considerations
|
||||
|
||||
### Token Management
|
||||
- Most models have 4k-128k token context windows
|
||||
- Message history can quickly consume tokens
|
||||
- Reserve tokens for:
|
||||
- System prompt: ~500-1000 tokens
|
||||
- Tool definitions: ~100-500 tokens per tool
|
||||
- Response: ~1000-2000 tokens
|
||||
- History: remaining tokens
|
||||
|
||||
### Rate Limiting
|
||||
- Discord: 5 requests per 5 seconds per channel
|
||||
- LLM APIs: Varies by provider (OpenAI: ~3500 RPM for GPT-4)
|
||||
- Implement queuing if needed
|
||||
|
||||
### Error Handling
|
||||
- API timeouts: Retry with exponential backoff
|
||||
- Tool execution failures: Return error message to model
|
||||
- Discord API errors: Log and notify user
|
||||
- Invalid tool calls: Validate before execution
|
||||
|
||||
### Security Considerations
|
||||
- **Tool access control**: Don't expose dangerous tools (file delete, system commands)
|
||||
- **Input validation**: Sanitize tool arguments
|
||||
- **Rate limiting**: Prevent abuse of expensive tools (web search)
|
||||
- **API key security**: Never log or expose API keys
|
||||
- **MCP filesystem access**: Restrict to safe directories only
|
||||
|
||||
### Cost Optimization
|
||||
- Use smaller models for simple queries (gpt-3.5-turbo)
|
||||
- Implement streaming for better UX
|
||||
- Cache common queries
|
||||
- Trim history aggressively
|
||||
- Consider LiteLLM's caching features
|
||||
|
||||
## Future Enhancements
|
||||
|
||||
### Short Term
|
||||
- [ ] Add slash commands for bot configuration
|
||||
- [ ] Implement conversation reset command
|
||||
- [ ] Add support for Discord threads
|
||||
- [ ] Stream responses for long outputs
|
||||
- [ ] Add reaction-based tool approval (user confirms before execution)
|
||||
|
||||
### Medium Term
|
||||
- [ ] Multi-modal support (voice, more image formats)
|
||||
- [ ] Per-user conversation isolation
|
||||
- [ ] Tool usage analytics and logging
|
||||
- [ ] Custom MCP server for Discord-specific tools
|
||||
- [ ] Web dashboard for bot management
|
||||
|
||||
### Long Term
|
||||
- [ ] Agentic workflows (multi-step tool usage)
|
||||
- [ ] Memory/RAG for long-term context
|
||||
- [ ] Multiple bot personalities per server
|
||||
- [ ] Integration with Discord's scheduled events
|
||||
- [ ] Voice channel integration (TTS/STT)
|
||||
|
||||
## Resources
|
||||
|
||||
### Documentation
|
||||
- **LiteLLM Docs**: https://docs.litellm.ai/
|
||||
- **LiteLLM Tools/Functions**: https://docs.litellm.ai/docs/completion/function_call
|
||||
- **MCP Specification**: https://modelcontextprotocol.io/
|
||||
- **MCP Server Examples**: https://github.com/modelcontextprotocol/servers
|
||||
- **Discord.py Docs**: https://discordpy.readthedocs.io/
|
||||
- **OpenAI API Docs**: https://platform.openai.com/docs/guides/function-calling
|
||||
|
||||
### Example MCP Servers
|
||||
- `@modelcontextprotocol/server-filesystem`: File operations
|
||||
- `@modelcontextprotocol/server-github`: GitHub integration
|
||||
- `@modelcontextprotocol/server-postgres`: Database queries
|
||||
- `@modelcontextprotocol/server-brave-search`: Web search
|
||||
- `@modelcontextprotocol/server-slack`: Slack integration
|
||||
- `@modelcontextprotocol/server-memory`: Persistent memory
|
||||
|
||||
### Tools for Development
|
||||
- **tiktoken**: Token counting (OpenAI tokenizer)
|
||||
- **litellm CLI**: `litellm --model gpt-4 --drop_params` for testing
|
||||
- **Postman**: Test LiteLLM API endpoints
|
||||
- **Docker**: Containerize LiteLLM proxy
|
||||
|
||||
## Questions to Resolve
|
||||
|
||||
1. **Which LiteLLM deployment?**
|
||||
- Self-hosted proxy (more control, more maintenance)
|
||||
- Hosted service (easier, potential cost)
|
||||
|
||||
2. **Which models to support?**
|
||||
- Single model (simpler)
|
||||
- Multiple models with fallback (more robust)
|
||||
- User-selectable models (more flexible)
|
||||
|
||||
3. **MCP server hosting?**
|
||||
- Same machine as bot
|
||||
- Separate server
|
||||
- Cloud functions
|
||||
|
||||
4. **System prompt strategy?**
|
||||
- Single global prompt
|
||||
- Per-guild prompts
|
||||
- User-configurable
|
||||
|
||||
5. **Tool approval flow?**
|
||||
- Automatic execution (faster but riskier)
|
||||
- User confirmation for sensitive tools (safer but slower)
|
||||
|
||||
6. **Conversation persistence?**
|
||||
- In-memory only (simple, lost on restart)
|
||||
- SQLite (persistent, moderate complexity)
|
||||
- Redis (distributed, more setup)
|
||||
|
||||
## Current Code Analysis
|
||||
|
||||
### v2/bot.py Strengths
|
||||
- Clean, simple structure
|
||||
- Proper async/await usage
|
||||
- Good image handling
|
||||
- Type hints in newer version
|
||||
|
||||
### v2/bot.py Issues to Fix
|
||||
- Line 44: Using synchronous `requests.get()` in async function
|
||||
- Lines 62-77: Embedding history in user message instead of proper conversation format
|
||||
- Line 41: `channel_history` dict declared but never used
|
||||
- No error handling for OpenAI API errors besides generic try/catch
|
||||
- No rate limiting
|
||||
- No conversation threading
|
||||
- History includes ALL channel messages, not just bot-relevant ones
|
||||
- No system prompt support
|
||||
|
||||
### scripts/discordbot.py Differences
|
||||
- Has system message (line 67) - better approach!
|
||||
- Slightly different message structure
|
||||
- Otherwise similar implementation
|
||||
|
||||
## Recommended Migration Path
|
||||
|
||||
**Step 1**: Quick wins (minimal changes)
|
||||
1. Add system prompt support using `scripts/discordbot.py` pattern
|
||||
2. Fix async image download (use aiohttp)
|
||||
3. Update env vars and client to point to LiteLLM
|
||||
|
||||
**Step 2**: Core refactor (moderate changes)
|
||||
1. Refactor message history to proper conversation format
|
||||
2. Implement token-aware history truncation
|
||||
3. Add basic tool support infrastructure
|
||||
|
||||
**Step 3**: Tool integration (significant changes)
|
||||
1. Define initial tool set
|
||||
2. Implement tool execution loop
|
||||
3. Add error handling for tool failures
|
||||
|
||||
**Step 4**: Polish (incremental improvements)
|
||||
1. Add slash commands for configuration
|
||||
2. Improve conversation management
|
||||
3. Add monitoring and logging
|
||||
|
||||
This approach allows you to test at each step and provides incremental value.
|
||||
|
||||
---
|
||||
|
||||
## Getting Started
|
||||
|
||||
When you're ready to begin implementation:
|
||||
|
||||
1. **Set up LiteLLM proxy**:
|
||||
```bash
|
||||
pip install litellm
|
||||
litellm --model gpt-4 --drop_params
|
||||
# Or use Docker: docker run -p 4000:4000 ghcr.io/berriai/litellm:main
|
||||
```
|
||||
|
||||
2. **Test LiteLLM endpoint**:
|
||||
```bash
|
||||
curl -X POST http://localhost:4000/v1/chat/completions \
|
||||
-H "Content-Type: application/json" \
|
||||
-d '{"model": "gpt-4", "messages": [{"role": "user", "content": "Hello!"}]}'
|
||||
```
|
||||
|
||||
3. **Start with system prompt**: Implement system prompt support first as low-risk improvement
|
||||
|
||||
4. **Iterate on tools**: Start with one simple tool, then expand
|
||||
|
||||
Let me know which phase you'd like to tackle first!
|
||||
@@ -0,0 +1,408 @@
|
||||
# LiteLLM Responses API with MCP tool integration
|
||||
|
||||
LiteLLM's `/v1/responses` endpoint enables automatic MCP tool execution through a single API call, eliminating the manual tool-calling loop required with chat.completions. When configured with `"require_approval": "never"`, LiteLLM handles tool discovery, execution, and response integration automatically—making Discord bot migration straightforward. The key differences from chat.completions are the `input` parameter (replacing `messages`) and native MCP tool support via a `"type": "mcp"` tool specification.
|
||||
|
||||
## Request and response format for /v1/responses
|
||||
|
||||
The Responses API (available in LiteLLM **1.63.8+**) uses `input` instead of `messages`. The `input` parameter accepts either a simple string or an array of message objects:
|
||||
|
||||
```python
|
||||
# Simple string input
|
||||
response = client.responses.create(
|
||||
model="anthropic/claude-3-5-sonnet-latest",
|
||||
input="What is the weather today?"
|
||||
)
|
||||
|
||||
# Array format (for multi-turn conversations)
|
||||
response = client.responses.create(
|
||||
model="anthropic/claude-3-5-sonnet-latest",
|
||||
input=[
|
||||
{"role": "user", "content": "Hello"},
|
||||
{"role": "assistant", "content": "Hi there!"},
|
||||
{"role": "user", "content": "Tell me about Python"}
|
||||
]
|
||||
)
|
||||
```
|
||||
|
||||
**Response structure** differs significantly from chat.completions. Instead of `choices[0].message.content`, responses use an `output` array:
|
||||
|
||||
```json
|
||||
{
|
||||
"id": "resp_abc123",
|
||||
"object": "response",
|
||||
"created_at": 1734366691,
|
||||
"status": "completed",
|
||||
"model": "claude-3-5-sonnet-latest",
|
||||
"output": [
|
||||
{
|
||||
"type": "message",
|
||||
"id": "msg_abc123",
|
||||
"status": "completed",
|
||||
"role": "assistant",
|
||||
"content": [
|
||||
{
|
||||
"type": "output_text",
|
||||
"text": "Here is the response text...",
|
||||
"annotations": []
|
||||
}
|
||||
]
|
||||
}
|
||||
],
|
||||
"usage": {"input_tokens": 18, "output_tokens": 98, "total_tokens": 116}
|
||||
}
|
||||
```
|
||||
|
||||
To extract text: `response.output[0].content[0].text`
|
||||
|
||||
## MCP tool specification format
|
||||
|
||||
MCP tools use `"type": "mcp"` with three critical parameters: `server_label`, `server_url`, and `require_approval`. The special value `"server_url": "litellm_proxy"` tells LiteLLM to act as an MCP gateway, handling all tool execution internally:
|
||||
|
||||
```python
|
||||
tools=[
|
||||
{
|
||||
"type": "mcp",
|
||||
"server_label": "my_mcp_server", # Identifier for the MCP server
|
||||
"server_url": "litellm_proxy", # LiteLLM handles MCP bridging
|
||||
"require_approval": "never", # Automatic execution
|
||||
"allowed_tools": ["tool1", "tool2"] # Optional: restrict available tools
|
||||
}
|
||||
]
|
||||
```
|
||||
|
||||
| Parameter | Purpose |
|
||||
|-----------|---------|
|
||||
| `server_label` | Identifies which configured MCP server to use (must match config.yaml) |
|
||||
| `server_url` | `"litellm_proxy"` for LiteLLM gateway, or direct URL like `"https://mcp.example.com/mcp"` |
|
||||
| `require_approval` | `"never"` for automatic execution; omit for approval-based flow |
|
||||
| `allowed_tools` | Whitelist of tool names to make available |
|
||||
|
||||
When `server_url="litellm_proxy"`, LiteLLM performs a **four-step automatic flow**: (1) fetches MCP tools and converts to OpenAI format, (2) sends tools to the LLM with your input, (3) executes any tool calls against MCP servers, and (4) returns the final response with tool results integrated.
|
||||
|
||||
## Streaming versus non-streaming responses
|
||||
|
||||
For **non-streaming**, pass `stream=False` (default) and receive the complete response object:
|
||||
|
||||
```python
|
||||
response = client.responses.create(
|
||||
model="gpt-4o",
|
||||
input="Hello",
|
||||
stream=False
|
||||
)
|
||||
text = response.output[0].content[0].text
|
||||
```
|
||||
|
||||
For **streaming**, set `stream=True` and iterate over events:
|
||||
|
||||
```python
|
||||
stream = client.responses.create(
|
||||
model="gpt-4o",
|
||||
input="Write a poem",
|
||||
stream=True
|
||||
)
|
||||
|
||||
full_text = ""
|
||||
for event in stream:
|
||||
if hasattr(event, 'type'):
|
||||
if event.type == "response.output_text.delta":
|
||||
print(event.delta, end="", flush=True)
|
||||
full_text += event.delta
|
||||
elif event.type == "response.completed":
|
||||
print("\n--- Done ---")
|
||||
```
|
||||
|
||||
Key streaming event types include `response.created`, `response.output_text.delta` (incremental text), `response.output_text.done`, and `response.completed`.
|
||||
|
||||
## Python SDK differences between responses.create() and chat.completions.create()
|
||||
|
||||
| Aspect | `responses.create()` | `chat.completions.create()` |
|
||||
|--------|---------------------|---------------------------|
|
||||
| Input parameter | `input` (string or array) | `messages` (array required) |
|
||||
| Response access | `response.output[0].content[0].text` | `response.choices[0].message.content` |
|
||||
| Conversation history | Built-in via `previous_response_id` | Manual message array management |
|
||||
| MCP tools | Native `"type": "mcp"` support | Standard function calling only |
|
||||
| Endpoint | `/v1/responses` | `/v1/chat/completions` |
|
||||
|
||||
**Client setup** is identical for both APIs:
|
||||
|
||||
```python
|
||||
from openai import OpenAI
|
||||
|
||||
client = OpenAI(
|
||||
base_url="http://localhost:4000", # Your LiteLLM proxy
|
||||
api_key="sk-your-litellm-key"
|
||||
)
|
||||
|
||||
# Responses API
|
||||
response = client.responses.create(model="gpt-4o", input="Hello")
|
||||
|
||||
# Chat Completions API (old way)
|
||||
response = client.chat.completions.create(
|
||||
model="gpt-4o",
|
||||
messages=[{"role": "user", "content": "Hello"}]
|
||||
)
|
||||
```
|
||||
|
||||
## Conversation history with the input parameter
|
||||
|
||||
Unlike chat.completions where you manually pass the full message history each time, the Responses API offers two approaches:
|
||||
|
||||
**Option 1: Use `previous_response_id`** for automatic context (recommended):
|
||||
```python
|
||||
# First message
|
||||
response1 = client.responses.create(model="gpt-4o", input="My name is Alice")
|
||||
|
||||
# Follow-up with context preserved automatically
|
||||
response2 = client.responses.create(
|
||||
model="gpt-4o",
|
||||
input="What's my name?",
|
||||
previous_response_id=response1.id # LiteLLM maintains context
|
||||
)
|
||||
```
|
||||
|
||||
**Option 2: Pass full history in input array** (manual approach):
|
||||
```python
|
||||
response = client.responses.create(
|
||||
model="gpt-4o",
|
||||
input=[
|
||||
{"role": "user", "content": "My name is Alice"},
|
||||
{"role": "assistant", "content": "Nice to meet you, Alice!"},
|
||||
{"role": "user", "content": "What's my name?"}
|
||||
]
|
||||
)
|
||||
```
|
||||
|
||||
The `input` array supports roles: `user`, `assistant`, `developer` (replaces `system` in newer models), and `tool`.
|
||||
|
||||
## The require_approval parameter and MCP options
|
||||
|
||||
**`require_approval: "never"`** enables fully automatic tool execution—LiteLLM returns the final response in a single API call:
|
||||
|
||||
```python
|
||||
response = client.responses.create(
|
||||
model="gpt-4o",
|
||||
input="Search for Python documentation",
|
||||
tools=[{
|
||||
"type": "mcp",
|
||||
"server_label": "search_server",
|
||||
"server_url": "litellm_proxy",
|
||||
"require_approval": "never" # No approval needed
|
||||
}]
|
||||
)
|
||||
# Response includes tool results integrated into final answer
|
||||
```
|
||||
|
||||
**Without `require_approval: "never"`**, you get an approval flow requiring two API calls:
|
||||
|
||||
```python
|
||||
# Step 1: Get approval request
|
||||
response = client.responses.create(
|
||||
model="gpt-4o",
|
||||
input="Search for docs",
|
||||
tools=[{"type": "mcp", "server_label": "search", "server_url": "litellm_proxy"}]
|
||||
)
|
||||
|
||||
# Extract approval request ID from response.output
|
||||
approval_id = None
|
||||
for output in response.output:
|
||||
if output.type == "mcp_approval_request":
|
||||
approval_id = output.id
|
||||
break
|
||||
|
||||
# Step 2: Approve and get final response
|
||||
final_response = client.responses.create(
|
||||
model="gpt-4o",
|
||||
input=[{"type": "mcp_approval_response", "approve": True, "approval_request_id": approval_id}],
|
||||
previous_response_id=response.id,
|
||||
tools=[{"type": "mcp", "server_label": "search", "server_url": "litellm_proxy"}]
|
||||
)
|
||||
```
|
||||
|
||||
## Restricting tools with allowed_tools
|
||||
|
||||
Control which MCP tools are available at **request time** or **server configuration level**:
|
||||
|
||||
**Request-level restriction** (per-call):
|
||||
```python
|
||||
tools=[{
|
||||
"type": "mcp",
|
||||
"server_label": "github_mcp",
|
||||
"server_url": "litellm_proxy",
|
||||
"require_approval": "never",
|
||||
"allowed_tools": ["list_repos", "get_file_contents"] # Only these tools available
|
||||
}]
|
||||
```
|
||||
|
||||
**Server-level restriction** (in config.yaml):
|
||||
```yaml
|
||||
mcp_servers:
|
||||
github_mcp:
|
||||
url: "https://api.github.com/mcp"
|
||||
allowed_tools: ["list_repos", "get_file_contents"] # Whitelist
|
||||
disallowed_tools: ["delete_repo", "force_push"] # Blacklist
|
||||
```
|
||||
|
||||
If both `allowed_tools` and `disallowed_tools` are specified, `allowed_tools` takes priority.
|
||||
|
||||
## Authentication headers
|
||||
|
||||
LiteLLM supports multiple authentication header formats:
|
||||
|
||||
| Header | Use Case |
|
||||
|--------|----------|
|
||||
| `Authorization: Bearer sk-...` | **Standard** - Used by OpenAI SDK automatically |
|
||||
| `x-litellm-api-key: Bearer sk-...` | **MCP connections** and custom scenarios |
|
||||
| `api-key: ...` | Azure OpenAI compatibility |
|
||||
|
||||
**For standard API calls** (Discord bot), use the OpenAI SDK default:
|
||||
```python
|
||||
client = OpenAI(
|
||||
base_url="http://localhost:4000",
|
||||
api_key="sk-your-key" # Sent as "Authorization: Bearer sk-your-key"
|
||||
)
|
||||
```
|
||||
|
||||
**For MCP tool headers** (when calling external MCP servers), use the `headers` parameter:
|
||||
```python
|
||||
tools=[{
|
||||
"type": "mcp",
|
||||
"server_label": "github",
|
||||
"server_url": "litellm_proxy",
|
||||
"require_approval": "never",
|
||||
"headers": {
|
||||
"x-litellm-api-key": "Bearer sk-your-litellm-key",
|
||||
"x-mcp-github-authorization": "Bearer ghp_your_github_token"
|
||||
}
|
||||
}]
|
||||
```
|
||||
|
||||
## Complete Discord bot migration example
|
||||
|
||||
Here's a full implementation pattern for migrating from chat.completions to responses with MCP:
|
||||
|
||||
```python
|
||||
from openai import OpenAI
|
||||
import os
|
||||
|
||||
class LiteLLMResponsesClient:
|
||||
"""Client wrapper for Discord bot using LiteLLM Responses API with MCP."""
|
||||
|
||||
def __init__(self, proxy_url: str, api_key: str):
|
||||
self.client = OpenAI(base_url=proxy_url, api_key=api_key)
|
||||
self.conversations = {} # user_id -> response_id mapping
|
||||
|
||||
def get_mcp_tools(self, server_label: str = "default") -> list:
|
||||
"""Define MCP tools configuration."""
|
||||
return [{
|
||||
"type": "mcp",
|
||||
"server_label": server_label,
|
||||
"server_url": "litellm_proxy",
|
||||
"require_approval": "never",
|
||||
"allowed_tools": ["search", "fetch_data", "analyze"] # Customize as needed
|
||||
}]
|
||||
|
||||
def chat(
|
||||
self,
|
||||
user_id: str,
|
||||
message: str,
|
||||
model: str = "anthropic/claude-3-5-sonnet-latest",
|
||||
use_mcp_tools: bool = True,
|
||||
stream: bool = False
|
||||
):
|
||||
"""Send a message and get response, with optional MCP tools and streaming."""
|
||||
|
||||
previous_id = self.conversations.get(user_id)
|
||||
|
||||
kwargs = {
|
||||
"model": model,
|
||||
"input": message,
|
||||
"stream": stream
|
||||
}
|
||||
|
||||
if previous_id:
|
||||
kwargs["previous_response_id"] = previous_id
|
||||
|
||||
if use_mcp_tools:
|
||||
kwargs["tools"] = self.get_mcp_tools()
|
||||
kwargs["tool_choice"] = "auto"
|
||||
|
||||
if stream:
|
||||
return self._handle_stream(user_id, **kwargs)
|
||||
else:
|
||||
response = self.client.responses.create(**kwargs)
|
||||
self.conversations[user_id] = response.id
|
||||
return self._extract_text(response)
|
||||
|
||||
def _handle_stream(self, user_id: str, **kwargs):
|
||||
"""Generator for streaming responses."""
|
||||
stream = self.client.responses.create(**kwargs)
|
||||
response_id = None
|
||||
|
||||
for event in stream:
|
||||
if hasattr(event, 'type'):
|
||||
if event.type == "response.created":
|
||||
response_id = event.response.id
|
||||
elif event.type == "response.output_text.delta":
|
||||
yield event.delta
|
||||
|
||||
if response_id:
|
||||
self.conversations[user_id] = response_id
|
||||
|
||||
def _extract_text(self, response) -> str:
|
||||
"""Extract text from Responses API response."""
|
||||
for output in response.output:
|
||||
if output.type == "message":
|
||||
for content in output.content:
|
||||
if content.type == "output_text":
|
||||
return content.text
|
||||
return ""
|
||||
|
||||
def clear_history(self, user_id: str):
|
||||
"""Clear conversation history for a user."""
|
||||
self.conversations.pop(user_id, None)
|
||||
|
||||
|
||||
# Discord bot integration example
|
||||
import discord
|
||||
|
||||
bot = discord.Bot()
|
||||
llm_client = LiteLLMResponsesClient(
|
||||
proxy_url=os.environ["LITELLM_PROXY_URL"],
|
||||
api_key=os.environ["LITELLM_API_KEY"]
|
||||
)
|
||||
|
||||
@bot.event
|
||||
async def on_message(message):
|
||||
if message.author.bot:
|
||||
return
|
||||
|
||||
if bot.user.mentioned_in(message):
|
||||
user_id = str(message.author.id)
|
||||
user_message = message.content.replace(f'<@{bot.user.id}>', '').strip()
|
||||
|
||||
# Non-streaming response with MCP tools
|
||||
response_text = llm_client.chat(
|
||||
user_id=user_id,
|
||||
message=user_message,
|
||||
use_mcp_tools=True
|
||||
)
|
||||
await message.reply(response_text)
|
||||
|
||||
# Run: bot.run(os.environ["DISCORD_TOKEN"])
|
||||
```
|
||||
|
||||
## Official documentation links
|
||||
|
||||
- **Responses API documentation**: https://docs.litellm.ai/docs/response_api
|
||||
- **MCP overview**: https://docs.litellm.ai/docs/mcp
|
||||
- **MCP usage guide**: https://docs.litellm.ai/docs/mcp_usage
|
||||
- **MCP permission management**: https://docs.litellm.ai/docs/mcp_control
|
||||
- **OpenAI provider Responses API**: https://docs.litellm.ai/docs/providers/openai/responses_api
|
||||
- **Streaming documentation**: https://docs.litellm.ai/docs/completion/stream
|
||||
- **Virtual keys and auth**: https://docs.litellm.ai/docs/proxy/virtual_keys
|
||||
|
||||
## Key migration considerations
|
||||
|
||||
The Responses API is marked as **BETA** in LiteLLM. Ensure you're running LiteLLM **1.63.8+** and using OpenAI SDK **1.66.1+** for full compatibility. Model names must include the provider prefix (e.g., `openai/gpt-4o`, `anthropic/claude-3-5-sonnet-latest`). Response IDs are encrypted per-user by default for security—users cannot access other users' conversation history unless you disable this with `disable_responses_id_security: true` in config.yaml.
|
||||
|
||||
The primary advantage for Discord bots is the automatic MCP tool execution loop. With chat.completions, you must manually detect tool calls, execute them, and send results back. With Responses API and `require_approval: "never"`, LiteLLM handles this entire flow internally, returning the final integrated response in a single call.
|
||||
+23
-3
@@ -1,4 +1,24 @@
|
||||
# Discord Bot Token - Get from https://discord.com/developers/applications
|
||||
DISCORD_TOKEN=your_discord_bot_token
|
||||
OPENAI_API_KEY=your_openwebui_api_key
|
||||
OPENWEBUI_API_BASE=http://your_openwebui_instance:port/api
|
||||
MODEL_NAME="Your_Model_Name"
|
||||
|
||||
# LiteLLM API Configuration
|
||||
LITELLM_API_KEY=sk-1234
|
||||
LITELLM_API_BASE=http://localhost:4000
|
||||
|
||||
# Model name (any model supported by your LiteLLM proxy)
|
||||
MODEL_NAME=gpt-4-turbo-preview
|
||||
|
||||
# System Prompt Configuration (optional)
|
||||
SYSTEM_PROMPT_FILE=./system_prompt.txt
|
||||
|
||||
# Maximum tokens to use for conversation history (optional, default: 3000)
|
||||
MAX_HISTORY_TOKENS=3000
|
||||
|
||||
# Enable debug logging (optional, default: false)
|
||||
# Set to 'true' to see detailed logs for troubleshooting
|
||||
DEBUG_LOGGING=false
|
||||
|
||||
# Enable MCP tools integration (optional, default: false)
|
||||
# Set to 'true' to allow the bot to use tools configured in your LiteLLM proxy
|
||||
# Tools are auto-executed without user confirmation
|
||||
ENABLE_TOOLS=false
|
||||
+510
-83
@@ -1,134 +1,561 @@
|
||||
import os
|
||||
import discord
|
||||
from discord.ext import commands
|
||||
from openai import OpenAI
|
||||
import os
|
||||
import base64
|
||||
from dotenv import load_dotenv
|
||||
import aiohttp
|
||||
from typing import Dict, Any, List
|
||||
import tiktoken
|
||||
import httpx
|
||||
|
||||
# Load environment variables
|
||||
load_dotenv()
|
||||
|
||||
# Configure OpenAI client to point to OpenWebUI
|
||||
# Get environment variables
|
||||
DISCORD_TOKEN = os.getenv('DISCORD_TOKEN')
|
||||
LITELLM_API_KEY = os.getenv('LITELLM_API_KEY')
|
||||
LITELLM_API_BASE = os.getenv('LITELLM_API_BASE')
|
||||
MODEL_NAME = os.getenv('MODEL_NAME')
|
||||
SYSTEM_PROMPT_FILE = os.getenv('SYSTEM_PROMPT_FILE', './system_prompt.txt')
|
||||
MAX_HISTORY_TOKENS = int(os.getenv('MAX_HISTORY_TOKENS', '3000'))
|
||||
DEBUG_LOGGING = os.getenv('DEBUG_LOGGING', 'false').lower() == 'true'
|
||||
ENABLE_TOOLS = os.getenv('ENABLE_TOOLS', 'false').lower() == 'true'
|
||||
|
||||
def debug_log(message: str):
|
||||
"""Print debug message if DEBUG_LOGGING is enabled"""
|
||||
if DEBUG_LOGGING:
|
||||
print(f"[DEBUG] {message}")
|
||||
|
||||
# Load system prompt from file
|
||||
def load_system_prompt():
|
||||
"""Load system prompt from file, with fallback to default"""
|
||||
try:
|
||||
with open(SYSTEM_PROMPT_FILE, 'r', encoding='utf-8') as f:
|
||||
return f.read().strip()
|
||||
except FileNotFoundError:
|
||||
return "You are a helpful AI assistant integrated into Discord."
|
||||
|
||||
SYSTEM_PROMPT = load_system_prompt()
|
||||
|
||||
# Configure OpenAI client to point to LiteLLM
|
||||
client = OpenAI(
|
||||
api_key=os.getenv('OPENAI_API_KEY'),
|
||||
base_url=os.getenv('OPENWEBUI_API_BASE') # e.g., "http://localhost:8080/v1"
|
||||
api_key=LITELLM_API_KEY,
|
||||
base_url=LITELLM_API_BASE # e.g., "http://localhost:4000"
|
||||
)
|
||||
|
||||
# Initialize tokenizer for token counting
|
||||
try:
|
||||
encoding = tiktoken.encoding_for_model("gpt-4")
|
||||
except KeyError:
|
||||
encoding = tiktoken.get_encoding("cl100k_base")
|
||||
|
||||
# Initialize Discord bot
|
||||
intents = discord.Intents.default()
|
||||
intents.message_content = True
|
||||
bot = commands.Bot(command_prefix="!", intents=intents)
|
||||
intents.messages = True
|
||||
bot = commands.Bot(command_prefix='!', intents=intents)
|
||||
|
||||
# Add a dictionary to store conversation histories
|
||||
conversation_histories = {}
|
||||
DEFAULT_HISTORY_LIMIT = 50
|
||||
MAX_HISTORY_LIMIT = 200
|
||||
# Message history cache - stores recent conversations per channel
|
||||
channel_history: Dict[int, List[Dict[str, Any]]] = {}
|
||||
|
||||
async def get_ai_response(prompt, channel_id=None, include_history=False, history_limit=DEFAULT_HISTORY_LIMIT):
|
||||
def count_tokens(text: str) -> int:
|
||||
"""Count tokens in a text string"""
|
||||
try:
|
||||
messages = []
|
||||
return len(encoding.encode(text))
|
||||
except Exception:
|
||||
# Fallback: rough estimate (1 token ≈ 4 characters)
|
||||
return len(text) // 4
|
||||
|
||||
# If history is requested and exists for this channel, include it
|
||||
if include_history and channel_id in conversation_histories:
|
||||
messages = conversation_histories[channel_id]
|
||||
async def download_image(url: str) -> str | None:
|
||||
"""Download image and convert to base64 using async aiohttp"""
|
||||
try:
|
||||
async with aiohttp.ClientSession() as session:
|
||||
async with session.get(url, timeout=aiohttp.ClientTimeout(total=10)) as response:
|
||||
if response.status == 200:
|
||||
image_data = await response.read()
|
||||
base64_image = base64.b64encode(image_data).decode('utf-8')
|
||||
return base64_image
|
||||
except Exception as e:
|
||||
print(f"Error downloading image from {url}: {e}")
|
||||
return None
|
||||
|
||||
# Add the current prompt
|
||||
messages.append({"role": "user", "content": prompt})
|
||||
async def execute_mcp_tool(tool_name: str, arguments: dict) -> str:
|
||||
"""Execute an MCP tool via LiteLLM's /mcp/call_tool endpoint"""
|
||||
import json
|
||||
|
||||
response = client.chat.completions.create(
|
||||
model=os.getenv('MODEL_NAME') or "us.anthropic.claude-3-5-sonnet-20241022-v2:0",
|
||||
messages=messages,
|
||||
model=os.getenv('MODEL_NAME') or "us.anthropic.claude-3-5-sonnet-20241022-v2:0",
|
||||
messages=messages,
|
||||
temperature=0.7,
|
||||
max_tokens=500
|
||||
try:
|
||||
base_url = LITELLM_API_BASE.rstrip('/')
|
||||
headers = {
|
||||
"Authorization": f"Bearer {LITELLM_API_KEY}",
|
||||
"Content-Type": "application/json",
|
||||
"Accept": "application/json"
|
||||
}
|
||||
|
||||
debug_log(f"Executing MCP tool: {tool_name} with args: {arguments}")
|
||||
|
||||
async with httpx.AsyncClient(timeout=60.0) as http_client:
|
||||
response = await http_client.post(
|
||||
f"{base_url}/mcp/call_tool",
|
||||
headers=headers,
|
||||
json={
|
||||
"jsonrpc": "2.0",
|
||||
"method": "tools/call",
|
||||
"params": {
|
||||
"name": tool_name,
|
||||
"arguments": arguments
|
||||
},
|
||||
"id": 1
|
||||
}
|
||||
)
|
||||
|
||||
# Store the conversation history if channel_id is provided
|
||||
if channel_id:
|
||||
if channel_id not in conversation_histories:
|
||||
conversation_histories[channel_id] = []
|
||||
conversation_histories[channel_id].extend([
|
||||
{"role": "user", "content": prompt},
|
||||
{"role": "assistant", "content": response.choices[0].message.content}
|
||||
])
|
||||
debug_log(f"MCP call_tool response status: {response.status_code}")
|
||||
|
||||
# Limit history based on specified or default limit
|
||||
if len(conversation_histories[channel_id]) > history_limit * 2: # multiply by 2 because each exchange has 2 messages
|
||||
conversation_histories[channel_id] = conversation_histories[channel_id][-(history_limit * 2):]
|
||||
if response.status_code == 200:
|
||||
result = response.json()
|
||||
debug_log(f"MCP tool result: {str(result)[:200]}...")
|
||||
|
||||
# Handle JSON-RPC response format
|
||||
if isinstance(result, dict):
|
||||
# Check for JSON-RPC result
|
||||
if "result" in result:
|
||||
rpc_result = result["result"]
|
||||
# MCP tool results have a "content" array
|
||||
if isinstance(rpc_result, dict) and "content" in rpc_result:
|
||||
content = rpc_result["content"]
|
||||
if isinstance(content, list) and len(content) > 0:
|
||||
# Handle text content blocks
|
||||
first_content = content[0]
|
||||
if isinstance(first_content, dict) and "text" in first_content:
|
||||
return first_content["text"]
|
||||
return json.dumps(content)
|
||||
return json.dumps(content) if content else "Tool executed successfully"
|
||||
return json.dumps(rpc_result)
|
||||
# Fallback for non-RPC format
|
||||
if "content" in result:
|
||||
content = result["content"]
|
||||
if isinstance(content, list) and len(content) > 0:
|
||||
first_content = content[0]
|
||||
if isinstance(first_content, dict) and "text" in first_content:
|
||||
return first_content["text"]
|
||||
return json.dumps(content)
|
||||
return json.dumps(content) if content else "Tool executed successfully"
|
||||
return json.dumps(result)
|
||||
return str(result)
|
||||
else:
|
||||
error_text = response.text
|
||||
debug_log(f"MCP call_tool error: {response.status_code} - {error_text}")
|
||||
return f"Error executing tool: {response.status_code} - {error_text}"
|
||||
|
||||
return response.choices[0].message.content
|
||||
except Exception as e:
|
||||
print(f"Error getting AI response: {e}")
|
||||
return "Sorry, I encountered an error while processing your request."
|
||||
debug_log(f"Exception calling MCP tool: {e}")
|
||||
import traceback
|
||||
debug_log(f"Traceback: {traceback.format_exc()}")
|
||||
return f"Error executing tool: {str(e)}"
|
||||
|
||||
async def get_available_mcp_tools():
|
||||
"""Query LiteLLM for available MCP tools and convert to OpenAI function format"""
|
||||
try:
|
||||
base_url = LITELLM_API_BASE.rstrip('/')
|
||||
headers = {"Authorization": f"Bearer {LITELLM_API_KEY}"}
|
||||
|
||||
async with httpx.AsyncClient(timeout=30.0) as http_client:
|
||||
# Get available MCP tools
|
||||
tools_response = await http_client.get(
|
||||
f"{base_url}/v1/mcp/tools",
|
||||
headers=headers
|
||||
)
|
||||
|
||||
if tools_response.status_code == 200:
|
||||
tools_data = tools_response.json()
|
||||
mcp_tools = tools_data.get("tools", []) if isinstance(tools_data, dict) else tools_data
|
||||
debug_log(f"Found {len(mcp_tools)} MCP tools")
|
||||
|
||||
# Convert MCP tools to OpenAI function calling format
|
||||
openai_tools = []
|
||||
for tool in mcp_tools:
|
||||
if isinstance(tool, dict) and tool.get("name") and tool.get("description"):
|
||||
openai_tool = {
|
||||
"type": "function",
|
||||
"function": {
|
||||
"name": tool["name"],
|
||||
"description": tool.get("description", ""),
|
||||
"parameters": tool.get("inputSchema", {"type": "object", "properties": {}})
|
||||
}
|
||||
}
|
||||
openai_tools.append(openai_tool)
|
||||
|
||||
debug_log(f"Converted {len(openai_tools)} tools to OpenAI format")
|
||||
return openai_tools
|
||||
else:
|
||||
debug_log(f"MCP tools endpoint returned {tools_response.status_code}")
|
||||
|
||||
except Exception as e:
|
||||
debug_log(f"Error fetching MCP tools: {e}")
|
||||
|
||||
return []
|
||||
|
||||
async def get_chat_history(channel, bot_user_id: int, limit: int = 50) -> List[Dict[str, Any]]:
|
||||
"""
|
||||
Retrieve chat history and format as proper conversation messages.
|
||||
Only includes messages relevant to bot conversations.
|
||||
Returns list of message dicts with proper role attribution.
|
||||
Supports both regular channels and threads.
|
||||
"""
|
||||
messages = []
|
||||
total_tokens = 0
|
||||
|
||||
# Check if this is a thread
|
||||
is_thread = isinstance(channel, discord.Thread)
|
||||
|
||||
debug_log(f"Fetching history - is_thread: {is_thread}, channel: {channel.name if hasattr(channel, 'name') else 'DM'}")
|
||||
|
||||
# For threads, we want ALL messages in the thread (not just bot-related)
|
||||
# For channels, we only want bot-related messages
|
||||
|
||||
message_count = 0
|
||||
skipped_system = 0
|
||||
|
||||
# For threads, fetch the context including parent message if it exists
|
||||
if is_thread:
|
||||
try:
|
||||
# Get the starter message (first message in thread)
|
||||
if channel.starter_message:
|
||||
starter = channel.starter_message
|
||||
else:
|
||||
starter = await channel.fetch_message(channel.id)
|
||||
|
||||
# If the starter message is replying to another message, fetch that parent
|
||||
if starter and starter.reference and starter.reference.message_id:
|
||||
try:
|
||||
parent_message = await channel.parent.fetch_message(starter.reference.message_id)
|
||||
if parent_message and (parent_message.type == discord.MessageType.default or parent_message.type == discord.MessageType.reply):
|
||||
is_bot_parent = parent_message.author.id == bot_user_id
|
||||
role = "assistant" if is_bot_parent else "user"
|
||||
content = f"{parent_message.author.display_name}: {parent_message.content}" if not is_bot_parent else parent_message.content
|
||||
|
||||
# Remove bot mention if present
|
||||
if not is_bot_parent and bot_user_id:
|
||||
content = content.replace(f'<@{bot_user_id}>', '').strip()
|
||||
|
||||
msg = {"role": role, "content": content}
|
||||
msg_tokens = count_tokens(content)
|
||||
|
||||
if msg_tokens <= MAX_HISTORY_TOKENS:
|
||||
messages.append(msg)
|
||||
total_tokens += msg_tokens
|
||||
message_count += 1
|
||||
debug_log(f"Added parent message: role={role}, content_preview={content[:50]}...")
|
||||
except Exception as e:
|
||||
debug_log(f"Could not fetch parent message: {e}")
|
||||
|
||||
# Add the starter message itself
|
||||
if starter and (starter.type == discord.MessageType.default or starter.type == discord.MessageType.reply):
|
||||
is_bot_starter = starter.author.id == bot_user_id
|
||||
role = "assistant" if is_bot_starter else "user"
|
||||
content = f"{starter.author.display_name}: {starter.content}" if not is_bot_starter else starter.content
|
||||
|
||||
# Remove bot mention if present
|
||||
if not is_bot_starter and bot_user_id:
|
||||
content = content.replace(f'<@{bot_user_id}>', '').strip()
|
||||
|
||||
msg = {"role": role, "content": content}
|
||||
msg_tokens = count_tokens(content)
|
||||
|
||||
if total_tokens + msg_tokens <= MAX_HISTORY_TOKENS:
|
||||
messages.append(msg)
|
||||
total_tokens += msg_tokens
|
||||
message_count += 1
|
||||
debug_log(f"Added thread starter: role={role}, content_preview={content[:50]}...")
|
||||
except Exception as e:
|
||||
debug_log(f"Could not fetch thread messages: {e}")
|
||||
|
||||
# Fetch history from the channel/thread
|
||||
async for message in channel.history(limit=limit):
|
||||
message_count += 1
|
||||
|
||||
# Skip system messages (thread starters, pins, etc.)
|
||||
if message.type != discord.MessageType.default and message.type != discord.MessageType.reply:
|
||||
skipped_system += 1
|
||||
debug_log(f"Skipping system message type: {message.type}")
|
||||
continue
|
||||
|
||||
# Determine if we should include this message
|
||||
is_bot_message = message.author.id == bot_user_id
|
||||
is_bot_mentioned = any(mention.id == bot_user_id for mention in message.mentions)
|
||||
is_dm = isinstance(channel, discord.DMChannel)
|
||||
|
||||
# In threads: include ALL messages for full context
|
||||
# In regular channels: only include bot-related messages
|
||||
# In DMs: include all messages
|
||||
if is_thread or is_dm:
|
||||
should_include = True
|
||||
else:
|
||||
should_include = is_bot_message or is_bot_mentioned
|
||||
|
||||
if not should_include:
|
||||
continue
|
||||
|
||||
# Determine role
|
||||
role = "assistant" if is_bot_message else "user"
|
||||
|
||||
# Build content with author name in threads for multi-user context
|
||||
if is_thread and not is_bot_message:
|
||||
# Include username in threads for clarity
|
||||
content = f"{message.author.display_name}: {message.content}"
|
||||
else:
|
||||
content = message.content
|
||||
|
||||
# Remove bot mention from user messages
|
||||
if not is_bot_message and is_bot_mentioned:
|
||||
content = content.replace(f'<@{bot_user_id}>', '').strip()
|
||||
|
||||
# Note: We'll handle images separately in the main flow
|
||||
# For history, we just note that images were present
|
||||
if message.attachments:
|
||||
image_count = sum(1 for att in message.attachments
|
||||
if any(att.filename.lower().endswith(ext)
|
||||
for ext in ['.png', '.jpg', '.jpeg', '.gif', '.webp']))
|
||||
if image_count > 0:
|
||||
content += f" [attached {image_count} image(s)]"
|
||||
|
||||
# Add to messages with token counting
|
||||
msg = {"role": role, "content": content}
|
||||
msg_tokens = count_tokens(content)
|
||||
|
||||
# Check if adding this message would exceed token limit
|
||||
if total_tokens + msg_tokens > MAX_HISTORY_TOKENS:
|
||||
break
|
||||
|
||||
messages.append(msg)
|
||||
total_tokens += msg_tokens
|
||||
debug_log(f"Added message: role={role}, content_preview={content[:50]}...")
|
||||
|
||||
# Reverse to get chronological order (oldest first)
|
||||
debug_log(f"Processed {message_count} messages, skipped {skipped_system} system messages")
|
||||
debug_log(f"Total messages collected: {len(messages)}, total tokens: {total_tokens}")
|
||||
return list(reversed(messages))
|
||||
|
||||
|
||||
async def get_ai_response(history_messages: List[Dict[str, Any]], user_message: str, image_urls: List[str] = None) -> str:
|
||||
"""
|
||||
Get AI response using LiteLLM chat.completions with manual MCP tool execution.
|
||||
|
||||
Uses manual tool execution loop since Responses API doesn't work with Bedrock + MCP.
|
||||
|
||||
Args:
|
||||
history_messages: List of previous conversation messages with roles
|
||||
user_message: Current user message
|
||||
image_urls: Optional list of image URLs to include
|
||||
|
||||
Returns:
|
||||
AI response string
|
||||
"""
|
||||
import json
|
||||
|
||||
# Build messages array
|
||||
messages = [{"role": "system", "content": SYSTEM_PROMPT}]
|
||||
messages.extend(history_messages)
|
||||
|
||||
# Build current user message
|
||||
if image_urls:
|
||||
content_parts = [{"type": "text", "text": user_message}]
|
||||
for url in image_urls:
|
||||
base64_image = await download_image(url)
|
||||
if base64_image:
|
||||
content_parts.append({
|
||||
"type": "image_url",
|
||||
"image_url": {"url": f"data:image/jpeg;base64,{base64_image}"}
|
||||
})
|
||||
messages.append({"role": "user", "content": content_parts})
|
||||
else:
|
||||
messages.append({"role": "user", "content": user_message})
|
||||
|
||||
try:
|
||||
# Build request parameters
|
||||
request_params = {
|
||||
"model": MODEL_NAME,
|
||||
"messages": messages,
|
||||
"temperature": 0.7,
|
||||
}
|
||||
|
||||
# Add MCP tools if enabled
|
||||
tools = []
|
||||
if ENABLE_TOOLS:
|
||||
debug_log("Tools enabled - fetching MCP tools")
|
||||
tools = await get_available_mcp_tools()
|
||||
|
||||
if tools:
|
||||
request_params["tools"] = tools
|
||||
request_params["tool_choice"] = "auto"
|
||||
debug_log(f"Added {len(tools)} tools to request")
|
||||
|
||||
debug_log(f"Calling chat.completions with {len(tools)} tools")
|
||||
response = client.chat.completions.create(**request_params)
|
||||
|
||||
# Handle tool calls if present
|
||||
response_message = response.choices[0].message
|
||||
tool_calls = getattr(response_message, 'tool_calls', None)
|
||||
|
||||
# Tool execution loop (max 5 iterations to prevent infinite loops)
|
||||
max_iterations = 5
|
||||
iteration = 0
|
||||
|
||||
while tool_calls and len(tool_calls) > 0 and iteration < max_iterations:
|
||||
iteration += 1
|
||||
debug_log(f"Tool call iteration {iteration}: Model requested {len(tool_calls)} tool calls")
|
||||
|
||||
# Add assistant's response with tool calls to messages
|
||||
messages.append({
|
||||
"role": "assistant",
|
||||
"content": response_message.content,
|
||||
"tool_calls": [
|
||||
{
|
||||
"id": tc.id,
|
||||
"type": "function",
|
||||
"function": {
|
||||
"name": tc.function.name,
|
||||
"arguments": tc.function.arguments
|
||||
}
|
||||
}
|
||||
for tc in tool_calls
|
||||
]
|
||||
})
|
||||
|
||||
# Execute each tool call via MCP
|
||||
for tool_call in tool_calls:
|
||||
function_name = tool_call.function.name
|
||||
function_args_str = tool_call.function.arguments
|
||||
|
||||
debug_log(f"Executing tool: {function_name}")
|
||||
|
||||
# Parse arguments
|
||||
try:
|
||||
args_dict = json.loads(function_args_str) if isinstance(function_args_str, str) else function_args_str
|
||||
except json.JSONDecodeError:
|
||||
args_dict = {}
|
||||
debug_log(f"Failed to parse tool arguments: {function_args_str}")
|
||||
|
||||
# Execute the tool via MCP
|
||||
tool_result = await execute_mcp_tool(function_name, args_dict)
|
||||
|
||||
# Add tool result to messages
|
||||
messages.append({
|
||||
"role": "tool",
|
||||
"tool_call_id": tool_call.id,
|
||||
"content": tool_result
|
||||
})
|
||||
|
||||
# Get next response from model
|
||||
debug_log("Getting model response after tool execution")
|
||||
request_params["messages"] = messages
|
||||
response = client.chat.completions.create(**request_params)
|
||||
|
||||
response_message = response.choices[0].message
|
||||
tool_calls = getattr(response_message, 'tool_calls', None)
|
||||
|
||||
if iteration >= max_iterations:
|
||||
debug_log(f"Warning: Reached max tool iterations ({max_iterations})")
|
||||
|
||||
final_content = response.choices[0].message.content
|
||||
debug_log(f"Final response: {final_content[:100] if final_content else 'None'}...")
|
||||
return final_content or "I received a response but it was empty. Please try again."
|
||||
|
||||
except Exception as e:
|
||||
error_msg = f"Error calling LiteLLM API: {str(e)}"
|
||||
print(error_msg)
|
||||
debug_log(f"Exception details: {e}")
|
||||
import traceback
|
||||
debug_log(f"Traceback: {traceback.format_exc()}")
|
||||
return error_msg
|
||||
|
||||
@bot.event
|
||||
async def on_message(message):
|
||||
# Ignore messages from the bot itself
|
||||
if message.author == bot.user:
|
||||
return
|
||||
|
||||
if bot.user in message.mentions:
|
||||
prompt = message.content.replace(f'<@{bot.user.id}>', '').strip()
|
||||
|
||||
if not prompt:
|
||||
await message.channel.send("Hello! How can I help you?")
|
||||
# Ignore system messages (thread starter, pins, etc.)
|
||||
if message.type != discord.MessageType.default and message.type != discord.MessageType.reply:
|
||||
return
|
||||
|
||||
# Check if the message includes a request for history
|
||||
include_history = False
|
||||
history_limit = DEFAULT_HISTORY_LIMIT
|
||||
should_respond = False
|
||||
|
||||
# Check for "with history", "with X history", or "X lines of chat" patterns
|
||||
import re
|
||||
if "with history" in prompt.lower() or re.search(r"\d+\s*lines of chat", prompt.lower()):
|
||||
include_history = True
|
||||
# Check for specific history limit
|
||||
if match := re.search(r"with (\d+) history", prompt.lower()):
|
||||
requested_limit = int(match.group(1))
|
||||
history_limit = min(requested_limit, MAX_HISTORY_LIMIT)
|
||||
prompt = re.sub(r"with \d+ history", "", prompt, flags=re.IGNORECASE)
|
||||
elif match := re.search(r"(\d+)\s*lines of chat", prompt.lower()):
|
||||
requested_limit = int(match.group(1))
|
||||
history_limit = min(requested_limit, MAX_HISTORY_LIMIT)
|
||||
prompt = re.sub(r"\d+\s*lines of chat", "", prompt, flags=re.IGNORECASE)
|
||||
else:
|
||||
prompt = prompt.lower().replace("with history", "")
|
||||
# Check if bot was mentioned
|
||||
if bot.user in message.mentions:
|
||||
should_respond = True
|
||||
|
||||
prompt = prompt.strip()
|
||||
# Check if message is a DM
|
||||
if isinstance(message.channel, discord.DMChannel):
|
||||
should_respond = True
|
||||
|
||||
# Check if message is in a thread
|
||||
if isinstance(message.channel, discord.Thread):
|
||||
# Check if thread was started from a bot message
|
||||
try:
|
||||
starter = message.channel.starter_message
|
||||
if not starter:
|
||||
starter = await message.channel.fetch_message(message.channel.id)
|
||||
|
||||
# If thread was started from bot's message, auto-respond
|
||||
if starter and starter.author.id == bot.user.id:
|
||||
should_respond = True
|
||||
debug_log("Thread started by bot - auto-responding")
|
||||
# If thread started from user message, only respond if mentioned
|
||||
elif bot.user in message.mentions:
|
||||
should_respond = True
|
||||
debug_log("Thread started by user - responding due to mention")
|
||||
except Exception as e:
|
||||
debug_log(f"Could not determine thread starter: {e}")
|
||||
# Default: only respond if mentioned
|
||||
if bot.user in message.mentions:
|
||||
should_respond = True
|
||||
|
||||
if should_respond:
|
||||
async with message.channel.typing():
|
||||
response = await get_ai_response(
|
||||
prompt,
|
||||
channel_id=str(message.channel.id),
|
||||
include_history=include_history,
|
||||
history_limit=history_limit
|
||||
)
|
||||
# Get chat history with proper conversation format
|
||||
history_messages = await get_chat_history(message.channel, bot.user.id)
|
||||
|
||||
# Remove bot mention from the message
|
||||
user_message = message.content.replace(f'<@{bot.user.id}>', '').strip()
|
||||
|
||||
# Collect image URLs from the message
|
||||
image_urls = []
|
||||
for attachment in message.attachments:
|
||||
if any(attachment.filename.lower().endswith(ext) for ext in ['.png', '.jpg', '.jpeg', '.gif', '.webp']):
|
||||
image_urls.append(attachment.url)
|
||||
|
||||
# Get AI response with proper conversation history
|
||||
response = await get_ai_response(history_messages, user_message, image_urls if image_urls else None)
|
||||
|
||||
# Send response (split if too long for Discord's 2000 char limit)
|
||||
if len(response) > 2000:
|
||||
# Split into chunks
|
||||
chunks = [response[i:i+2000] for i in range(0, len(response), 2000)]
|
||||
for chunk in chunks:
|
||||
await message.channel.send(chunk)
|
||||
await message.reply(chunk)
|
||||
else:
|
||||
await message.channel.send(response)
|
||||
await message.reply(response)
|
||||
|
||||
await bot.process_commands(message)
|
||||
|
||||
@bot.command(name='clearhistory')
|
||||
async def clear_history(ctx):
|
||||
channel_id = str(ctx.channel.id)
|
||||
if channel_id in conversation_histories:
|
||||
conversation_histories[channel_id] = []
|
||||
await ctx.send("Conversation history has been cleared.")
|
||||
else:
|
||||
await ctx.send("No conversation history exists for this channel.")
|
||||
@bot.event
|
||||
async def on_ready():
|
||||
print(f'{bot.user} has connected to Discord!')
|
||||
|
||||
|
||||
def main():
|
||||
# Get the Discord token from environment variables
|
||||
discord_token = os.getenv('DISCORD_TOKEN')
|
||||
if not discord_token:
|
||||
raise ValueError("Discord token not found in environment variables")
|
||||
if not all([DISCORD_TOKEN, LITELLM_API_KEY, LITELLM_API_BASE, MODEL_NAME]):
|
||||
print("Error: Missing required environment variables")
|
||||
print(f"DISCORD_TOKEN: {'✓' if DISCORD_TOKEN else '✗'}")
|
||||
print(f"LITELLM_API_KEY: {'✓' if LITELLM_API_KEY else '✗'}")
|
||||
print(f"LITELLM_API_BASE: {'✓' if LITELLM_API_BASE else '✗'}")
|
||||
print(f"MODEL_NAME: {'✓' if MODEL_NAME else '✗'}")
|
||||
return
|
||||
|
||||
# Run the bot
|
||||
bot.run(discord_token)
|
||||
print(f"System Prompt loaded from: {SYSTEM_PROMPT_FILE}")
|
||||
print(f"Max history tokens: {MAX_HISTORY_TOKENS}")
|
||||
bot.run(DISCORD_TOKEN)
|
||||
|
||||
if __name__ == "__main__":
|
||||
main()
|
||||
@@ -0,0 +1,6 @@
|
||||
discord.py>=2.0.0
|
||||
openai>=2.0.0
|
||||
python-dotenv>=1.0.0
|
||||
aiohttp>=3.8.0
|
||||
tiktoken>=0.5.0
|
||||
httpx>=0.25.0
|
||||
@@ -0,0 +1,18 @@
|
||||
You are a helpful AI assistant integrated into Discord. Users will interact with you by mentioning you, sending direct messages, or chatting in threads.
|
||||
|
||||
Key behaviors:
|
||||
- Be concise and friendly in your responses
|
||||
- Use Discord markdown formatting when helpful (code blocks, bold, italics, etc.)
|
||||
- When users attach images, analyze them and provide relevant insights
|
||||
- Keep track of conversation context from the chat history provided
|
||||
- In threads, you have access to the full conversation context - reference previous messages when relevant
|
||||
- In regular channels, you only see messages where you were mentioned
|
||||
- If you're unsure about something, acknowledge it honestly
|
||||
- Provide helpful and accurate information
|
||||
|
||||
Tool capabilities:
|
||||
- You have access to various tools and integrations (like GitHub, file systems, etc.) that can help you accomplish tasks
|
||||
- When appropriate, use available tools to provide more accurate and helpful responses
|
||||
- If you use a tool, explain what you're doing so users understand the process
|
||||
|
||||
You are an AI assistant, not a human. Be transparent about your capabilities and limitations.
|
||||
@@ -0,0 +1,24 @@
|
||||
# Discord Bot Token - Get from https://discord.com/developers/applications
|
||||
DISCORD_TOKEN=your_discord_bot_token
|
||||
|
||||
# LiteLLM API Configuration
|
||||
LITELLM_API_KEY=sk-1234
|
||||
LITELLM_API_BASE=http://localhost:4000
|
||||
|
||||
# Model name (any model supported by your LiteLLM proxy)
|
||||
MODEL_NAME=gpt-4-turbo-preview
|
||||
|
||||
# System Prompt Configuration (optional)
|
||||
SYSTEM_PROMPT_FILE=./system_prompt.txt
|
||||
|
||||
# Maximum tokens to use for conversation history (optional, default: 3000)
|
||||
MAX_HISTORY_TOKENS=3000
|
||||
|
||||
# Enable debug logging (optional, default: false)
|
||||
# Set to 'true' to see detailed logs for troubleshooting
|
||||
DEBUG_LOGGING=false
|
||||
|
||||
# Enable MCP tools integration (optional, default: false)
|
||||
# Set to 'true' to allow the bot to use tools configured in your LiteLLM proxy
|
||||
# Tools are auto-executed without user confirmation
|
||||
ENABLE_TOOLS=false
|
||||
@@ -0,0 +1,139 @@
|
||||
import os
|
||||
import discord
|
||||
from discord.ext import commands
|
||||
from openai import OpenAI
|
||||
import base64
|
||||
import requests
|
||||
from io import BytesIO
|
||||
from collections import deque
|
||||
from dotenv import load_dotenv
|
||||
|
||||
# Load environment variables
|
||||
load_dotenv()
|
||||
|
||||
# Get environment variables
|
||||
DISCORD_TOKEN = os.getenv('DISCORD_TOKEN')
|
||||
OPENAI_API_KEY = os.getenv('OPENAI_API_KEY')
|
||||
OPENWEBUI_API_BASE = os.getenv('OPENWEBUI_API_BASE')
|
||||
MODEL_NAME = os.getenv('MODEL_NAME')
|
||||
|
||||
# Configure OpenAI client to point to OpenWebUI
|
||||
client = OpenAI(
|
||||
api_key=os.getenv('OPENAI_API_KEY'),
|
||||
base_url=os.getenv('OPENWEBUI_API_BASE') # e.g., "http://localhost:8080/v1"
|
||||
)
|
||||
|
||||
# Configure OpenAI
|
||||
# TODO: The 'openai.api_base' option isn't read in the client API. You will need to pass it when you instantiate the client, e.g. 'OpenAI(base_url=OPENWEBUI_API_BASE)'
|
||||
# openai.api_base = OPENWEBUI_API_BASE
|
||||
|
||||
# Initialize Discord bot
|
||||
intents = discord.Intents.default()
|
||||
intents.message_content = True
|
||||
intents.messages = True
|
||||
bot = commands.Bot(command_prefix='!', intents=intents)
|
||||
|
||||
# Message history cache
|
||||
channel_history = {}
|
||||
|
||||
async def download_image(url):
|
||||
response = requests.get(url)
|
||||
if response.status_code == 200:
|
||||
image_data = BytesIO(response.content)
|
||||
base64_image = base64.b64encode(image_data.read()).decode('utf-8')
|
||||
return base64_image
|
||||
return None
|
||||
|
||||
async def get_chat_history(channel, limit=100):
|
||||
messages = []
|
||||
async for message in channel.history(limit=limit):
|
||||
content = f"{message.author.name}: {message.content}"
|
||||
|
||||
# Handle attachments (images)
|
||||
for attachment in message.attachments:
|
||||
if any(attachment.filename.lower().endswith(ext) for ext in ['.png', '.jpg', '.jpeg', '.gif', '.webp']):
|
||||
content += f" [Image: {attachment.url}]"
|
||||
|
||||
messages.append(content)
|
||||
return "\n".join(reversed(messages))
|
||||
|
||||
async def get_ai_response(context, user_message, image_urls=None):
|
||||
messages = [{"role": "user", "content": []}]
|
||||
|
||||
# Add text content
|
||||
text_content = f"##CONTEXT##\n{context}\n##ENDCONTEXT##\n\n{user_message}"
|
||||
messages[0]["content"].append({"type": "text", "text": text_content})
|
||||
|
||||
# Add image content if present
|
||||
if image_urls:
|
||||
for url in image_urls:
|
||||
base64_image = await download_image(url)
|
||||
if base64_image:
|
||||
messages[0]["content"].append({
|
||||
"type": "image_url",
|
||||
"image_url": {
|
||||
"url": f"data:image/jpeg;base64,{base64_image}"
|
||||
}
|
||||
})
|
||||
|
||||
try:
|
||||
response = client.chat.completions.create(
|
||||
model=MODEL_NAME,
|
||||
messages=messages
|
||||
)
|
||||
return response.choices[0].message.content
|
||||
except Exception as e:
|
||||
return f"Error: {str(e)}"
|
||||
|
||||
@bot.event
|
||||
async def on_message(message):
|
||||
# Ignore messages from the bot itself
|
||||
if message.author == bot.user:
|
||||
return
|
||||
|
||||
should_respond = False
|
||||
|
||||
# Check if bot was mentioned
|
||||
if bot.user in message.mentions:
|
||||
should_respond = True
|
||||
|
||||
# Check if message is a DM
|
||||
if isinstance(message.channel, discord.DMChannel):
|
||||
should_respond = True
|
||||
|
||||
if should_respond:
|
||||
async with message.channel.typing():
|
||||
# Get chat history
|
||||
history = await get_chat_history(message.channel)
|
||||
|
||||
# Remove bot mention from the message
|
||||
user_message = message.content.replace(f'<@{bot.user.id}>', '').strip()
|
||||
|
||||
# Collect image URLs from the message
|
||||
image_urls = []
|
||||
for attachment in message.attachments:
|
||||
if any(attachment.filename.lower().endswith(ext) for ext in ['.png', '.jpg', '.jpeg', '.gif', '.webp']):
|
||||
image_urls.append(attachment.url)
|
||||
|
||||
# Get AI response
|
||||
response = await get_ai_response(history, user_message, image_urls)
|
||||
|
||||
# Send response
|
||||
await message.reply(response)
|
||||
|
||||
await bot.process_commands(message)
|
||||
|
||||
@bot.event
|
||||
async def on_ready():
|
||||
print(f'{bot.user} has connected to Discord!')
|
||||
|
||||
|
||||
def main():
|
||||
if not all([DISCORD_TOKEN, OPENAI_API_KEY, OPENWEBUI_API_BASE, MODEL_NAME]):
|
||||
print("Error: Missing required environment variables")
|
||||
return
|
||||
|
||||
bot.run(DISCORD_TOKEN)
|
||||
|
||||
if __name__ == "__main__":
|
||||
main()
|
||||
@@ -0,0 +1,5 @@
|
||||
discord.py>=2.0.0
|
||||
openai>=1.0.0
|
||||
python-dotenv>=1.0.0
|
||||
aiohttp>=3.8.0
|
||||
tiktoken>=0.5.0
|
||||
Reference in New Issue
Block a user