Features
Everything Iron LLM Gateway provides out of the box. All features are self-hosted — your data stays on your infrastructure.
Unified Proxy
CoreSingle endpoint that routes to OpenAI, Anthropic, GitHub Copilot, or your own models.
- Route via
/v1/llm/*— model name is auto-detected from the request body - Model prefix routing:
anthropic/claude-3-5-sonnet,openai/gpt-4o - Round-robin load balancing across multiple API keys per provider
- Dedicated endpoints:
/v1/openai/*,/v1/anthropic/*,/v1/copilot/* - Supports streaming (SSE) and non-streaming responses
- Token limits auto-adjusted per model (max_tokens / max_completion_tokens)
# Unified endpoint – auto-detects provider from model name
client = OpenAI(
base_url="https://llm-gateway-api.ironcode.cloud/v1/llm",
api_key="rpk_your_proxy_key",
)
response = client.chat.completions.create(
model="gpt-4o", # → routes to OpenAI
# model="claude-3-5-sonnet-20241022", # → routes to Anthropic
messages=[{"role": "user", "content": "Hello"}],
)PII Detection & Filtering
SecurityScan every request for sensitive data before it reaches the LLM provider.
- Built-in patterns: OpenAI/Anthropic/AWS/GitHub API keys, JWT tokens, private keys, database URLs, passwords
- Two actions:
block(reject request) orreplace(redact and forward) - Custom rules via GRL (Go Rules Language) — create regex or function-based detectors
- Rule salience (priority ordering) — higher priority rules evaluated first
- IP allowlist/blocklist with CIDR support
- All violations logged with count and type per request
# Example: request containing an API key gets blocked
POST /v1/llm/chat/completions
{
"messages": [{
"role": "user",
"content": "My key is sk-proj-abc123..."
}]
}
# Response when blocked:
# 400 Bad Request
# { "error": "Request blocked: API key detected in content" }
# Or with replace action — key is redacted before forwarding:
# "content": "My key is [REDACTED]"Encrypted Key Management
SecurityStore provider API keys encrypted at rest. Issue proxy keys to users so real keys are never exposed.
- Provider keys (OpenAI, Anthropic, etc.) stored with AES-256-GCM encryption
- Master key derived via Argon2 — set
ENCRYPTION_KEYenv var (min 32 chars) - Proxy keys issued to users: format
rpk_[48 random chars] - Proxy keys stored as SHA-256 hash — plain key shown only once at creation
- Dashboard shows safe preview:
rpk_XXXXXXXX... - Last-used timestamp tracked per proxy key
# Set encryption key before starting the server
ENCRYPTION_KEY=your-32-char-minimum-secret-key
# Users authenticate with proxy keys
curl https://llm-gateway-api.ironcode.cloud/v1/openai/chat/completions \
-H "Authorization: Bearer rpk_your_proxy_key" \
-d '{"model": "gpt-4o", "messages": [...]}'OAuth Authentication
AuthSign in with GitHub or Google. Role-based access separates admins from regular users.
- GitHub OAuth: Authorization Code flow via
/api/auth/github/oauth/callback - Google OAuth: Authorization Code flow via
/api/auth/google/callback - JWT tokens issued on login — configurable expiry via
JWT_EXPIRY_HOURS - Roles:
admin(full access) anduser(own keys & logs only) - Avatar and display name fetched from provider profile
- Configure via environment variables:
GITHUB_CLIENT_ID,GOOGLE_CLIENT_ID, etc.
# Environment variables for OAuth setup JWT_SECRET=your-jwt-secret JWT_EXPIRY_HOURS=24 GITHUB_CLIENT_ID=your-github-app-client-id GITHUB_CLIENT_SECRET=your-github-app-secret GITHUB_REDIRECT_URI=https://your-domain/auth/callback/github GOOGLE_CLIENT_ID=your-google-client-id GOOGLE_CLIENT_SECRET=your-google-secret GOOGLE_REDIRECT_URI=https://your-domain/auth/callback/google
Full Audit Logs
ObservabilityEvery request is logged with full context — provider, model, tokens, latency, violations, and client IP.
- Fields: provider, endpoint, model, HTTP status, latency (ms), client IP, user agent
- Token metrics: input tokens, output tokens, total tokens per request
- Violation tracking: blocked flag, replaced flag, violation count
- User attribution: proxy key name and user ID linked to each request
- Configurable retention: set
RETENTION_DAYSenv var - Browse and filter logs from the dashboard
/logspage
# Query logs via API (JWT required)
GET /api/logs?limit=50&provider=openai&blocked=true
# Response
{
"data": [{
"id": "...",
"timestamp": "2025-01-01T10:00:00Z",
"provider": "openai",
"model": "gpt-4o",
"status_code": 200,
"latency_ms": 342,
"input_tokens": 150,
"output_tokens": 89,
"blocked": false,
"client_ip": "1.2.3.4"
}]
}Usage & Cost Tracking
Observability65+ built-in model prices. Track token usage per user and per model with configurable pricing overrides.
- 65+ built-in model prices for OpenAI, Anthropic, and Google models
- Input and output pricing per 1K tokens, configurable per model
- Exact-match and prefix-match for pricing overrides (e.g.
gpt-4oorgpt-4) - Per-user aggregated usage visible in
/usagedashboard - Admin pricing config via
/admin/pricing— add, edit, delete model prices - Trigger manual cost recalculation via
POST /api/usage/recalculate
# Configure custom pricing (admin only)
POST /api/admin/pricing
{
"model": "my-fine-tuned-gpt4",
"input_per_1k": 0.005,
"output_per_1k": 0.015
}
# View usage for current user
GET /api/usage
# Response
{
"data": {
"total_tokens": 125000,
"total_cost_usd": 1.23,
"by_model": { "gpt-4o": { ... } }
}
}GRL Rules Engine
AdvancedHot-reloadable rules engine. Write regex patterns, IP filters, or custom detection functions.
- Rule types: regex pattern match, IP allowlist/blocklist (CIDR), custom detection functions
- Built-in functions:
@func:detect_api_keys,@func:detect_passwords,@func:detect_jwt_tokens, etc. - Actions per rule:
block(reject) orreplace(redact matched content) - Salience (priority) ordering — higher salience rules run first
- Severity levels: low, medium, high, critical
- Rules stored in DB and hot-reloaded atomically — no restart needed
// Example GRL rule (via dashboard /rules or POST /api/rules)
{
"name": "Block internal API keys",
"description": "Block requests containing internal service tokens",
"pattern": "internal-svc-[a-zA-Z0-9]{32}",
"action": "block",
"severity": "high",
"salience": 100,
"enabled": true
}