Home/Features

Features

Everything Iron LLM Gateway provides out of the box. All features are self-hosted — your data stays on your infrastructure.

Unified Proxy

Core

Single endpoint that routes to OpenAI, Anthropic, GitHub Copilot, or your own models.

  • Route via /v1/llm/* — model name is auto-detected from the request body
  • Model prefix routing: anthropic/claude-3-5-sonnet, openai/gpt-4o
  • Round-robin load balancing across multiple API keys per provider
  • Dedicated endpoints: /v1/openai/*, /v1/anthropic/*, /v1/copilot/*
  • Supports streaming (SSE) and non-streaming responses
  • Token limits auto-adjusted per model (max_tokens / max_completion_tokens)
# Unified endpoint – auto-detects provider from model name
client = OpenAI(
    base_url="https://llm-gateway-api.ironcode.cloud/v1/llm",
    api_key="rpk_your_proxy_key",
)

response = client.chat.completions.create(
    model="gpt-4o",          # → routes to OpenAI
    # model="claude-3-5-sonnet-20241022",  # → routes to Anthropic
    messages=[{"role": "user", "content": "Hello"}],
)

PII Detection & Filtering

Security

Scan every request for sensitive data before it reaches the LLM provider.

  • Built-in patterns: OpenAI/Anthropic/AWS/GitHub API keys, JWT tokens, private keys, database URLs, passwords
  • Two actions: block (reject request) or replace (redact and forward)
  • Custom rules via GRL (Go Rules Language) — create regex or function-based detectors
  • Rule salience (priority ordering) — higher priority rules evaluated first
  • IP allowlist/blocklist with CIDR support
  • All violations logged with count and type per request
# Example: request containing an API key gets blocked
POST /v1/llm/chat/completions
{
  "messages": [{
    "role": "user",
    "content": "My key is sk-proj-abc123..."
  }]
}

# Response when blocked:
# 400 Bad Request
# { "error": "Request blocked: API key detected in content" }

# Or with replace action — key is redacted before forwarding:
# "content": "My key is [REDACTED]"

Encrypted Key Management

Security

Store provider API keys encrypted at rest. Issue proxy keys to users so real keys are never exposed.

  • Provider keys (OpenAI, Anthropic, etc.) stored with AES-256-GCM encryption
  • Master key derived via Argon2 — set ENCRYPTION_KEY env var (min 32 chars)
  • Proxy keys issued to users: format rpk_[48 random chars]
  • Proxy keys stored as SHA-256 hash — plain key shown only once at creation
  • Dashboard shows safe preview: rpk_XXXXXXXX...
  • Last-used timestamp tracked per proxy key
# Set encryption key before starting the server
ENCRYPTION_KEY=your-32-char-minimum-secret-key

# Users authenticate with proxy keys
curl https://llm-gateway-api.ironcode.cloud/v1/openai/chat/completions \
  -H "Authorization: Bearer rpk_your_proxy_key" \
  -d '{"model": "gpt-4o", "messages": [...]}'

OAuth Authentication

Auth

Sign in with GitHub or Google. Role-based access separates admins from regular users.

  • GitHub OAuth: Authorization Code flow via /api/auth/github/oauth/callback
  • Google OAuth: Authorization Code flow via /api/auth/google/callback
  • JWT tokens issued on login — configurable expiry via JWT_EXPIRY_HOURS
  • Roles: admin (full access) and user (own keys & logs only)
  • Avatar and display name fetched from provider profile
  • Configure via environment variables: GITHUB_CLIENT_ID, GOOGLE_CLIENT_ID, etc.
# Environment variables for OAuth setup
JWT_SECRET=your-jwt-secret
JWT_EXPIRY_HOURS=24

GITHUB_CLIENT_ID=your-github-app-client-id
GITHUB_CLIENT_SECRET=your-github-app-secret
GITHUB_REDIRECT_URI=https://your-domain/auth/callback/github

GOOGLE_CLIENT_ID=your-google-client-id
GOOGLE_CLIENT_SECRET=your-google-secret
GOOGLE_REDIRECT_URI=https://your-domain/auth/callback/google

Full Audit Logs

Observability

Every request is logged with full context — provider, model, tokens, latency, violations, and client IP.

  • Fields: provider, endpoint, model, HTTP status, latency (ms), client IP, user agent
  • Token metrics: input tokens, output tokens, total tokens per request
  • Violation tracking: blocked flag, replaced flag, violation count
  • User attribution: proxy key name and user ID linked to each request
  • Configurable retention: set RETENTION_DAYS env var
  • Browse and filter logs from the dashboard /logs page
# Query logs via API (JWT required)
GET /api/logs?limit=50&provider=openai&blocked=true

# Response
{
  "data": [{
    "id": "...",
    "timestamp": "2025-01-01T10:00:00Z",
    "provider": "openai",
    "model": "gpt-4o",
    "status_code": 200,
    "latency_ms": 342,
    "input_tokens": 150,
    "output_tokens": 89,
    "blocked": false,
    "client_ip": "1.2.3.4"
  }]
}

Usage & Cost Tracking

Observability

65+ built-in model prices. Track token usage per user and per model with configurable pricing overrides.

  • 65+ built-in model prices for OpenAI, Anthropic, and Google models
  • Input and output pricing per 1K tokens, configurable per model
  • Exact-match and prefix-match for pricing overrides (e.g. gpt-4o or gpt-4)
  • Per-user aggregated usage visible in /usage dashboard
  • Admin pricing config via /admin/pricing — add, edit, delete model prices
  • Trigger manual cost recalculation via POST /api/usage/recalculate
# Configure custom pricing (admin only)
POST /api/admin/pricing
{
  "model": "my-fine-tuned-gpt4",
  "input_per_1k": 0.005,
  "output_per_1k": 0.015
}

# View usage for current user
GET /api/usage

# Response
{
  "data": {
    "total_tokens": 125000,
    "total_cost_usd": 1.23,
    "by_model": { "gpt-4o": { ... } }
  }
}

GRL Rules Engine

Advanced

Hot-reloadable rules engine. Write regex patterns, IP filters, or custom detection functions.

  • Rule types: regex pattern match, IP allowlist/blocklist (CIDR), custom detection functions
  • Built-in functions: @func:detect_api_keys, @func:detect_passwords, @func:detect_jwt_tokens, etc.
  • Actions per rule: block (reject) or replace (redact matched content)
  • Salience (priority) ordering — higher salience rules run first
  • Severity levels: low, medium, high, critical
  • Rules stored in DB and hot-reloaded atomically — no restart needed
// Example GRL rule (via dashboard /rules or POST /api/rules)
{
  "name": "Block internal API keys",
  "description": "Block requests containing internal service tokens",
  "pattern": "internal-svc-[a-zA-Z0-9]{32}",
  "action": "block",
  "severity": "high",
  "salience": 100,
  "enabled": true
}

Ready to get started?