Files
free-ide-proxy/AGENTS.md
T
Renato 05a52d29e2 feat: multi-backend proxy with auto rate-limit jailing
Refactor from monolithic proxy to multi-backend architecture
(Backend interface: OpenCode Zen + Kilo gateway).

- Add auto rate-limit jailing: models returning HTTP 429 are
  hidden from /v1/models immediately and re-probed every 30 min
- Backend interface supports model aliases, custom headers, and
  per-backend routing
- Full Anthropic Messages API support with OpenAI format conversion
- 9 free models across two backends with Claude name aliases
2026-06-26 19:08:01 +02:00

6.3 KiB

Free IDE Proxy

OpenAI/Anthropic-compatible proxy that aggregates the free models from multiple opencode-based editors (currently OpenCode Zen and Kilo Code) under a single endpoint. Each upstream lives behind a Backend implementation and the proxy routes by model name.

Build & Run

go build -o free-ide-proxy .
./free-ide-proxy                     # no auth
./free-ide-proxy -api-key "secret"   # with auth

Auth reads OPENCODE_PROXY_KEY env var; -api-key flag takes precedence. Defaults: 127.0.0.1:6446. Use -host 0.0.0.0 for external access.

Architecture

Client → internal/server/ (HTTP, auth, routing) → internal/proxy/ (model resolve, backend router, session mgmt, forwarding, Anthropic↔OpenAI conversion) → Backend → upstream
  • Pure stdlib Go 1.22 — no router, no deps beyond stdlib
  • Entrypoint: main.go wires proxy.NewProxy(key) into server.NewServer(p)
  • Multi-backend: internal/proxy/backend.go defines the Backend interface. The Proxy holds an ordered list and routes each request to the first backend that recognises the model. Unknown models fall back to the default backend (opencode) and are forwarded as-is.
  • Backends: opencode.go (OpenCode Zen, opencode.ai/zen/v1, Bearer public) and kilo.go (Kilo gateway, api.kilo.ai/api/gateway, no auth). Add a new file implementing Backend and register it in NewProxy (models.go) to support another upstream.
  • Auth middleware in server.go:88 checks Authorization: Bearer or x-api-key header
  • Model aliases live in each backend (opencode.go and kilo.go) — short names, company names, and Claude model names → full IDs
  • Session rotation: session IDs rotate every 30 min per user key (proxy.SessionID())
  • Auto rate-limit jailing: models returning HTTP 429 are automatically hidden from /v1/models and re-probed every 30 min until they recover. See models.go (modelHealth, healthLoop, pingModel, markRateLimited, markActive, isRateLimited, Models() filtering).
  • Pass-through: unknown model names are forwarded as-is to the default backend, not rejected

Endpoints

Method Path Description
GET /health Health check (returns hardcoded "1.0.0")
GET /v1/models List available free models (auth required if key set)
GET /v1/models/{id} Single model detail
POST /v1/chat/completions OpenAI Chat Completions (streaming + non-streaming)
POST /v1/messages Anthropic Messages API
POST /v1/v1/messages Same as above — workaround for Claude Code double-path bug

Anthropic Messages API

Fully implemented in internal/proxy/anthropic.go. Converts Anthropic → OpenAI format upstream, then converts responses back to Anthropic SSE (streaming) or JSON (non-streaming).

Claude Code quirk: Claude Code appends /v1/messages to ANTHROPIC_BASE_URL automatically. Set ANTHROPIC_BASE_URL to http://127.0.0.1:6446 (without /v1 suffix) to avoid double-path. The proxy also handles /v1/v1/messages as a fallback (server.go:62).

Claude model name aliases are defined in models.go:77-88 — Claude Code validates model names client-side, so you must set ANTHROPIC_MODEL to a valid Claude name like claude-sonnet-4-6 (maps to deepseek-v4-flash-free).

OpenCode Zen backend (opencode.goopencode.ai/zen/v1, Bearer public)

Model ID Aliases Context Max Output
deepseek-v4-flash-free deepseek, deepseek-v4, ds 1M 384K
big-pickle pickle 200K 32K
mimo-v2.5-free mimo, mimo-v2.5, xiaomi 1M 32K
north-mini-code-free north, north-mini, cohere 256K 64K
nemotron-3-ultra-free nemotron, nemotron-3, nvidia 1M 16K

Kilo gateway backend (kilo.goapi.kilo.ai/api/gateway, no auth)

Model ID Aliases Context Max Output
stepfun/step-3.7-flash:free stepfun, stepfun-free 256K 32K
poolside/laguna-m.1:free poolside, poolside-free, laguna 256K 32K
nvidia/nemotron-3-ultra-550b-a55b:free 1M 16K
openrouter/free openrouter 256K 32K

Source of truth is the code (opencode.go, kilo.go), not README.md. The README may list models that differ from the code (e.g. minimax-m2.5-free, kimi-k2.5-free, gpt-5-nano, nemotron-3-super-free, qwen3.6-plus-free are NOT in the code). The code controls what works.

Quirks & Gotchas

  • No tests exist. Any change is untested unless you add them.
  • No CI/CD, no Makefile, no lint config. All manual.
  • In-memory only. Session and health state dies on restart. The health map auto-rebuilds via initial scan.
  • deploy/start-proxy.bat and deploy/CREDENCIALES.md contain a hardcoded API key — do not commit them.
  • opencode-proxy.exe~ in root is a backup artifact; ignore.
  • Upstream Zen API expects x-opencode-request, x-opencode-session, x-opencode-client, x-opencode-project headers (set in opencode.go OpenCodeBackend.Headers). Kilo needs no auth — its gateway returns HTTP 200 for free models with no Authorization header.
  • Server logs request bodies to stdout (truncated to 500 chars in server.go:265-271). Verbose by design.
  • Health endpoint returns hardcoded version "1.0.0" in server.go:126. Update both main.go:12 and server.go:126 when bumping.
  • Upstream timeout: proxy.go sets a 10-minute HTTP client timeout — adjust if Zen models are slow on first call.
  • Unknown models are forwarded as-is via the default backend (opencode) — useful when new models appear upstream before code is updated.
  • Reasoning models (DeepSeek, Nemotron, StepFun) need max_tokens ≥ 500 — first tokens go to reasoning, not visible content. With max_tokens:30 they return empty content; that is model behaviour, not a proxy bug.
  • Rate limit jailing: proxy.go:164 and anthropic.go:713 check for 429 and call markRateLimited immediately. healthLoop() in models.go does a full scan on startup, then re-checks caged models every 30 min. Models() filters out rate-limited models — they reappear once recovered.