Files
free-ide-proxy/AGENTS.md
T
Renato 05a52d29e2 feat: multi-backend proxy with auto rate-limit jailing
Refactor from monolithic proxy to multi-backend architecture
(Backend interface: OpenCode Zen + Kilo gateway).

- Add auto rate-limit jailing: models returning HTTP 429 are
  hidden from /v1/models immediately and re-probed every 30 min
- Backend interface supports model aliases, custom headers, and
  per-backend routing
- Full Anthropic Messages API support with OpenAI format conversion
- 9 free models across two backends with Claude name aliases
2026-06-26 19:08:01 +02:00

88 lines
6.3 KiB
Markdown

# Free IDE Proxy
OpenAI/Anthropic-compatible proxy that aggregates the **free** models from multiple opencode-based editors (currently OpenCode Zen and Kilo Code) under a single endpoint. Each upstream lives behind a `Backend` implementation and the proxy routes by model name.
## Build & Run
```bash
go build -o free-ide-proxy .
./free-ide-proxy # no auth
./free-ide-proxy -api-key "secret" # with auth
```
Auth reads `OPENCODE_PROXY_KEY` env var; `-api-key` flag takes precedence.
Defaults: `127.0.0.1:6446`. Use `-host 0.0.0.0` for external access.
## Architecture
```
Client → internal/server/ (HTTP, auth, routing) → internal/proxy/ (model resolve, backend router, session mgmt, forwarding, Anthropic↔OpenAI conversion) → Backend → upstream
```
- **Pure stdlib Go 1.22** — no router, no deps beyond stdlib
- **Entrypoint:** `main.go` wires `proxy.NewProxy(key)` into `server.NewServer(p)`
- **Multi-backend:** `internal/proxy/backend.go` defines the `Backend` interface. The `Proxy` holds an ordered list and routes each request to the first backend that recognises the model. Unknown models fall back to the default backend (opencode) and are forwarded as-is.
- **Backends:** `opencode.go` (OpenCode Zen, `opencode.ai/zen/v1`, `Bearer public`) and `kilo.go` (Kilo gateway, `api.kilo.ai/api/gateway`, no auth). Add a new file implementing `Backend` and register it in `NewProxy` (`models.go`) to support another upstream.
- **Auth middleware** in `server.go:88` checks `Authorization: Bearer` or `x-api-key` header
- **Model aliases** live in each backend (`opencode.go` and `kilo.go`) — short names, company names, and Claude model names → full IDs
- **Session rotation:** session IDs rotate every 30 min per user key (`proxy.SessionID()`)
- **Auto rate-limit jailing:** models returning HTTP 429 are automatically hidden from `/v1/models` and re-probed every 30 min until they recover. See `models.go` (`modelHealth`, `healthLoop`, `pingModel`, `markRateLimited`, `markActive`, `isRateLimited`, `Models()` filtering).
- **Pass-through:** unknown model names are forwarded as-is to the default backend, not rejected
## Endpoints
| Method | Path | Description |
|--------|------|-------------|
| `GET` | `/health` | Health check (returns hardcoded `"1.0.0"`) |
| `GET` | `/v1/models` | List available free models (auth required if key set) |
| `GET` | `/v1/models/{id}` | Single model detail |
| `POST` | `/v1/chat/completions` | OpenAI Chat Completions (streaming + non-streaming) |
| `POST` | `/v1/messages` | Anthropic Messages API |
| `POST` | `/v1/v1/messages` | Same as above — workaround for Claude Code double-path bug |
## Anthropic Messages API
Fully implemented in `internal/proxy/anthropic.go`. Converts Anthropic → OpenAI format upstream, then converts responses back to Anthropic SSE (streaming) or JSON (non-streaming).
**Claude Code quirk:** Claude Code appends `/v1/messages` to `ANTHROPIC_BASE_URL` automatically. Set `ANTHROPIC_BASE_URL` to `http://127.0.0.1:6446` (without `/v1` suffix) to avoid double-path. The proxy also handles `/v1/v1/messages` as a fallback (`server.go:62`).
**Claude model name aliases** are defined in `models.go:77-88` — Claude Code validates model names client-side, so you must set `ANTHROPIC_MODEL` to a valid Claude name like `claude-sonnet-4-6` (maps to `deepseek-v4-flash-free`).
## Available Models (from `opencode.go` + `kilo.go`)
### OpenCode Zen backend (`opencode.go` → `opencode.ai/zen/v1`, `Bearer public`)
| Model ID | Aliases | Context | Max Output |
|----------|---------|---------|------------|
| `deepseek-v4-flash-free` | `deepseek`, `deepseek-v4`, `ds` | 1M | 384K |
| `big-pickle` | `pickle` | 200K | 32K |
| `mimo-v2.5-free` | `mimo`, `mimo-v2.5`, `xiaomi` | 1M | 32K |
| `north-mini-code-free` | `north`, `north-mini`, `cohere` | 256K | 64K |
| `nemotron-3-ultra-free` | `nemotron`, `nemotron-3`, `nvidia` | 1M | 16K |
### Kilo gateway backend (`kilo.go` → `api.kilo.ai/api/gateway`, no auth)
| Model ID | Aliases | Context | Max Output |
|----------|---------|---------|------------|
| `stepfun/step-3.7-flash:free` | `stepfun`, `stepfun-free` | 256K | 32K |
| `poolside/laguna-m.1:free` | `poolside`, `poolside-free`, `laguna` | 256K | 32K |
| `nvidia/nemotron-3-ultra-550b-a55b:free` | — | 1M | 16K |
| `openrouter/free` | `openrouter` | 256K | 32K |
**Source of truth is the code (`opencode.go`, `kilo.go`), not `README.md`.** The README may list models that differ from the code (e.g. `minimax-m2.5-free`, `kimi-k2.5-free`, `gpt-5-nano`, `nemotron-3-super-free`, `qwen3.6-plus-free` are NOT in the code). The code controls what works.
## Quirks & Gotchas
- **No tests exist.** Any change is untested unless you add them.
- **No CI/CD, no Makefile, no lint config.** All manual.
- **In-memory only.** Session and health state dies on restart. The health map auto-rebuilds via initial scan.
- **`deploy/start-proxy.bat`** and **`deploy/CREDENCIALES.md`** contain a hardcoded API key — do not commit them.
- **`opencode-proxy.exe~`** in root is a backup artifact; ignore.
- Upstream Zen API expects `x-opencode-request`, `x-opencode-session`, `x-opencode-client`, `x-opencode-project` headers (set in `opencode.go` `OpenCodeBackend.Headers`). Kilo needs no auth — its gateway returns HTTP 200 for free models with no `Authorization` header.
- **Server logs request bodies to stdout** (truncated to 500 chars in `server.go:265-271`). Verbose by design.
- **Health endpoint** returns hardcoded version `"1.0.0"` in `server.go:126`. Update both `main.go:12` and `server.go:126` when bumping.
- **Upstream timeout:** `proxy.go` sets a 10-minute HTTP client timeout — adjust if Zen models are slow on first call.
- **Unknown models are forwarded as-is** via the default backend (opencode) — useful when new models appear upstream before code is updated.
- **Reasoning models** (DeepSeek, Nemotron, StepFun) need `max_tokens` ≥ 500 — first tokens go to reasoning, not visible content. With `max_tokens:30` they return empty content; that is model behaviour, not a proxy bug.
- **Rate limit jailing:** `proxy.go:164` and `anthropic.go:713` check for 429 and call `markRateLimited` immediately. `healthLoop()` in `models.go` does a full scan on startup, then re-checks caged models every 30 min. `Models()` filters out rate-limited models — they reappear once recovered.