Claude Code on Muapi
Environment variables
| Variable | Value |
|---|---|
ANTHROPIC_BASE_URL | https://api.muapi.ai/anthropic |
ANTHROPIC_AUTH_TOKEN | Your Muapi API key (Claude Code sends it as a Bearer token) |
ANTHROPIC_MODEL | Optional. An id from the models endpoint, e.g. claude-sonnet-5-5 |
Don't include /v1/messages in the base URL; Claude Code adds it.
Quick start (shell)
export ANTHROPIC_BASE_URL=https://api.muapi.ai/anthropic
export ANTHROPIC_AUTH_TOKEN="$MUAPI_API_KEY"
claude
Settings file
~/.claude/settings.json (or .claude/settings.local.json). Values here are literal, so prefer shell exports if you don't want the key in a file. Never commit credentials to a project-level .claude/settings.json.
{
"env": {
"ANTHROPIC_BASE_URL": "https://api.muapi.ai/anthropic",
"ANTHROPIC_AUTH_TOKEN": "<your Muapi API key>"
}
}
Models
curl -s -H "Authorization: Bearer $MUAPI_API_KEY" https://api.muapi.ai/anthropic/v1/models
Unrestricted (abliterated) models
The models endpoint also lists Muapi's tool-capable unrestricted models (for example qwen-3-8-27b-obliterated, mimo-v2-6-flash-abliterated, glm-5-3-abliterated, abliterated-model-large-v2). Claude Code can run on them: Muapi translates between Claude Code's protocol and the model's chat API for you.
export ANTHROPIC_BASE_URL=https://api.muapi.ai/anthropic
export ANTHROPIC_AUTH_TOKEN="$MUAPI_API_KEY"
export ANTHROPIC_DEFAULT_HAIKU_MODEL=qwen-3-8-27b-obliterated
export ANTHROPIC_SMALL_FAST_MODEL=qwen-3-8-27b-obliterated
claude --model qwen-3-8-27b-obliterated
Set the two HAIKU/SMALL_FAST variables to the same model id. Claude Code sends background requests (titles, summaries) to a small "Haiku" model name, and without them those requests fail with a 404.
What to expect:
- Quality varies by model. These are smaller open models; on multi-step coding tasks they can skip a step (for example forget to create one of several files). Check results, and try a larger model for harder work.
- Cost grows with every turn. There is no prompt caching, so each turn is billed for the whole conversation so far. A typical turn costs roughly one cent to a few cents on the smaller models and more on the large ones; the Models endpoint and your dashboard show exact usage.
- Output per response is capped at 16,384 tokens (some models lower), and the model's context window is smaller than Claude's, so very long sessions need
/compactsooner. - Only models with tool calling are offered here; Claude Code cannot run without tools.
- Image input works only on vision-capable models.
Verify
Run /status in Claude Code to confirm the base URL. Use claude --debug to see requests.
Notes
- If Claude Code asks you to sign in to Anthropic, quit the terminal completely and reopen it from a shell where the variables are exported.
- A
401means the key or header is wrong; a402means your balance can't cover the request's worst case (the full prompt plus the maximum output is held first and the unused part is returned), so keep a buffer above the cost of one turn.