Skip to content

Configuration ​

Template and runtime configuration ​

Edit the checked-in config.yaml template, then run make up to render it and restart the proxy. The rendered file is ~/.cli-proxy-api/proxy/config.yaml, mounted at /data/config.yaml inside the container.

The client key is a 1Password reference using OP_VAULT and OP_ITEM_NAME. The management password placeholder is replaced with a bcrypt hash. Do not put a real key, token, password, or hash of a real password into the public template.

Panel edits to the runtime configuration are overwritten on the next make up. Port persistent changes back into the template, replacing any sensitive values with secret references.

Environment settings ​

VariableDefaultPurpose
OP_ACCOUNTmy.1password.com1Password sign-in account
OP_VAULTPrivateVault containing the credential item
OP_ITEM_NAMEcli-proxy-apiCredential item name
HARA_ENV_FILE$LOCAL_DIR/credentials.envOptional private file with default credential item names
LOCAL_DIR~/.cli-proxy-apiRuntime state outside the checkout
LOCAL_NAMEharaportless LAN alias
TS_HOSTNAMEcliproxyTailscale container device name
TS_AUTHKEYunsetOptional private enrollment key
URLhttps://hara.localTarget for Claude and operations Make commands
HARA_KEY_MAX_AGE_DAYS30Keychain cache lifetime
BACKUP_DIR$LOCAL_DIR/backupsOptional private client backup destination

LOCAL_DIR and BACKUP_DIR should be absolute paths outside the repository. The Codex profile has its own base URL; URL does not change it.

Routing and limits ​

The template enables round-robin routing, one-hour session affinity, and affinity for subagents. The quota service raises account priority when unused weekly capacity will expire soon. Panel priority edits are overwritten by that service on its next hourly run.

Cooldown state is persisted next to provider credentials, so a restart does not relearn every limit. Claude and Codex use model-level cooling, allowing another model on the same account to remain eligible. When all accounts are cooling down, the template limits waits between retry rounds to five seconds.

Codex speed setting ​

The template supplies service_tier: priority for gpt-* and codex-* models using the Codex protocol. Fast mode can consume more usage. Remove that payload rule if you prefer the upstream standard tier. First-token bootstrap buffering stays off; a slow or overloaded account is not silently hidden by delaying the first token.

Logs and plugins ​

File logging is enabled with a 512 MB total cap and up to ten request error logs. Debug and request logging are disabled. Logs still deserve private treatment because account identifiers and operational details may appear.

The plugin runtime is enabled. Only install providers or plugin binaries you trust. Plugins in the container's plugins directory are lost when the container is recreated. Consult the upstream configuration for persistence options before adding plugins.

Upgrade the proxy ​

The proxy image is pinned in local/compose.yaml. Read upstream release notes, back up private runtime state securely, update the tag, and run:

bash
make up
make status
make ops models
make ops smoke
make ops ws-smoke

Check the configuration against the release you deploy. The upstream default branch's option reference can differ from the pinned image.

hara · built on CLIProxyAPI