Configuration
Settings live in ~/.ex5/config.toml. Override the whole data root with EX5_HOME. Project-level overrides go in <repo>/.ex5/local.toml. Precedence: CLI flag / env > config file > built-in default.
You normally do not write this by hand — `ex5 login --key ex5-live-…` generates it from the model catalogue.
# ~/.ex5/config.toml default_model = "ex5/ex5-security-27b" [providers.ex5] type = "openai" # OpenAI Chat Completions base_url = "https://ex5.ai/v1" api_key = "ex5-live-…" [models."ex5/ex5-security-27b"] provider = "ex5" model = "ex5-security-27b" max_context_size = 131072 max_output_size = 8192 # ÇIKIŞ tavanı — bağlamdan ayrı, aşağıya bak capabilities = ["tool_use"] display_name = "EX5-Security 27B" [thinking] enabled = false
Output limit
max_output_size is separate from max_context_size on purpose. Our models are served by vLLM, where the prompt and the reply share one window (max_model_len). If a client sets max_tokens to the full context size, the request is rejected and every agent step fails. Leave it as login wrote it unless you know the server accepts more.
Thinking
Reasoning models produce a chain of thought before answering. When the model declares supports_reasoning, ex5 renders that stream in a separate collapsed block instead of printing it as the answer. Toggle it with [thinking].enabled; leaving it off keeps turns cheaper and the tool loop faster.
Other providers
ex5 ships pointed at ex5.ai but is not locked to it. Add any provider as a sibling table and any model under [models.<alias>]. Supported type values: openai, openai_responses, anthropic, google-genai, vertexai, kimi. Anything that speaks OpenAI Chat Completions works — OpenRouter, Together, a local vLLM, Ollama — by pointing base_url at it.
[providers.openai] type = "openai" base_url = "https://api.openai.com/v1" api_key = "sk-…" [models."openai/gpt-x"] provider = "openai" model = "gpt-x" max_context_size = 128000 max_output_size = 16384 capabilities = ["tool_use"]
