LLMs

The ai() layer owns provider clients and model traffic. It keeps Ax programs focused on signatures while one provider surface handles chat, streaming, embeddings, media, usage normalization, thinking controls, routing, balancing, tracing, and provider-specific behavior.

C++

auto client = axllm::OpenAICompatibleClient({
  {"api_key", api_key},
  {"model", "gpt-4.1-mini"},
});

Provider Setup

Create provider clients near the application boundary, keep keys in environment variables, and pass the client into forward(), agents, flows, or optimizers.

Generated Package Provider Path

The C++ package exposes the AxIR-supported provider surface. Public examples use OpenAI-compatible clients, while internal fixtures cover provider normalization without credentials.

IllustrativeGenerated-package equivalent. Prefer checked-in package examples for copy/paste runnable code.

C++

auto client = axllm::OpenAICompatibleClient({
  {"api_key", api_key},
  {"model", "gpt-4.1-mini"},
});

Use the generated package examples for exact provider API runs, stream mapping, Responses audio mapping, and realtime event folding for this language.

Model Catalog

Use the model catalog before runtime when a UI or router needs model choices, costs, and capabilities. It can filter for text, code, embedding, and audio models.

IllustrativeGenerated-package equivalent. Prefer checked-in package examples for copy/paste runnable code.

C++

// TypeScript exposes the bundled model catalog helper.
// Generated packages publish capability metadata in axir-capabilities.json.

flowchart LR
  A[Model catalog] --> B[Capability filter]
  B --> C[Text]
  B --> D[Embeddings]
  B --> E[Audio]
  C --> F[Route or select model]
  D --> F
  E --> F

Routing And Balancing

Routing has two distinct jobs. The multi-service router combines provider model lists and dispatches the model key the caller already chose; it does not select a model. AxBalancer handles equivalent services behind shared model aliases, preserving the Ax request shape while applying capability filters and provider failover.

Every supported language can opt into adaptive AxBalancer routing. It learns transient provider failure rate and successful latency, then weighs the probability of a failure or deadline miss against estimated request cost. This is operational provider selection, not semantic prompt-to-model selection, so every model behind an alias must be an acceptable substitute.

Embeddings

Embeddings live on the same provider client surface. Use them for retrieval indexes, memory search, context lookup, and similarity workflows while keeping embedding model selection separate from generation model selection.

IllustrativeGenerated-package equivalent. Prefer checked-in package examples for copy/paste runnable code.

C++

// Implement embedding calls through the generated AxAI client surface when present.
// Use package conformance coverage to confirm current support for this language.

Audio, Realtime, And Responses

Ax maps batch transcription, batch speech, conversational audio, OpenAI Responses audio, and realtime event folding where supported. Direct ax(...) programs can pass media to compatible models; agents usually transcribe audio before planner/executor/responder stages.

IllustrativeGenerated-package equivalent. Prefer checked-in package examples for copy/paste runnable code.

C++

// Realtime audio over WebSocket — build with -DAXLLM_ENABLE_REALTIME=ON
axllm::OpenAIResponsesClient client(axllm::object({{"model", "gpt-realtime-2"}, {"api_key", std::getenv("OPENAI_APIKEY")}}));
axllm::Value request = axllm::object({{"model", "gpt-realtime-2"}, {"chat_prompt", axllm::array({axllm::object({{"role", "user"}, {"content", "Say hello."}})})}, {"audio", axllm::object({{"output", axllm::object({{"voice", "alloy"}})}})}});
axllm::Value result = client.realtime_chat(request, nullptr);  // one merged turn: transcript + base64 PCM audio
// Realtime models also route transparently through chat(); chat() accepts input_audio parts; transcribe()/speak() do batch STT/TTS.

Thinking And Context Caching

Thinking controls expose provider-specific reasoning budgets through one Ax option. Context caching marks stable prompt regions so providers with prefix caching can reuse expensive context.

IllustrativeGenerated-package equivalent. Prefer checked-in package examples for copy/paste runnable code.

C++

// Thinking budgets are provider-specific runtime options.
// Trace usage and provider metadata before relying on a budget in production.

flowchart TB
  A[Stable context field] --> B[Cache breakpoint]
  C[User query] --> D[Generation]
  B --> D
  E[thinkingTokenBudget] --> D
  D --> F[Usage + trace]

Production Notes

Keep provider keys outside source code.
Prefer model aliases like fast, smart, or cheap when app callers should not know provider model IDs.
Trace request latency, retries, token usage, cost, route choice, media mode, and model key.
Keep public provider examples separate from internal conformance fixtures.
Use OpenAI-compatible clients for generated-language package examples when that is the supported provider path.

See ai() LLM models and ai() API.