ai() LLM Models
Use ai() to create provider clients and keep model traffic behind one Ax request shape.
client := ax.NewOpenAICompatibleClient(map[string]ax.Value{
"api_key": apiKey,
"model": "gpt-4.1-mini",
})
program := ax.NewAx("question:string -> answer:string", nil)
out, err := program.Forward(ctx, client, map[string]ax.Value{"question": "What is Ax?"}, nil)What It Does
ai() selects a provider implementation from configuration and returns a client that Ax programs can call. The client handles chat, streaming, embeddings, media where supported, usage normalization, provider options, model keys, routing hooks, tracing, and runtime defaults.
flowchart LR A["Model key or alias"] --> B["Model catalog"] B --> C["Capability filter"] C --> D["Provider client"] D --> E["Request mapping"] E --> F["Provider API"] F --> G["Response normalization"] G --> H["Usage + trace"]
Core Call Shape
Create the client once near the application boundary, then pass it into forward(), streamingForward(), agents, flows, or optimizers.
client = ai(provider options)
result = program.forward(client, inputs)Common Patterns
- Use a provider
nameand environment-backed API key. - Set a default model in provider config when the app has one obvious model.
- Define model aliases when callers should choose
fast,smart, orcheapinstead of provider model IDs. - Use OpenAI-compatible
apiURLfor compatible providers. - Use model catalog helpers before runtime when the UI needs provider/model selectors.
- Use routers or balancers when provider fallback is part of the product.
Vertex routing and OpenAI prompt caching
Gemini and Anthropic Vertex clients accept a project, location, and optional endpoint. Ax resolves global, US/EU multi-region, and regional hosts; an explicit base URL takes precedence. Generated packages take a caller-supplied bearer access token and leave ADC refresh to the host application.
GPT-5.6 OpenAI Chat requests can opt into stable explicit prompt-cache
breakpoints. Give AxGen a stable promptCacheKey plus contextCache; those
forward options reach the provider in every language. Cache reads and writes
are normalized separately for usage and catalog-backed cost estimates.
Adaptive balancing
AxBalancer keeps its existing ordered failover behavior by default. Set strategy.type to adaptive to rank equivalent providers per chat request using learned reliability, successful latency, a deadline, and estimated cost. Configure badOutcomeCost in the same currency or unit as the route cost estimate.
Use the native stats-store option for authoritative decision state. The built-in in-memory store can be shared by balancers in one process; multi-process applications can implement AxBalancerStatsStore with an atomic Redis or database update. The routing-event hook is best-effort telemetry, not routing state. Stable route keys are required with a shared store, and namespace plus slice keep unrelated traffic from learning from each other.
Adaptive balancing does not inspect prompt meaning or decide which model is best for a task. The application defines acceptable substitutes through shared logical aliases.
Provider clients
Generated Package Provider Path
The Go package exposes the AxIR-supported provider surface. Public examples use OpenAI-compatible clients, while internal fixtures cover provider normalization without credentials.
client := ax.NewOpenAICompatibleClient(map[string]ax.Value{
"api_key": apiKey,
"model": "gpt-4.1-mini",
})
program := ax.NewAx("question:string -> answer:string", nil)
out, err := program.Forward(ctx, client, map[string]ax.Value{"question": "What is Ax?"}, nil)Vertex Gemini
client := ax.NewGoogleGeminiClient(map[string]ax.Value{
"api_key": required("GOOGLE_VERTEX_ACCESS_TOKEN"),
"project_id": required("GOOGLE_PROJECT_ID"),
"region": required("GOOGLE_REGION"),
"model": model,
})Use the generated package examples for exact provider API runs, prompt-cached AxGen calls, stream mapping, Responses audio mapping, and realtime event folding for this language.
Embeddings and audio
// Implement embedding calls through the generated AxAI client surface when present.
// Use package conformance coverage to confirm current support for this language.// Realtime audio over WebSocket — needs Go 1.23+
client := ax.NewOpenAIResponsesClient(map[string]ax.Value{"model": "gpt-realtime-2", "api_key": os.Getenv("OPENAI_APIKEY")})
request := map[string]ax.Value{"model": "gpt-realtime-2", "chat_prompt": ax.Array(ax.Object("role", "user", "content", "Say hello.")), "audio": ax.Object("output", ax.Object("voice", "alloy"))}
final, _ := client.RealtimeChat(context.Background(), request, nil, nil) // one merged turn: transcript + base64 PCM audio
// Realtime models also route transparently through Chat(); Chat() accepts input_audio parts; Transcribe()/Speak() do batch STT/TTS.Practical Notes
- Prefer provider factories over direct provider classes in new code.
- Use model catalog and provider-scoring helpers when choosing between providers.
- Use a multi-service router to dispatch caller-selected model keys; use a balancer for fallback or adaptive operational routing across equivalent services.
- Keep public provider examples separate from internal conformance fixtures.
- Trace provider requests, token usage, estimated cost, and routing decisions in production.
See ai() API.