Skip to content

What this module does not do

Documentation usually describes what software does. This page is the other half: what chat is not for, which combinations do not work, and where a value you set is ignored. Each item is a decision or a known gap, not an oversight waiting to be reported.

One client talks to one provider

A ChatClient is built for exactly one provider and one model. There is no way to ask one client to use Claude for prose and Gemini for images, or to route different calls to different providers by content. Construct one client per provider and choose between them in your own code.

A fallback composite is not an exception to this. It uses one provider at a time and moves to the next only when the active one fails; it is a resilience mechanism, not a router.

There is no image generation, embedding or audio synthesis

The module sends text and attachments to a chat completion endpoint and returns text. Images, PDFs and (on Gemini) audio and video are input only, described in Providers.

There is no API here for generating an image, producing embeddings, transcribing audio or synthesising speech, and no plan to add one behind ChatClient. Those are different endpoints with different shapes, and squeezing them into a chat interface would make the interface worse without making them usable.

Configuration is read once, at construction

chat.New reads Config and hands the resolved values to the provider. Nothing re-reads them. A host that hot-reloads its own configuration does not change a live client's provider, model, timeout or credentials. Build a new client and swap it, which also avoids mutating conversation state mid-session.

Nothing bounds a retry sequence as a whole

Config.RequestTimeout bounds a single request and Config.MaxRetries bounds the number of attempts, but there is no setting for "give up after N seconds total". A provider returning long Retry-After hints, retried twice, can occupy the caller for considerably longer than the request timeout suggests. Cancel the context if you need a wall-clock bound.

The module does not pace calls or track quota

There is no rate limiter, token-budget accounting or spend cap. Usage() and Config.UsageObserver report what was consumed after the fact; they do not refuse a call. When a provider rate-limits, its own retry runs and, if that is exhausted, the failure surfaces as chat.ErrThrottled. Pacing and budgets are the caller's, and this is a deliberate boundary: the module has no view of a caller's other traffic against the same key.

Structured output is not checked at construction

Config.ResponseSchema has no marker interface and no construction-time check in the core. This is unlike Config.Stateless, Config.Effort and Config.CacheTTL, each of which is type-asserted and fails construction when the provider cannot carry it.

So a custom provider that ignores the field is not detected, and chat.New will hand back a client that returns unstructured text.

It is also not true that every shipped provider honours it. Claude, OpenAI, Gemini and claude-local do, the last through the CLI's --json-schema. On Bedrock it depends on the model: no Amazon-owned model supports structured output at all. chat-bedrock therefore does the check the core does not, and refuses at construction when the selected model is known not to support it. That is a provider compensating for a gap here rather than the gap being closed, and CapTools has the same shape.

If you register your own provider and accept a schema, honour it or return an error from your factory. Nothing upstream will catch it for you.

Capability reporting cannot see past a deployment name

Azure addresses a deployment, whose name is chosen by whoever created it, so my-gpt4-prod resolves to no entry in any table. Guessing the model from a substring would be a confident claim about somebody else's deployment.

From chat-openai v0.12.1 the answer is an honest SupportUnknown for everything model-dependent, which proceeds untouched and so refuses nothing. Before that it reported SupportYes for tools, structured output, multimodal, caching and effort: adapter-level claims presented where an endpoint-level answer belongs, and CapMultimodal said yes while the core refused every attachment, which cannot both be true.

The remaining limitation is real but smaller: you cannot be told in advance that a deployment lacks something. Declaring your endpoint's capabilities is what closes that, and is tracked in issue 17.

The information does exist: Azure returns the real model in its responses. It arrives too late to be useful, because capability resolution is a pure function of the model string and runs at construction, before any request has been sent. Making it vary over a client's life would break the contract every other provider keeps.

Unknown proceeds untouched, so nothing is refused that would have worked. What you lose is the ability to be told in advance that a deployment cannot do something. openai-compatible has the same limitation for the same structural reason, and issue 17 is where a way for a consumer to declare what their backend supports would be settled.

claude-local is missing more than it has

ProviderClaudeLocal shells out to the claude CLI, which constrains it sharply. It does not support:

Capability Behaviour on claude-local
Tool calling SetTools returns an error; MCP-based tools are a future release
Streaming does not implement chat.StreamingChatClient; use Chat
Persistence does not implement chat.PersistentChatClient; there is no internal state to snapshot
Prompt caching does not implement chat.CachingChatClient; setting Config.CacheTTL is a construction error
Sampling no --temperature or --top-p flag exists at any CLI version, so setting either is a construction error
Media accepts no attachments; any chat.Media is rejected with chat.ErrMediaUnsupported
Token accounting reports usage only when the binary does; otherwise Usage{Known: false} (do not cost a run from it)

Four Config fields are also ignored rather than refused, because the CLI has nowhere to put them: MaxTokens, RequestTimeout, HTTPClient and MaxRetries. The subprocess is bounded only by the context you pass, so a claude-local call has no timeout of its own, so pass a context.WithTimeout if you need one.

What it does support is Config.Effort (all five levels map to the CLI's --effort), Config.Stateless, Config.SystemPrompt, Config.Model and Config.ResponseSchema.

A fallback composite carries less than you might expect

Four constraints, all consequences of composing clients through the public interface rather than reaching inside them:

  • It is never safe for concurrent use, even when every underlying client is stateless and individually safe. The composite advances its active provider mid-call. Give each worker its own composite.
  • Replay is lossy. Only user turns are recorded and replayed into a fallback provider. Assistant turns and the tool-call/tool-result interleaving cannot be injected through ChatClient, so a conversation that did heavy tool use resumes with less context than it had. WithStrictToolContext turns that into a fast failure instead.
  • Streaming fails over only before the first visible event. Once an EventTextDelta or EventToolCallStart has reached your callback it cannot be un-emitted, so any later error is terminal.
  • Model, token, credentials and base URL do not carry across members. On the config-driven paths each provider self-resolves those, so a single global model name is not applied to the whole chain. Pin a model per provider by passing explicit Config values.

Capability tables lag new models

chat.CapabilitiesFor answers from a table each provider module generates by querying and probing the vendor. A model released after that table was generated reports SupportUnknown, which the module treats as "proceed and let the provider answer", so a stale table costs nothing, but it cannot warn you either. Only a confident SupportNo changes behaviour.

Two of the five providers can never do better than Unknown: claude-local is a CLI with no models endpoint, and openai-compatible points at an arbitrary server.

Prompt caching is a hint, never a guarantee

AddCached asks a provider to cache a block. Every provider declines below a per-model minimum (1,024 tokens on Gemini and OpenAI, between 512 and 4,096 on Claude) and most decline silently: the call succeeds, the content is sent, nothing is cached, and no error is returned. Usage().CachedTokens staying at zero is the only reliable signal. Claude also accepts at most four cached blocks per request.

Media formats are limited by the sniffer, not the vendor

An attachment's type is sniffed from its bytes with the standard library's http.DetectContentType, and only types that sniffer can positively identify are on the allowlist. Formats it cannot name (mov, flv, wmv, 3gpp, flac, m4a) are rejected with chat.ErrMediaRejected even where the vendor would accept them. A declared Media.MIMEType cannot override this; it is only ever a cross-check against the sniffed family.

Provider modules must match the core's minor version

The core and the three provider modules version independently but are not independently compatible: the provider-authoring API (ResolveAPIKey, UsageTracker, DispatchToolExecution, ValidateMediaSet, the capability markers) may change in a minor release while the module is pre-1.0.

The failure mode that bites is a core newer than the provider. Go's minimum-version selection will raise the core when your module (or any other dependency) requires a newer chat than the provider modules do, and a provider built against the older authoring API may then fail to compile. Requesting the core at @latest while pinning providers is the usual way to arrive there. Install the provider and let it choose the core, rather than requiring chat directly.

Specs are not documentation

Some pages link to a numbered spec in the project wiki. Those are point-in-time decision records: they say what was decided and what was rejected on the day, and they are not updated as the code moves. If a spec and this site disagree, this site is the one maintained against the code. If this site and the code disagree, the code wins and the page is wrong.