AI integration: providers, aifunc, markers, codeswap
How a CX+AI program reaches a model, and how code a model writes gets into a running program. Everything below is shipped behaviour, checked against the compiler and runtime that this release builds.
One principle shapes the whole surface: code a model writes lands in a SEPARATE code array; the code you wrote stays untouched. Dispatch flips between the two at a marker or an AI-function entry and flips back at its exit, so an AI body can be swapped in, swapped out or replaced while the program runs without ever rewriting the program's own instructions. The mechanism is the register VM's codeswap; see UNIFIED_REGISTER_ISA.md for the machine it runs on.
ai_set_provider("anthropic"); // or "openai", "google", "mistral", "cohere",
// "xai", "deepseek", "groq", "ollama",
// "kwaai", "custom"
json cfg; // a key is OPTIONAL -- this provider's own
cfg = parseDoc("{\"models\":[{\"provider\":\"anthropic\",\"key\":\"sk-...\"}]}");
ai_set_options(cfg); // environment variable is read otherwise
ai_init(); // 1 when a key is resolvable, 0 otherwise
println(ai_get_error()); // the last error, as a string
CX+AI is provider-agnostic: the table names ten providers plus a custom endpoint today, and adding one needs no change to the runtime you ship. Each cloud provider reads a conventional environment variable (ANTHROPIC_API_KEY, OPENAI_API_KEY, GROQ_API_KEY, and so on); the two local ones need no key at all. custom is the escape hatch: your own endpoint, same call.
The two keyless rows are the ones that answer on your own machine. ollama speaks to a local Ollama daemon. kwaai speaks to a KwaaiNet node's shard API -- an OpenAI-compatible door in front of inference distributed across the network's block servers -- and defaults to http://localhost:8080, which is KwaaiNet's own default port for it. A node that also joins the network with kwaainet run-node already has 8080 taken by the p2p listener, so start the API on another port and say so:
json cfg;
cfg = parseDoc("{\"models\":[
{\"provider\":\"kwaai\", \"url\":\"http://localhost:8090/v1/chat/completions\"},
{\"provider\":\"ollama\"}
]}");
ai_set_options(cfg); // the network first, your own machine second
ai_set_provider("kwaai");
The port belongs to the kwaai section and reaches nothing else. That matters here more than anywhere: the two models in this example run on different machines, and an endpoint that followed the chain would send your local fallback to the shard API as well.
Name your provider. ai_init() will also pick one out of your environment when you have not chosen, but the explicit call is the path to rely on.
json cfg;
cfg = parseDoc("{\"models\":[{\"provider\":\"groq\"},{\"provider\":\"ollama\"}]}");
ai_set_options(cfg);
The order you write is the order tried. models is an array rather than an object for exactly this reason — json promises nothing about the order of an object's members, so an object could not carry a chain at all.
When a call comes back as an error the next section is tried, and so on to the end. A section carries its own endpoint, model and key, so there is nothing to snapshot and nothing to restore: falling through to the second model cannot change what the first one is configured to do.
json u = ai_usage(); // what it has cost, as a document
ai_usage_reset(); // clear the counters (the budgets survive)
A budget is a field of the model that spends it — "totallimit": 200000 in that model's section — so two models on one provider keep two budgets, and a ceiling can never be applied to a model you did not mean.
The numbers are the provider's own. CX reads the usage block out of each reply in whatever spelling that provider uses — Anthropic's input_tokens, OpenAI's prompt_tokens, Ollama's prompt_eval_count, Google's usageMetadata, Cohere's meta.tokens — so what you read is what you will be billed for, not an estimate CX made up.
ai_usage() answers as one document rather than as six accessors:
{ "in": 918, "out": 47, "calls": 3, "uncounted": 1,
"providers": { "anthropic": { "in": 918, "out": 47, "calls": 2,
"uncounted": 0, "cap": 200000 },
"ollama": { "in": 0, "out": 0, "calls": 1,
"uncounted": 1, "cap": 0 } },
"last": { "provider": "ollama", "in": 0, "out": 0, "counted": 0 } }
uncounted is not a rounding detail. A reply that carries no usage block at all increments it and adds nothing to the totals, because zero is a legal token count: if CX invented a zero, "the provider did not say" and "the call was free" would be the same answer, and a budget built on that would silently never fire. last.counted says whether the last call's two numbers are real.
A cap is a spending guard, so it fires before the request. A call whose provider has reached its cap never reaches the network; it returns the empty answer every AI error returns, with ai_get_error() naming the provider, what it has spent and what its ceiling is.
And because that refusal is an ordinary error, the fallback chain carries it. This is the whole reason the cap is per provider rather than per process:
json cfg;
cfg = parseDoc("{\"models\":[
{\"provider\":\"anthropic\", \"totallimit\":200000},
{\"provider\":\"ollama\"}
]}");
ai_set_options(cfg);
ai_set_provider("anthropic");
Past the budget the paid model drops out, the next section answers, and nothing in your program has to notice. A process-wide cap could not do this — it would refuse the local model too, and the caller would get nothing instead of an answer.
A cap limits spending, not answering. A remembered rule and a cached answer both come back before anything reaches the network, so a capped model does not silence them. Only a call that genuinely needs the wire fails — and it fails loudly, never silently and never by wandering onto a model your document did not name, which would be worse than having no cap at all.
The counters are per process. A daily budget is your program's file and your program's idea of a day; the runtime counts, you decide when to reset.
Everycxexcerpt below that carries a region name is lifted from one program that runs. It ships asexamples/reference/01_manual_fragments.cx; each excerpt names the region it comes from, and the doc gate refuses the pair if they ever drift apart.
key = "task-signature-or-hash";
hit = ai_cache_get(key);
if (hit == "") {
result = ai_call(prompt);
ai_cache_put(key, result, 3600); // ttl seconds
} else { result = hit; }
req = ai_call_async(prompt); // returns request id (or <0 on error)
while (ai_async_ready(req) == 0) { delay(50); }
result = ai_async_result(req); // blocks if still pending; auto-cleans up
There is a second cache you do not call, and it is keyed on who answered. ai_call() remembers its own answers, and the key is the provider, the model and the whole prompt -- so the same question asked of two models is two questions, and the answer you get back came from the model you asked. Each provider the fallback chain walks to is asked about its own cache before a request is built for it, which is what lets a chain whose primary is down answer from the fallback's memory instead of paying for it again. A reply is stored under whoever actually answered, which after a fallback is not the provider you named.
Two consequences worth knowing before you measure anything. The same prompt now caches once per provider-and-model, so a hit rate read across a rotating chain falls roughly by the number of models in play -- that is the price of the answer being honest about its source. And ai_cache_put/ai_cache_get are a separate namespace: entries you place by hand carry no provider and no model, so your own keys can no longer collide with a query's, in either direction.
(Both behaviours arrived in the same change. Before it, the automatic cache was keyed on the first 95 characters of the prompt and compared against the whole of it -- so nothing longer than that could ever be found again, and a repeated question was paid for every time. If you have a number for this cache from an earlier release, it was measuring a cache that almost never hit.)
Declare a function and give it a task instead of a body:
aifunc function add.i(a.i, b.i) {
"Return a + b."
}
aifunc function clamp.i(x.i, lo.i, hi.i) {
"Return x clamped to [lo, hi]."
}
// The call site is an ordinary call:
result = add(3, 7); // -> 10
y = clamp(50, 0, 100); // -> 50
The body is a single string literal, and it is the task. aifunc is a modifier written before function, and the model is asked when the function is called — so the key is a run-time requirement, and two calls can differ.
Any number of int / string / float / long parameters, mixed freely.
aifunc_ctandaifunc_rtretired in 3.2.008.0, and this manual was wrong about them until it did. It taught that_ctran the model at compile time and_rtat start-up. Neither was true: the parser set ONE flag for both spellings and nothing downstream ever asked which had been written, so there was not even a distinction that had decayed — there was one behaviour with two names from the start. A compile-time round trip does not exist in CX v3 at all;cx.exelinks no provider code and could not reach a model without pulling the HTTP layer into the compiler. Both spellings now refuse with CX-E0046, namingaifunc, rather than aliasing to it — a second name for one behaviour is the thing being removed.
An AI function that could not reach a provider does not quietly return the zero of its type. It writes one line naming the function, what went wrong, and the environment variable the provider was looking for, and leaves the message readable through ai_get_error() — the same door every other call on this surface reports through:
cx ai: the AI function 'add' got no answer -- no api key for anthropic
its return value is the ZERO OF ITS TYPE and not a result; read ai_get_error() to handle this.
set ANTHROPIC_API_KEY in the environment, or give that model a "key" in ai_set_options().
An AI function works the same on the native backend and on the register VM, and that is arranged by construction rather than by agreement: the prompt is composed in one place in the runtime, and both backends drive it. Neither emitter knows the prompt's shape, so there is no second version of it to drift.
(Until 3.2.008.0 the register VM refused an AI function by name, CX-E1038 -- which was itself an improvement on what it replaced, a body that emitted nothing and returned the zero of its type with ai_get_error() empty. A refusal was still a capability that worked on one backend and not the other, so it went.)
A marker is a region of ordinary CX that runs as written, and can be replaced while the program runs.
#pragma named val
val.i = 0;
function runMarker.v() {
\{5: "Set val to a value"
val = 42; // the default body; this is what runs first
\5:}
}
runMarker();
// val = 42
generatecode("Set val to 99. Push 99 then store into val.", 5);
replacecode(5);
runMarker();
// val = 99 -- the generated body has been swapped in for marker 5
\{N: opens a marker and carries its description; \N:} closes it. N is a positive integer and is unique across the program. Until something replaces it, a marker costs nothing: the region is compiled inline and the open and close are no-ops. After replacecode(N), entering the marker runs the replacement body instead and returns to the instruction after the close.
Markers nest. Peers each have their own slot; nested markers stack. An inner marker is bypassed while its outer one is replaced, because the replacement body does not contain it.
| Builtin | What it does |
|---|---|
generatecode(task.s, N.i) | One model round-trip; the response is held for marker N. Pass "" to reuse the marker's own description as the prompt. |
replacecode(N.i) | Parses the held response and activates the swap for marker N. |
setcode(body.s) | Supplies the body yourself, as a string — no model involved. This is how a deterministic offline test drives a marker. |
clonemarker(srcId.i, dstId.i) | Copies one marker's active body onto another declared marker site. |
A marker body is not a function: no parameters, no return value, and the evaluation stack must be as it was on entry. Reach your program's data through the named slots below.
An aifunc can write a C-rule at runtime and hand it to arm() — an AI authoring entity behaviour as data. Rule text spells handle metadata exactly like every other CX program:
arm(e, "{ int n = e[\"skus\"]->count; ... }") // the metadata surface, in a rule
There is nothing special to remember and nothing to translate. ->count, ->cap, ->type, ->valid and ->id all read the same in rule text as in program text, because the rule front-end lowers them through the same shared builder to the same call the sandbox already accepts.
The call spellings do NOT work here either. On a TYPED handle, jsonSize(...), queueCount(q), listSize(xs) and the rest are CX-E1110 in rule text just as they are everywhere else (Language Reference §11) -- the reply names the field to write instead. They are not retired: on a handle held in a plain int there is no -> form at all, so the call is the only spelling and still answers. Until v3.187.0 rule text was the one exemption — because -> parses as a field read on a dereference and the rule sandbox declines that whole node class — so an AI trained on documentation older than v3.187.0 will write the call form. Regenerate the prompt, do not translate.
Why this matters more in a rule than anywhere else: a rule that fails to compile makes arm() return 0 and the rule simply never fires — there is no error at the call site. Wrong metadata spelling in rule text is behaviour that silently does nothing, in either direction.
What -> does NOT open in a rule. Only the metadata fields. e->hull is still refused (it is not a spelling of e["hull"]), pointers are still refused, and a metadata field on something that is not a container handle is still an error rather than a 0.
#pragma named val
val.i = 0;
#pragma named marks a global as reachable from generated bytecode, and it accumulates across lines rather than replacing what came before. Container forms — #pragma named array NAME.t[N], list NAME.t, map NAME.t — emit the declaration and the binding together.
A named slot is not copied into the AI machine and back; it is read and written in place, through a table whose entries point at the real storage.
Underneath the declaration form there is a set of builtins that do each step by hand. They are supported, and they are what an offline test drives:
#pragma named val
val.i = 0;
rt_createaifunc("setval42", "Set val to 42."); // @todo VERIFY -- see note below
result = ai_call_func("setval42", 0);
src.s = "PUSHI 99\nSTORE val\nRET\n";
ai_parse_bytecode("setval99", src); // a body you wrote yourself
ai_call_func("setval99", 0);
println(ai_dump_asm("setval99")); // read it back as text
✅ SETTLED 2026-09-03 (3.2.008.0), and the answer is that there was never a compile-time path at all. The 2026-07-23 note below asked which of two compile-time claims this example meant, and refused to guess — correctly. Both were wrong. Thert_-prefixed family is honestly named: every one of these calls runs the model at RUN time. And theaifunc_ctdeclaration in §3, named as "the compile-time path that certainly exists", did not exist either —_ctand_rtset the same parser flag and emitted the same wrapper.cx.exelinks no provider code, so nothing in this manual's compile-time story could have run. Both spellings retired; §3 says what actually happens. The original note, kept because it was right to ask: the family isrt_createaifunc/rt_callaifunc/rt_runaifunc/rt_destroyaifunc/rt_aifunc_count— allrt_-prefixed, with noct_variant.
rt_createaifunc, ai_call_func, ai_parse_bytecode and ai_dump_asm all work today. (ai_set_named / ai_get_named were RETIRED 2026-09-03: a #pragma named slot IS the program's own variable, so hp = 75 writes what the bytecode's FETCH hp reads -- and unlike the retired pair, assignment carries floats and strings and lets the compiler catch a misspelled name.) New code should prefer the declaration form of §3 and the markers of §4.
setcode + replacecode drive a marker with a body you supply, so a test can exercise the whole swap without a provider, a key or a network:
function runM.v() {
\{7: "default"
val = 1;
\7:}
}
runM(); // val = 1
setcode("PUSHI 99\nSTORE val\n");
replacecode(7);
runM(); // val = 99
This is the recommended shape for a deterministic test. The declaration form of §3 also tries to generate at compile time, which a test usually does not want.
Generated bytecode does not have a machine of its own. It is translated instruction for instruction into the register VM's own instruction set and runs there — the same VM your compiled program runs on — so there is one executor behind the declaration form, the markers and the direct surface alike. Named slots reach host storage in place through that translation, and CX's containers (arrays, lists and maps) are bridged into it rather than copied.
The HTTP layer is provider-shaped rather than provider-specific: one call carries the provider id, the model, the key, the system and user prompts, the token cap, the temperature and a timeout, and each provider supplies its own request body and response reader. That is what makes "adding one needs no change to the runtime you ship" a description rather than a promise.
#pragma named, and the rest of the language the excerpts above are written in.