CX+AI

CX+AI Language Reference

The language, end to end

CX+AI Language Reference

Version 3.137.1 | July 2026 | Generation 3 (pure C)

This replaces the V1.x reference. CX's implementation moved from PureBasic (three generations: alfa → beta/v02 → retired 2026-05-30) to pure C. The language most users write hasn't changed shape as much as the machine under it has — this document covers both, and calls out the handful of places where user-visible syntax genuinely moved (struct field access, marker syntax, new return-type suffixes).

Table of Contents

  1. Overview
  2. Getting Started
  3. Syntax Basics
  4. Variables and Types
  5. Operators
  6. Control Flow
  7. Functions
  8. Arrays
  9. Structures
  10. Pointers and Function Pointers
  11. Collections — Lists, Maps, and Queues
  12. Strings
  13. Math Functions
  14. Type Conversion
  15. Files, Filesystem, and Memfiles
  16. Archives, Compression, and Networking
  17. Graphics and GUI
  18. JSON
  19. XML and Regex
  20. SQL and SQLite
  21. Testing
  22. Cryptography
  23. Date/Time
  24. System
  25. Pragma Directives
  26. Markers and Runtime Code Modification
  27. AI Integration
  28. Rules and Automation
  29. Embedding Files and the Asset Resolver
  30. The 2D UI Layer (CGI2D) and Where It's Going
  31. Libraries — Status
  32. Appendix: Complete Example

1. Overview

What is CX?

CX is C eXpanded — the C language made easier by being clever, never by losing power. It is not a replacement for C: it compiles through C, retains every escape hatch, and runs on any C compiler you choose.

The design goal is the minimum code to achieve the objective. The mechanism is to put complexity into the variable and its type suffix, so the source stays flat and readable:

The implementation

CX is a C-like language whose compiler and runtime are now written entirely in C (gcc/clang, C99/C11). It compiles to native binaries via a C transpiler path — there is no more "compile to portable bytecode, always run on a VM" model. A CX program is, by default, straight-line native C. Only the parts you explicitly mark as AI-mutable, or that the AI itself generates at runtime, need an embedded register VM at all — and even then, once that bytecode can be decompiled back to source (§26), it can be recompiled and run as native C too, rather than interpreted indefinitely.

Key facts, current build:

Compiler pipeline (current):

Source (.cx) → Preprocessor → Scanner → Parser → AST
   → emit_c (native path)  ─┐
   → emit_risc (VM path)   ─┴→ backend cc (gcc/clang/msvc) → native .exe

There's no separate post-processor/optimizer/FixJMP/serialize stage anymore — each emitter writes its final form directly, and the "bytecode file you distribute and always VM-execute" model (.ocx as the only way to ship a program) is gone. .ocx still exists as an internal artifact for diagnostics and for the register-VM path, but a normal build produces a real executable.


2. Getting Started

Command Line

cx <file.cx> [options]
OptionDescription
--buildTranspile + backend-compile to an exe; don't run
--runBuild then execute; process exit code propagates; -- args… passes argv
-o FILEOutput path
-t, --terminalHeadless — skip opening a debug console
--keep-cWith --run, keep the generated .c
-P pcode=riscRequest the register-VM path. For a program with an explicit main(), this still emits native C by default and embeds the bytecode only as inert metadata for inspection — the VM interpreter itself only actually runs for the top-level-statement (no-main) program shape. Check with grep cx_risc_run_embedded on the emitted C if you need to know which you got.
--emit-riscDump the register-VM bytecode listing (function names, [_local] slots, escaped string literals, jump targets — this is today's ASM listing)
--decompileReconstruct CX source from register-VM bytecode
--run-liftMode 4. Decompile a program's own register-VM bytecode back to CX, recompile it through the native path, and run that in-process — recovers most of the VM's interpretation cost for whatever the decompiler can faithfully lift (§26)
--target wasmBuild the same source as a web page instead of a native executable — emits <name>.html + <name>.js + <name>.wasm. A target is a toolchain, not a second code generator: same emit_c output, a different C compiler, and nothing in the source knows the difference. Graphics needs no extra instruction, and #pragma rules c still works, so a rule authored at run time compiles and fires in the tab. Needs the emsdk toolchain and refuses loudly without it (CX-E2005, naming the environment variables and the activation command) — it never quietly builds a native binary instead
--sharedBuild a shared library instead of an executable — .dll on Windows, .so on Linux, .dylib on macOS (also #pragma BuildShared yes). No main() is emitted: the runtime starts when the library is LOADED and shuts down when it is unloaded, so an export can be called the moment the host has the handle. Every CX function is exported, which makes the functions you write the library's API — and on Windows it is also what keeps the runtime's own symbols private, since a DLL that declares no exports of its own has all of them taken. Top-level statements still run, at load, under the platform's loader lock: that is the module body, so keep it to declarations and cheap setup and put real work in a function the host calls when it is ready. gcc and clang only, and the combinations that cannot produce a library — --run, --target, -P pcode=risc, a C compiler with no __attribute__((constructor)) — are refused by name, each with its own code (CX-E2009 a run, CX-E2010 a cross target, CX-E2011 the VM whole-program mode, CX-E2012 the toolchain) rather than half honoured. See Conclave/tests-internal/com_spike/ for a CX shared library that loads into Microsoft Outlook as a COM add-in
--tokens, --astFront-end dumps
-P key=valInject a pragma from the CLI (repeatable)
--version, -VPrint the baked version
-h, --helpHelp

--emit-cisc no longer exists — there is one VM. .ocx compile-only output and the old -C/--no-source/--no-od flags from the PureBasic era are gone; use --build -o FILE to produce a binary.

File Types

Hello World

println("Hello, World!");
cx --run hello.cx

Hello, AI-Generated World

aifunc function fib.i(n.i) {
   "Return the n-th Fibonacci number recursively."
}

println(fib(10));   // 55

aifunc asks the configured LLM when the function is called, through the same provider fallback chain as every other AI call (§27). (The aifunc_ct / aifunc_rt spellings retired in 3.2.008.0 with CX-E0046. They were never two behaviours: both set one parser flag, and CX v3 has no compile-time model round trip at all -- cx.exe links no provider code.)


3. Syntax Basics

Comments

// single-line

/* multi-line
   comment */

Statements

Statements end with ;. Blocks use { }.

Calls Are Expressions — Builtins Nest

A call is an expression, so it can be an argument to another call, part of a larger expression, or the right-hand side of an assignment. That holds for builtins and for your own functions alike, to any depth, and it is the ordinary way to write CX rather than a trick — there is no temporary you are obliged to introduce for the middle step:

function demo.v() {
    list rows.s;
    listAdd(rows, "  Widget , Tools , 3  ");
    line.s = rtrim(listGet(rows, 0));                  // a call as another call's argument
    qty.i  = (int)(trim(stringsplit(line, ",", 2)));  // three deep
    printf("%-10s %3d\n", trim(stringsplit(line, ",", 0)), qty);
}

The one thing to keep in mind is the return TYPE: a call is whatever its declaration says it is (§7), so strtoi(...) is an int wherever it appears and trim(...) is a string. Nesting changes nothing about that — it only saves the variable you would otherwise have declared to hold the intermediate value.

Case Folding

CX source is effectively case-insensitive: the scanner lowercases identifiers before they reach code generation, so myVar, MYVAR, and MyVar name the same variable. (The compiler itself is written in case-sensitive C — this folding is a parser behavior, not a C-language one.)

There is no opt-out — no directive, no flag, no mode. The PB era had a switch for it; its off mode also stopped folding CALL names, so parseDoc() failed as an unknown call and no real program could run under it. It was retired in v3.172.0 and the fold has been unconditional ever since.

The consequence worth naming is that two declarations differing only in case are ONE variable, silently — a global A and a global a share a cell, with no diagnostic. Single-letter names collide with scratch names soonest, so prefer matA / matB over A / B.

Preprocessor

#define MAX_SIZE 100
#include "other_file.cx"

#cinclude "engine/hud.c"     // splice raw C at file scope — new; see §7

#cinclude "file.c" / #cinclude <header.h> spices a verbatim C file into the generated C at file scope (before main), so it can define global C functions/types/data. Pair it with foreign declarations to call into it from CX.

A macro name folds to UPPERCASE. Identifiers fold down (§ above), so MAX, Max and max are one name whichever of them declared it — but the canonical spelling of a #defined name is the upper-case one, and that is deliberate rather than incidental: "since variables fold to lowercase then macros go uppercase; its an aid". The case tells a reader which kind of name they are looking at, with nothing to look up.

One name means one declaration. A #define, a _json block, a _rules block and a variable share ONE name space, so #define MAX 10 beside a variable max is CX-E1121, naming both lines — including a variable declared inside a function, which the preprocessor reaches by refusing to expand a macro reference standing in declaration position (a name written with its type suffix, max.i). Its one limit: a C-typed declaration — int max; — has no suffix to stand on, so that collision is still refused but arrives as CX-E0005 "expected variable name after type", naming the symptom rather than the cause.

A #defined name is substituted before the parser sees it, so assigning to one is assigning to a constant — #define K 5 followed by K = 7 becomes 5 = 7, and is refused with CX-E1061. (New in v3.223.0: this used to reach gcc as "lvalue required as left operand of assignment" against generated C on native, and a wrong "native-only construct" on the register VM.)


4. Variables and Types

Scalar Suffixes

SuffixTypeNotes
.i64-bit intdefault when no suffix is given
.fdouble-precision float
.l64-bit longfor when you want the width explicit; int64_t/uint64_t/longlong all fold to this
.sstringvalue-isolated, UTF-8, GC-managed (§12)
.vvoidreturn-type only — not a variable type
.ppointera slot handle, not a raw memory address (§10)

C-style declarations are also accepted (int n = 0; string s = "hi"; float f = 0.0;); the two styles freely mix in one file. C-stdint aliases (int8_t..uint32_t, size_t) fold to .i; unsigned/signed is accepted with a warning (no signedness tracking).

count.i = 0;
pi.f = 3.14159;
message.s = "Hello";

x = 42;            // inferred int
name = "Alice";    // inferred string

Global vs. Local

File-scope declarations are global; declarations inside a function body are local to it. Params and in-body declarations shadow a same-named global, as in C.

Undeclared names inside a function

A name you never declare is still legal — CX declares it on first use. The rule is one line: an undeclared name binds a same-named module-global if there is one, and otherwise becomes a fresh local. That is what CX has always done, and the compiler now tells you every time it makes that call, because whether a bare i is your loop counter or the file's global i was invisible in the source and is the kind of thing that costs an afternoon.

what you wrotewhat happenswhat the compiler says
x.i = 0 — an explicit declarationa fresh local, shadowing any globalnothing (this is the fix)
a parameter named xa fresh localnothing
bare x = 1, no module-global xa fresh local, type inferred (int if it can't be)nothing
bare x = 1, module-global x existsbinds the global — the write leaves the functionCX-W1010
bare x read before anything writes it, declared nowhereassumed int 0CX-W1012
bare x = 1 where x is a module-global list/map/array/json/queuerefusedCX-E1056
bare x = 1 where x is a function's namerefusedCX-E1057
x read on a line above this function's own declaration of xrefusedCX-E1058
x["k"] = 1, x declared nowheremints a json document — see §18nothing
x["k"] read, x declared nowhere and never writtenrefusedCX-E1091
x where a struct x { … } also existsworks — the variable is renamed in the emitted Cnothing

CX-E1056, CX-E1057 and CX-E1058 are errors rather than warnings because there is no useful reading of them: before they were diagnosed, the container case crashed the program and the json case looped forever, and a warning followed by a segfault would be worse than the silence it replaced. CX-E1091 is the same judgement one operator over — a subscript write to an unknown name has a document to mint, and a read has nothing to read.

A variable may share a struct type's name

struct point { x.i  y.i }

struct point point;    // a struct variable named after its own type
point.i = 5;           // ...or a plain int of that name, at module or function scope

C keeps typedefs and variables in one namespace, so neither of these can be spelled literally in the generated C. CX renames the variable to cxv_point there and leaves you the name you wanted — you never see the renamed form unless you read the emitted C. It costs nothing at runtime and applies to every kind: scalars, arrays, lists, maps, sortndx, struct variables, module-globals and function locals alike.

Until v3.221.8 this was true of every use of such a name but only some of its declarations, so a few shapes reached the C compiler with the two spellings disagreeing and failed with messages about a cxv_ name you never wrote.

Read it after you declare it (CX-E1058)

function f.v() {
    printf("%d\n", a);    // CX-E1058: 'a' is read before its own declaration on line 3
    a.i = 7;
}

This is C's rule — a name is usable from its declaration onward — and where the surface is C's, CX answers as C does. It is about position, not existence: declaring a mid-function is perfectly ordinary CX, and every read after the declaration is untouched.

It is worth knowing what this replaced, because the old behaviour was silent. Native builds declarations at the top of their block, so the early read used to answer that fresh local's 0 with no warning at all; the register VM declares in source order, so it refused with "unknown identifier" — about a name declared on the very next line. And when a module-global of the same name existed, the two backends ran the program and printed different numbers (native 0 from the hoisted local, the VM the global's value), both at exit code 0. One diagnostic at the read, on both backends, is the whole fix.

The compiler reports the first early read of each name per function — the second is the same mistake in a second place, and the line it names is the declaration you need to move.

This is a different question from the undeclared lane above: x = 1 with no type tag is not a declaration, so it stays in the ruled mint-or-bind world and never fails a build. CX-E1058 needs an actual declaration, further down.

#pragma localdeclares and _local — C discipline on demand

If you would rather an undeclared name always be local, say so:

#pragma localdeclares on      // whole file; default off

_local function tick.v() { ... }      // or one function at a time
function tick.v() _local { ... }      // same thing, either position

_forcelocal function tick.v() { ... } // heritage spelling, identical meaning

In a strict function an undeclared name mints a local that shadows any same-named global, and the compiler reports each one with CX-W1011 naming what it shadowed. Strict mode never turns a program into a build failure — it changes which way the compiler guesses, it does not add a way to lose. It also dissolves the two refusals above by construction: a local int has no container to crash into. An explicit declaration silences either warning.

_forcelocal is a heritage alias for _local — same bit, same behaviour, either position. _local is the canonical spelling; write that in new code. It used to mean something else: it forced a function to get its own call frame so locals survived recursion. That is automatic on both backends now (native locals are C automatic storage; the register VM's CALL/RET save and restore the callee's window), including recursion through a function pointer, so the marker had stopped doing anything at all. Rather than leave a spelling that silently did nothing, it was pointed at the nearest thing it looks like it means. Existing _forcelocal functions keep compiling and keep framing exactly as before; what changes is that undeclared names inside them now mint locals and report CX-W1011.

Module-level code is untouched by the pragma: out at file scope x = 5 is the global's declaration, so there is nothing to be strict about.

A nested block is its own scope (v3.292.0)

A declaration inside { } belongs to that block, on both backends, exactly as in C. It shadows an enclosing declaration of the same name, and the name comes back at the closing brace:

Every cx excerpt below that carries a region name is lifted from one program that runs. It ships as examples/reference/01_manual_fragments.cx; each excerpt names the region it comes from, and the doc gate refuses the pair if they ever drift apart.
int x; x = 7;
if (ready) {
    int x; x = 9;          // a DIFFERENT x, this block's own
    println(str(x));       // 9
}
println(str(x));           // 7 -- the outer x was never written

Three consequences worth knowing, all of them C's:

x.i = 9 is not a declaration in this sense — it is a typed assignment, and inside a block it writes the enclosing x. Write int x; when you mean a new one.

Until v3.292.0 this was true natively and false on the register VM, which allocated one slot per name per function: the inner declaration wrote the outer variable, on every type, with no diagnostic from either backend.

Constants

#define MAX_ITEMS 100
#define TAX_RATE 0.08

5. Operators

Standard C precedence throughout: arithmetic (+ - * / %), comparison (== != < <= > >=), bitwise (& | ^ ~ << >>), logical (&& || !), assignment (= += -= *= /= %= ++ --), ternary (?:), address-of (&).

There is no ** power operator — CX is C, and C has no such operator. Use pow(base, exp). This documented one until 2026-07-23, which is worth a warning because the mistake is not caught: a = 2 ** 3; parses as 2 * (*3), emits a dereference of address 3, and segfaults with no diagnostic. Logged as a defect (it should be a clear CX error, per FP4).

A few things worth knowing that weren't true of the older interpreter:

Evaluation order is DEFINED: left to right, on both backends — wherever two or more operands can each change something. C leaves the order of a + b's operands and of a call's arguments unspecified, which is exactly what permits an implementation to pin it, and CX pins it in that case so the program cannot answer differently on native and on the register VM. Where only ONE operand changes something, it is NOT pinned and the backends CAN differ — see the measured note below the example:

int n = 0;
function bump.i() { n = n + 1; return n; }

printf("%d %d\n", bump(), bump());   // always "1 2" -- never "2 1"
int r = two(bump(), bump());         // two() receives 1 then 2
s = str(take(q)) + str(take(q));     // the queue is drained left to right

The compiler pins it only where the order is observable to the pin — two or more operands or arguments that can each change something (a call to one of your functions, or a builtin marked as mutating). Everything else costs nothing, which is why an expression with at most one such call emits exactly what it always did. (v3.245.0 operands; v3.259.0 call arguments.)

THE LIMIT OF THAT, MEASURED, AND IT IS NOT ONLY ++. "Two or more that can each change something" is the rule as implemented, so ONE mutating call beside a plain READ of what it mutates is NOT pinned — and the two backends answer differently. A -> metadata read changes nothing, so it does not count towards the two: ``cx function demo.v() { queue q.i; queueInit(q, 4, 0); queuePush(q, 9); printf("count=%d take=%d\n", q->count, queueTake(q)); } ` prints count=1 take=9 on the register VM — the written order — and count=0 take=9 natively, where the argument list is C's and gcc evaluates it right to left, draining the queue before the count is read. The same holds for a container read beside a call that mutates that container. **Do not write both in one argument list**: read into a variable on its own line first, which costs nothing and answers the same on both backends. **The other gap: ++/-- inside one expression.** f(++a, ++a) and ++a + ++a modify a` twice with no sequence point between, which is undefined in C rather than merely unspecified — CX has not pinned it and does not diagnose it. Assign to a variable first.

6. Control Flow

if/else if/else, while, C-style for, switch/case/default, break/continue — all unchanged from a C programmer's expectations.

foreach

There is one form. The collection's own cursor is the iterator — no loop variable is introduced — so the body pulls the current element explicitly:

foreach myList { v = listGet(myList); print(v); }
foreach nums   { print(arrGet(nums)); }         // works on a grown array too
foreach ages   { print(mapKey(ages), " = ", mapValue(ages)); }
foreach ranks  { print(get(ranks)); }           // sortndx, in sorted order

Iterable kinds: list, map, array (fixed, multi-dim, or grown), sortndx, json. A queue is a container but is not iterable this way — take from it (queueTake) in a while instead.

For json, foreach walks the members of whatever the handle names, and a parsed document is deref'd to its root — so foreach doc over parseDoc("[10,20,30]") runs three times, the same three ->count reports. It is the same walk whether you hold the document or a node inside it:

Rows of objects — the shape a sqlQuery returns — are read by SUBSCRIPTING the cursor, exactly as any other json handle is read. jsonGet(rows) names the current row; ["customer"] reads a member of it; the value coerces at the sink, so no jsonAs* call is needed in a printf argument or a typed assignment:

doc = parseDoc("{\"items\":[10,20,30]}");
foreach doc            { ... }              // the document's members
items = doc["items"];
foreach items          { n = n + jsonGet(items); }              // 10, 20, 30
foreach doc { row = jsonGet(doc); foreach row { ... } }         // nesting is fine

rows = sqlQuery("select customer, spend from orders");
foreach rows {                                                  // rows of objects
    printf("  %-6s %d\n", jsonGet(rows)["customer"], jsonGet(rows)["spend"]);
}
sName = rows[0]["customer"];                                    // by index, outside a foreach

Reach for jsonAsStr / jsonAsInt / jsonAsFloat where the sink does NOT coerce — most often an argument to another builtin, which takes the value as written: jsonArrInt(args, jsonAsInt(rows[0]["id"])) stores the id, while jsonArrInt(args, rows[0]["id"]) stores the node's handle instead.

One limit: the cursor lives on the node, so you cannot run two iterations of the same node at once (foreach doc { foreach doc { } }). Different nodes — the nested form above — and a second pass afterwards are both fine. (Before v3.219.6, foreach over a parsed document ran exactly once, visiting its root container instead of the members.)

Anything else is refused at the .cx line, identically on both backends:

you wroteyou get
foreach v in c { } or foreach (v in c) { }CX-E0044 — the foreign spelling; neither has ever been CX
foreach n { } where n is an int, string, queue, …CX-E1053 — names the type it was declared with
foreach ghost { }CX-E1054 — not declared anywhere in the program

(Until v3.219.3 this section documented foreach (item in myList) as a current form. It never parsed — and the three invalid shapes above were silent no-ops on at least one backend, which is why they now have codes.)

An open design question (not yet built, don't rely on it) is a foreach (x : c) colon form — noted here only so you don't go looking for it.

Codeswap Markers

Syntax changed from the PureBasic era — markers are now backslash-braced:

\{5: "description"
   ... default code ...
\5:}

Full detail in §26.


7. Functions

Declaration and Return Types

function add.i(a.i, b.i) { return a + b; }
function greet.s(name.s) { return "Hello, " + name; }
function log.v(msg.s)    { print(msg); return; }   // .v: bare return only

Return-type suffixes: .i (default), .s, .f, .v (void), .p (pointer), .l (long), .StructName (struct return — new; the function must declare it: function origin.Point() { Point z; return z; }).

A .v function makes return <expr>; a compile error — use bare return; to exit early. A .v function need not end in return; at all; the bare form exists for the EARLY exit.

AN UNSUFFIXED FUNCTION IS INT, AND INT IS NOW HELD TO (v3.3.007.0). Omitting the suffix still means .i — that default is unchanged, there is no mandatory suffix and no inference — but all three ways of putting a non-int through it are refused, where previously only the first was:

you wroterefused with
function f() { return "hi"; }CX-E1076 — a string has no lossless numeric reading; use (int)s or (float)s
function f(x.i) { return x / 2.0; }CX-E1123 — declare the function .f, or write (int)x to make the truncation explicit
function f() { println("hi"); }CX-E1124 — return a value on every path, or declare the function .v

CX-E1124 covers both spellings of one defect: falling off the end, and an explicit bare return; inside a non-void function.

What those two rows did before v3.3.007.0 is the reason they are refusals now. The middle one answered 2.000 natively and 2.500 on the register VM — one program, two answers, neither of them announced. The last gave undefined bytes natively and a quiet 0 on the VM, and the quiet 0 is the more expensive of the two: it is a plausible value for a counter, an index or a length, so a program carries it a long way before anyone asks where it came from.

The check is deliberately conservative and refuses only where it can prove no value comes back. A switch, a _C{} escape and a loop that never ends are all accepted without further analysis — it would rather miss a case than reject a correct program.

ORDER DOES NOT MATTER. A function may be called from a line ABOVE its own definition, with no forward declaration and no prototype — the whole file is parsed before anything is emitted, on both backends:

greet("world");
function greet.v(who.s) { println("hello, " + who); }

That is one of the places CX is deliberately not C. Put your entry point wherever it reads best.

Parameters

All of these are accepted, and can mix in one signature:

function calc(x, y.f, label.s)                 // untyped defaults to int
function calc(int x, float y, string label)     // C-style
function f(struct Point p)                      // struct param — byref (see below)
function f(byref n.i)                           // explicit byref primitive
function f(byval n.i)                           // explicit byval (redundant; it's the default)
function connect(host.s, port.i = 8080)          // default value, trailing params only

A call must supply between the required and the declared number of arguments — defaults make it a range, so connect("h") and connect("h", 99) are both fine and connect() is not. Outside that range the call is refused with CX-E1062, naming what you passed and what the function takes. (New in v3.223.0. Native used to hand this to gcc, which answered "too many arguments to function" at a line in generated C you cannot open; the register VM compiled the call and failed at RUN time with an arity error. CX-E5012 still covers the calls no compiler can count — through a function pointer or a fn value.)

A function that returns a function value declares .fn (or .fp), and then the value prints as <fn> wherever it lands — in the call, or in a variable bound from it:

function pick.fn()  { return &add; }
print(pick());                       // <fn>, both backends
function raw.i()    { return &add; }
print(raw());                        // a raw number — see below

(New in v3.227.0. The declaration is what CX reads: a function returning a function while declaring .i keeps the raw lane — an address natively, a function index on the register VM — because inferring the type from the body would be guessing, and two returns in one function can disagree. Before this, even the DECLARED form printed the raw lane.)

An untyped parameter is an int, and CX takes that literally on both backends. Three consequences, all decided at the CALL:

what you passwhat happens
an intordinary CX. Nothing is said.
a container or a stringCX-E1063 — the parameter has no type, and a handle is not an int
a floatit runs, TRUNCATED to the int, and says so with CX-W1013
function total.v(xs)        { foreach xs { } }   // CX-E1063 at the call
function total.v(list xs.i) { foreach xs { } }   // the fix — one suffix away

function passf.f(x) { return x; }
print(passf(2.5));                             // CX-W1013; prints 2.000
function passf.f(x.f) { return x; }
print(passf(2.5));                             // 2.500

(CX-E1063 new in v3.224.0, widened to strings in v3.228.0; CX-W1013 new in v3.228.0. Before this the register VM RAN the refused shapes — its slots are typed at runtime — while native handed gcc a cx_list_int * for a long long; and the float lane printed 2.000 natively against 2.500 on the VM, silently. The truncation is not a defect being papered over: an untyped parameter IS an int, so C truncates, and CX answers as C answers — the warning exists because the collapse is lossy, not because it is wrong. The check is per CALL, so a function nothing in the file calls is not judged.)

Byref Semantics — the one thing most different from plain C

Compound values — struct, list, map, array — are byref by rule when passed as parameters. Primitives are byval unless you write byref. A list/map parameter is a heap handle: the callee shares the same container, and mutations propagate to the caller with no byref keyword needed.

function fill.v(list xs.i) { listAdd(xs, 1); }   // caller's list gains the element

Array parameters historically required an honest compile error; the current compiler instead promotes a passed array to dynamic at the call site (v3.10.0+): the array becomes a cx_array grown handle, and inside the callee re-stating array a.i[n] resizes the caller's array, a->count reads the live length, and element reads/writes propagate. This is precise — only a user-function call with an array-typed param triggers it; builtins like sort/len still operate on fixed arrays in place. (Re-measured against the shipped 3.271.0 compiler on 2026-08-14, both backends identically: a callee's a[0] = 99 is visible to the caller, and a callee's array a[8] leaves the caller reading ->count 8 with the new element in place. This paragraph used to end by telling you to go and check that yourself — which is the sentence a reference writes when nobody has.)

A brace initialiser survives the promotion (v3.279.0). The declaration's values are written into the grown handle at the declaration, so passing the array changes where it lives and nothing else:

function peek.i(array a.i) { return a[4]; }
function demo.v() {
    array nums.i[5] = {10, 20, 30, 40, 50};
    printf("%d %d\n", nums[4], peek(nums));   // 50 50
}

Before v3.279.0 the promotion dropped that initialiser: both numbers printed 0, including the caller's own nums[4], on both backends — the same line answered 50 if you removed the call. It reached int, float and string elements, multi-dimensional arrays and file-scope arrays alike.

Reassigning one container variable to another (a = b, both containers) is dropped with a compile warning — containers share by reference, so a plain rebind would alias two owners onto one object (a double-free). Pass byref into a function instead. This covers json too (v3.279.0 — until then a json a = b; rebound silently, in both declaration forms, which is exactly the aliasing the rule exists to prevent).

Every container family may be a parameter, and the natural spelling works inside the callee (v3.280.0). A map parameter takes a subscript store, and a sortndx parameter takes its own verbs:

function tally.v(map m.i)     { m["hits"] = m["hits"] + 1; }
function rank.v(sortndx s.i)  { add(s, 5); add(s, 1); }

Every VERB of the family works on a parameter too, not only the two shown: mapPut/mapGet/mapHas/mapDelete on a map parameter, listAdd/listGet/listGetAt/listSort on a list parameter, queuePush/queueTake on a queue parameter, and ->count on any of them — the callee holds the caller's container, so there is nothing to hand back and no separate spelling to learn.

Before v3.280.0 both of these compiled natively and were refused by the register VM — m["k"] = 1 as cannot compile subscript-assign, and add(s, 5) as unknown call add, because the VM learns a parameter's family from one enumeration and those two families were missing from it. mapPut(m, ...) had always worked, which is what made the subscript form a trap rather than an unsupported feature.

Function Pointers

fp is the canonical keyword for a function-pointer type (fn is kept as an alias):

fp callback = &add;
result = callback(3, 4);

array *ops[2];
ops[0] = &add;  ops[1] = &sub;
result = ops[0](10, 5);

A function NAME on its own is not a value. f means the function; f() calls it; &f passes it. Writing the bare name where a value is wanted is refused (CX-E1106) rather than quietly compiled to the function's address as an integer, which is what it used to do — and into an .i parameter that produced a plausible number and no diagnostic anywhere. Two places keep the bare form, because in both it is a name rather than a value: a comparator argument (sort(xs, descI), sortndx ranks.i by descCmp), and a call into C — qsort(x, n, sizeof(double), cmp_dbl) is correct C, C has function pointers, and CX keeps all of C.

What CX does with your main()

CX synthesises the program's real entry point, so a main() you write always moves aside — and what happens next depends on whether your file has anything else at the top level.

Your fileWhat runs
main() and nothing else at top levelyour main() is the program (this is what lets verbatim C source run unchanged)
main() plus module-level statementsthe module-level statements run first, then your main() is called

The second case warns (CX-W1018) so the arrangement is never a surprise. Before v3.311.0 it did something else entirely: your main() was renamed and never called, with no diagnostic and exit 0.

Your return value is the program's exit code, in both shapes and on both backends — the same thing it means in C. A main.v() has no value to exit with, so the program exits 0.

argc and argv are forwarded when your signature takes two parameters — and the two spellings are handed two different things, because they asked for two different things:

Your signatureWhat the second parameter is
int main(int argc, char **argv)the real process vector, exactly as in C
function main.i(argc.i, argv.i)the args document (see below)

Before v3.2.009.0 the CX spelling also got the raw char ** in an int slot — C's convention leaking into CX. It now gets a value CX can actually read. (You rarely need either: args is pre-bound and always there.) Both backends run it: the register VM has the same command line the native build does, so its entry call is built with the same two arguments.

The command line is a json document — args

args is already there. It is a writable json array of the arguments you passed, handed to the program pre-allocated — no call, no declaration, nothing to open:

if (args->doc->count == 0) { println("usage: report <file> [--wide]"); return 1; }
path.s = args[0];

args holds the user arguments only, so args->doc->count is 0 for a bare invocation and args[0] is the first thing you actually typed.

argc() and argv(n) still work, unchanged and C-identical — they read that same document. So there are two views of one datum, and they are deliberately one apart:

argv(0)the program path
argv(n)is args[n-1]
argc()is args->doc->count + 1

That is the one thing here worth reading twice. argv(0) is the path because that is what C means by it — which is also why there is no progpath() builtin: argv(0) already is that.

Both views are live. Write through args and argv(n) says the new value; append to args and argc() moves. Patching your own command line before you parse it is a supported thing to do.

Reading past the end answers null, not "" — so jsonIsNull(args[n]) tells "there is no such argument" apart from an argument that is genuinely the empty string. A negative index is out of range for the same reason and answers the same way.

Two places have no command line at all, and they say so rather than pretending: a --shared library entered through an export, and a _CM{} program that has not called cx_rt_init(argc, argv). There args is empty and argc() answers 0 — a number a real process never produces, since C guarantees at least the program path. A _CM{} program that does call cx_rt_init gets its real command line, and hears nothing.

If you declare your own args, it is yours — the pre-bound one is simply not there, exactly as any shadowed name behaves.

Turning it into parameters — SplitParameters()

One call parses the whole command line:

json p = SplitParameters("-,--", "=", "--out,--in");

There is no separator argument: the OS split the line before your program started, so quoting is its business and this never parses a quote.

You get one entry per parameter, in order — positional arguments included, with flag: false:

report . --out=R:/tmp -v --in data.csv

[{"flag":false,"name":"report"},
 {"flag":false,"name":"."},
 {"flag":true,"name":"out","option":"R:/tmp"},
 {"flag":true,"name":"v"},
 {"flag":true,"name":"in","option":"data.csv"}]

Three states, kept distinct, because collapsing any two would make one readout mean two things:

flag absentnot in the array at all
-vpresent, no option key — jsonIsNull(p[i]["option"]) is 1
--out=present, option is "" — jsonIsNull is 0

A value-taking flag sitting last, with nothing after it to consume, is reported by name and left with no option — never given an invented empty one, which you could not tell from --out=.

Anything this does not do, you do by patching args before you call it: it is ordinary writable json.

A main() in the C escape takes the entry — _CM{}

A main written in raw C is a declaration: you are taking the program's entry. Say so by spelling the block _CM{}, and CX steps aside — the block is spliced at file scope and no main is synthesized.

string tag;
tag = "ready";              // module-level statements still run, first

function status.s() { return tag; }

_CM {
    int main(int argc, char **argv) {
        cx_rt_init(argc, argv);          // hand CX the command line (optional)
        cx_str_handle s = status();
        printf("%s\n", cx_str_peek(&s, (size_t *)0));
        return 0;                        // your return value is the exit code
    }
}

Nothing is lost by taking the entry. The runtime startup and your module-level statements run before your main, from the C runtime's static-initialiser list, and the shutdown runs at exit — the same mechanism --shared uses, and for the same reason: the entry belongs to somebody else, so a prologue cannot be decorated onto it. The one thing that does not arrive by itself is the captured command line: a static initialiser has no arguments, so args is empty and argc() answers 0 until you call cx_rt_init(argc, argv) yourself. Call it and everything catches up — including an args your module-level statements already touched, which is topped up in place rather than replaced. Skip it and the runtime says so on the way out, because an empty args would otherwise be indistinguishable from a program nobody passed anything to. Your own main parameters are the real ones either way.

A plain _C{} at top level that defines a main means exactly the same thing and CX says so once (CX-W1020) — an entry that moves without you saying so is a surprise, and _CM{} is the spelling that removes the line. Three shapes are refused rather than guessed at: _CM{} inside a function body (CX-E1107 — it is spliced where it stands, so there is no file scope to take), _CM{} with no main in it (CX-E1108), and a C entry beside a CX function main (CX-E1109 — two entries, one program). Inside a function body a _C{} main is still just a nested function nothing calls, and still says so (CX-W1019).

On the register VM the entry belongs to the host that starts the VM, so a C entry block is declined there by name — the program-scale version of the rule that already makes any _C{} function native-only.


8. Arrays

Two Storage Kinds — Chosen at Compile Time, Not by You

array nums.i[10];        // all dims compile-time constant -> STATIC: raw C stack array, zero allocation
array a.i[n];            // n is a runtime variable        -> GROWN: cx_array on the GC heap

A static array is a plain C-stack buffer — no allocation, raw C semantics, no bounds check unless you ask for one. A grown array lives on the GC heap, can be resized, and its element reads/writes are bounds-checked by default in the lenient sense: an out-of-range read returns 0, a write is silently dropped ("still C"). Add #pragma checks on (aliases: check, checkbounds) to turn that into a hard runtime error instead.

arr and dim are both aliases for array. A constant dimension must be a literal or a single #define — [N*2] is rejected; precompute into one define.

array IS THE WHOLE STATEMENT — it declares AND it redimensions (v3.285.0)

One statement, one meaning, wherever it appears. array names an array of a size; writing it again for a name already declared in the same scope RESIZES that array. There is no second keyword for the second act:

function probe.v() {
    array data.i[3];
    data[0] = 1;
    array data.i[6];                           // the SAME array, resized
    printf("%d %d\n", data->count, data[0]);   // 6 1 -- grown, elements preserved
}

redim was the older spelling of that second line and is retired: it is refused at compile time with CX-E1048, naming array, on both backends. It was never a statement of its own — the parser aliased it onto this same production — so nothing it did has gone away, only the second name for doing it.

What the one statement means:

``cx array d.i[4]; d[0] = 1; d[3] = 7; array d.i[4] = {9}; // 9 0 0 0 -- the literal, then the type's zero ``

``cx array d.i[2]; d[0] = 4; if (ready) { array d.i[6]; // a NEW array, six long d[5] = 55; } printf("%d %d\n", d->count, d[0]); // 2 4 -- the outer array is untouched ``

A RE-STATEMENT AT THE SAME SCOPE still RESIZES, and the difference is the block, not the spelling: array d.i[2]; array d.i[6]; on consecutive lines is one array grown to six. (Until v3.292.0 the nested case resized too. The register VM had no block scope for any type — an inner-block int x shadowed natively and reused the slot on the VM — so shadowing arrays alone would have been a one-backend behaviour, which is never acceptable. Block scope landed for every type at once, and this clause came with it.)

``cx function grow.v(array a.i) { array a.i[8]; a[7] = 99; } function probe.v() { array d.i[3]; grow(d); printf("%d %d\n", d->count, d[7]); // 8 99 } ``

An array that is never re-stated, never passed to a user function and never appended to stays a raw C stack array with no allocation. That is the point of the distinction, and it is not widened by this rule.

A subscript that is one index short is a ROW (v3.270.0)

A declarator has N dimensions; an expression supplies K indices. K == N is an element; K < N is a row — a pointer to the remaining slice, exactly as in C.

char roman[13][3] = {"M\0", "CM\0", "D\0", /* ... */};
printf("%s", roman[i]);      // roman[i] is a `char *` -- a C string
printf("%d", roman[i][0]);   // roman[i][0] is a character

int grid[4][8];
int *row;
row = grid[2];               // a real row pointer
printf("%d", row[5]);        // == grid[2][5]

Depth here is arithmetic, not a special case: char cube[2][3][4] subscripted twice is still a char *, because it is still one index short.

Two limits worth knowing, both deliberate:

Grown Arrays Double as Lists

array nums.i[0];              // the [0] forces it heap-backed / grown
arrAdd(nums, 10); arrAdd(nums, 20);
print(nums[0]);                        // index
foreach nums { print(arrGet(nums)); }  // iterate

And a list or grown array doubles as an ordered map (PHP/Lua-style): mapPut(c, key, value) appends the value into the sequence and indexes it by string key; the keyed value is still visible to foreach and to ->count.

An element nobody wrote reads as the element type's zero

A cell that exists but has never been assigned answers 0, 0.0 or "" by the array's declared element type, on both backends and however the cell came to exist — declared, grown by arrAdd, or invented by a re-statement:

function first.s(array a.s) { return a[0]; }
function demo.v() {
    array ws.s[2];
    println("[" + ws[1] + "]");          // []
    println("[" + first(ws) + "]");      // []
}

It is the same sentence lists pad a sparse store with (§11) and the analogue of the real nulls a sparse JSON store pads with (§18) — one rule, stated once, true of the write path and the read path alike.

(Until v3.280.0 the register VM answered the integer 0 for the string case, because a cell's zero BITS are the int zero and nothing had said otherwise. Native answered "". The wrong value then corrupted the concatenation it landed in — the line above printed 0], having lost its own [ — which is why a split like this is silent-wrong rather than cosmetic. It was invisible until v3.279.0, because before that a passed array lost its initialiser and every element of a passed string array was empty.)

Universal Bulk Operations

One name, works over any sequence (fixed array, grown array, list — both backends, one implementation):

FunctionDoes
fill(c, v)set every element
sum(c) / minof(c) / maxof(c)int-preserving numeric reduce (any float element promotes the whole result)
avg(c) / average(c)always float
find(c, v)linear first-match → index or −1
search(c, v)binary search, sorted arrays
contains(c, v)membership → 0/1 (on a string, stays the substring test)
reverse(c)in-place

Also works on maps (reduces over the values; key membership is mapHas/mapContains). arrSum/arrMin/arrMax/arrAvg remain as accepted synonyms on arrays specifically.

On a json value (v3.220.0)

The reduces and the finders work over a json array's elements or an object's member values — the same members foreach walks — and the subject can be a handle or a subscript:

json d;
d = parseDoc("{\"units\":[12,40,7],\"skus\":[\"widget\",\"gizmo\"]}");
print(sum(d["units"]));                 // 59.000  -- float, see below
print(maxof(d["units"]));               // 40.000
print(find(d["units"], 7));             // 2       -- an index, so int
print(contains(d["skus"], "gizmo"));    // 1

Two differences from the sequence contract, both deliberate:

Before v3.220.0 none of these worked: natively every one was CX-E1051, and on the register VM every one answered silently — 0 for the reduces, −1 for find, 0 for contains, and a no-op for the mutators.


9. Structures

Definition and Access — dot, not backslash

Struct field access is C-style ., not the old PureBasic-style \:

struct Point { x.i; y.i; }
struct Pair  { a.i; b.f[3]; }     // inline array field

Point p1;
p1.x = 3;
p1.y = 4;

A struct cannot take a name the type system already owns (CX-E1131). The suffix letters (.i .c .b .w .l .s .f .d .v .p) and the type keywords (int, word, string, ...) are refused at the declaration, naming the collision -- a struct named P would make a later ix.P mean the native pointer type, never your struct, which is the silent class CX refuses. The same reasoning covers an index: sortndx ix.p by age is CX-E1132 at the by, because a native element type has no members (by over a native element names a comparator function, never a field).

A struct declaration is already zeroed — every field, including strings and inline arrays, reads back empty before you write to it. CX makes Go's choice here rather than C's, so there is no = {} to write and none to forget; if (p->next == 0) is a real test on a link you have not set yet.

Nested structs and arrays-of-structs work the same shape as always, with . throughout. Compounds are byref by rule as parameters (§7); a struct-typed return needs the .StructName suffix.

Declaring by type tag, and what assignment means

The declare-on-first-use suffix takes a struct type in the tag position, exactly as it takes .i or .s — and assignment between two struct variables COPIES, as it does in C (both true on both backends since v3.280.0):

struct Pt { x.i; y.i; }
function demo.v() {
    struct Pt p;
    p.x = 3;
    q.Pt = p;            // declares q as a Pt, and copies p into it
    p.x = 99;
    printf("%d %d\n", q.x, p.x);   // 3 99 -- q is its own object
}

(Before v3.280.0 the register VM refused the q.Pt = ... form outright — CX-E1038, reading Pt as a field name — and, for the explicit struct Pt q; q = p; form, it moved the buffer pointer instead of the bytes, so the two names silently became one object. Native did neither.)

Mistakes here are refused at the .cx line, identically on both backends:

you wroteyou get
p.zz where the struct has no zz (read or write)CX-E1020 — names the struct and the field
p = { 1, 2, 3 } into a 2-field structCX-E1021 — the value count must match the field count
p = { nosuch: 1 }CX-E1020 — a named entry must name a real field
struct NoSuch v; where no such struct is definedCX-E1060 — the type, not the construct, is missing
p.Point = {}; — the type-tag form given an empty literalCX-E1113 — the declaration is already zeroed; write struct Point p;

(CX-E1113 is new in v3.1.045.0, and it closes a divergence rather than adding a rule: addr.Address = {}; compiled natively as a zeroed declaration while the register VM refused it as a member access, so it was a live spelling with one backend and no entry in this manual. Because the declaration was already zeroed, the empty literal only ever named a guarantee that existed without it. A populated literal is untouched — p1.Point = {10, 20}; still declares and initialises, on both backends.)

(Before v3.223.0 the first three were native-only — on the register VM a spare value was silently dropped and p.zz was reported as a "native-only construct", which is true of neither. CX-E1060 is new in v3.223.0: it used to be gcc's "unknown type name" against generated C on native, and the same wrong "native-only" on the VM.)

jsonToStruct — filling a struct from JSON in one line

struct hero { int hp; float speed; string name; }
json d = parseDoc(fileRead("save.json"));
hero h;
h.speed = 1.0;                  // a default
jsonToStruct(d["hero"], h);   // fills h.hp / h.speed / h.name from doc["hero"]

This expands at compile time into the field-by-field assignments you'd otherwise hand-write — not a reflection-based runtime call. Field mapping is by exact (lowercased) name; a missing/null json key leaves that field untouched (set a default first); extra json keys are ignored; nested struct fields recurse; array/list/map fields are a compile-time error (fill those yourself). Returns the count of fields filled, so you can check completeness — but only as a statement or assignment target, not nested inside a larger expression.


10. Pointers and Function Pointers

x = 42;
ptr = &x;         // address-of
value = *ptr;      // dereference

A .p value is a slot handle, not a raw memory address — this is a real change from the PureBasic era's peek_i/poke_i/realptr/realaccess model, which assumed literal process memory. If your program does low-level pointer arithmetic or raw memory poking, verify the current equivalents against your build; that corner of the language changed shape with the VM redesign (§ARCHITECTURE) and isn't something this rewrite can responsibly guess at in detail.

Function pointers (fp, see §7) are the well-confirmed, current pointer-adjacent feature: address-of a function, call through the pointer, arrays of function pointers, and functions-as-values passed to higher-order functions.

C-spelled pointers, at any depth (v3.264.0)

Writing the C declarator gives you a genuine C pointer with the pointee you wrote, and the number of stars is a number, not a list of supported cases — **, *** and deeper all ride one implementation, at every position that declares a pointer:

double v;  double *p;  double **pp;  double ***ppp;
v = 7.0;  p = &v;  pp = &p;  ppp = &pp;
printf("%.3f\n", (***ppp));      // 7.000

Works as a function local, a module-scope global, a struct field, a parameter and a return type; on both backends for locals and globals. Pointer arithmetic strides by element, because the pointee is real.

Two limits, and each is a loud refusal rather than a wrong number:

Pointers to STRUCTS (v3.265.0)

The pointee can be one of your own struct types, in either spelling, and every way of reaching through it works:

struct Pt { x.i; y.i; }

Pt *find(Pt *first) { return first; }     // parameter and return type

int main() {
    Pt v;  Pt *p;  struct Pt *q;          // `T *` and `struct T *` are synonyms
    v.x = 7;  p = &v;  q = find(p);
    printf("%d %d %d %d\n", (*p).x, p->x, p.x, p[0].x);   // 7 7 7 7 — all four
    q->x = 9;                              // writes through to v
    return 0;
}

p.x is CX's own spelling and lowers to p->x; use whichever reads better. Depth composes here too — Pt **pp behaves exactly as double **pp does.

Self-referential structs (v3.266.0)

A struct field may be a pointer — including a pointer to the struct being declared. That is what C allows and for the same reason: a pointer is one word wide whatever it points at, so the size is known while the type is still incomplete.

struct node {
    val.i;
    node *next;               // points at the struct being declared
}                             // `struct node *next;` is the same declaration

node *push(node *head, int val) {      // C's declarator: see the note below
    node *n;
    n = malloc(16);
    n->val = val;  n->next = head;
    return n;
}

int main() {
    node *p;  int total;
    p = 0;
    p = push(p, 1);  p = push(p, 2);  p = push(p, 4);
    total = 0;
    while (p != 0) { total = total + p->val; p = p->next; }   // 7
    printf("%d\n", total);
    return 0;
}

With that one field come the linked list, the binary tree and the graph. Reaching through the field works in every spelling — n.next->val, (*n.next).val, n.next.val, n.next[0].val — and chains to any depth: a.next->next->next->val. The link is zero-initialised like any other field, so if (p->next == 0) is a real terminator test. Depth composes: node **p and node ****p as fields need no extra syntax.

Two structs may point at each other, and a field may name a struct defined later — no forward declaration is required:

struct vertex { id.i;  edge *first; }     // `edge` is defined below
struct edge   { to.i;  vertex *target;  edge *next; }

Worked example: Examples/407 self-referential structs.cx (list, tree, adjacency list).

Not yet: an array of pointers as a field (T *v[N], CX-E1081).

The .T* suffix spelling (v3.267.0)

A pointer type can also be written in CX's suffix form, at a function's return type and at a struct field. node *next; and next.node*; are the same declaration — two doors onto one fact — so you can mix them inside one struct.

struct item {
    val.i;
    next.item*;                        // same as `item *next;`
}

function last.item*(item *h) {         // same as `item *last(item *h)`
    while (h->next != 0) { h = h->next; }
    return h;
}

Depth is a number, not a set of cases: .item** and .item**** work for the same reason .item* does. .p remains the untyped pointer suffix (a pointer whose pointee nobody named); .T* is the typed one.

Everywhere else — a parameter, a local, a list/map/array element type — the suffix pointer is refused with CX-E1083, which says coming, not wrong. Use C's form at those positions: function get.i(node *p) has always worked.

Returning a pointer through a scalar is an error (v3.267.0)

function push.i(node *head) returning a node * is CX-E1084 at that line. An address is not a number CX converts for you, and before this the mismatch was your C compiler's to report, somewhere in generated code you never wrote. Declare the return as a pointer — .T* or node *push(...), both exact — or say you meant the address as an integer with an explicit cast.

Two things it deliberately leaves alone: an explicit cast, because an explicit conversion is obeyed exactly; and &fn, a function's address, which CX parks in an integer slot by design (fp h = &add1; is the same lowering at the assignment position).

A bare return &x; on a data object through a non-pointer return type is CX-E1084 too, since v3.267.1 — C refuses it and so does CX. It built and ran before that, so this is the one part of the floor that removes something rather than moving where it is reported. The fix is the same one C wants: a cast, or a pointer return type.

Struct pointers on the register VM (v3.268.0)

They work. A struct-pointer variable, a pointer field, a walk, a recursive tree and mutual reference all compile and run on the register VM and print what native prints — p->x and a.x are the same instruction there, because a struct pointer on that backend is the struct handle. Depth composes for free: a.next->next->next->val costs one field load per arrow.

Three things are still native-only, and each says why:

Worked example: Examples/408 suffix pointer types.cx runs byte-identically on both backends.

new and delete — nodes from CX's own heap (v3.269.0)

struct node { val.i; node *next; }

node *push(node *head, int v) {
    node *n;
    n = new node;           // one ZEROED node; the size comes from the shape
    n->val  = v;
    n->next = head;
    return n;
}
...
delete p;                   // give it back

new T allocates one zeroed instance of struct T and yields a T *. It is an expression, so it works anywhere a struct-pointer expression works — assignment, initialiser, argument, return. delete p; is a statement that releases one.

You never write a byte count. The size comes from the shape, so adding a field to node cannot leave a stale malloc(16) behind it somewhere else in the file. The node also carries its own identity: its string fields are released with it, and under #pragma checks on the shutdown leak report names the struct type that leaked rather than a byte count:

LEAK #2  type=struct  size=16  rc=1  node

A node forgotten anywhere in the program comes back by name — something a raw malloc block can never do, because the libc heap keeps no books CX can read.

delete sets its operand to 0, which C does not do. That is deliberate: it makes if (p != 0) mean what it reads, turns a use-after-delete into a named null-deref instead of a silent wrong value, and keeps the two backends byte-identical on a program that touches the pointer afterwards. And because of it, delete on a pointer that is already 0 is a no-op — like C's free(NULL) — so a "delete if held" needs no guard.

One spelling caution: delete is a contextual keyword — the sortndx element verb delete(ranks, 3) and a user function named delete both still work, separated by one token of lookahead. The cost of that rule: **delete (p); — with parentheses — reads as a call to a function named delete**, not as a delete of a parenthesised operand. Write delete p;.

Both backends. Examples/407 self-referential structs.cx — a linked list, a binary search tree and a mutually-referential graph, all built with new — runs identically native and on the register VM.

Checked mode: what a null pointer does

Checks are off by default on both backends; #pragma checks on enables them on both; a program behaves identically native or on the register VM either way.

That single sentence is the whole rule, and it has a cost side and a safety side. By default a field access through a null struct pointer does what C does — it dereferences null, and the OS ends the program. Nothing is spent asking whether the pointer was null, on either backend, because the default is C's own bargain: speed first, and you were the one who wrote p->x.

Turn checks on and the same access is a named error at the .cx line instead:

#pragma checks on
struct node { val.i; node *next; }
node *p;
printf("%d\n", p->val);      // CX-E5024: p: null pointer dereference

The register VM says the same thing in the same words — null pointer dereference — under its own code (CX-E5021 for a read, CX-E5022 for a write), because it can name the access but not the variable. Grep the phrase, not the code, and you will find it on either backend.

The pragma is decided at COMPILE time, so a program built without it carries no guard at all — the register VM's bytecode is byte-for-byte what it would have been if the feature did not exist. One deliberate exception on both backends: the guard covers a named pointer (p->x, p.x, p->arr[i]), not a chained intermediate (a->next->val where a->next is null), because a guard on the chain could only name the wrong thing in its own message.

The same pragma is what turns grown-array bounds violations into hard errors (§"Static and grown arrays") — one switch, one meaning, both backends. It is also what makes access through a deleted container a named error rather than the typed zero it answers by default (CX-E5044; §containerDelete).

Not built yet: new T[n] (an array of nodes) — filed, with delete[]'s design question attached.

Two allocators, one rule

CX has exactly two ways to get heap memory, and they are different heaps:

new T / delete p;malloc() / free()
HeapCX's GC — every allocation carries a typed chunk headerraw libc — the C escape hatch, obeyed exactly as C
Backendsboth — native and the register VMnative only (CX-E1013 on the VM)
Contentszeroeduninitialised, as C's malloc
Sizefrom the struct's shape — you never write a byte countyou write the byte count, as in C
Leak reportvisible by type under #pragma checks on (LEAK #2 type=struct … node)invisible — the GC keeps no books on it
Node churnfree-list reuse; a 400k-cycle alloc/free benchmark ran ~4.4× faster than malloc/free (measured on the Windows dev box — advisory, not a gate)plain libc speed

The rule is C++'s own: delete what you new, free what you malloc, and never cross. Crossing hands one heap's pointer to the other allocator — undefined behaviour, exactly as passing a new[] pointer to free is in C++. CX does not police the pairing at run time.

Guidance: CX code uses new / delete. Reach for malloc only when C interop requires it — a buffer a C library will own, realloc, or free — and then release it the C way.

A cast's pointee is the type you wrote (v3.270.0)

(const double *)p casts to a pointer to double, and dereferencing it reads a double:

int cmp_dbl(const void *a, const void *b) {
    double x = *(const double*)a - *(const double*)b;
    return x < 0 ? -1 : x > 0;
}

Before v3.270.0 the base type was thrown away at the * and every non-struct pointer cast was spelled long long, so that subtraction was comparing IEEE-754 bit patterns. char *, int *, double *, void * and your own struct types all now cast to what they say. float * becomes double * and long * becomes long long * — the same deliberate collapses a declaration makes, because CX has one 8-byte float and one 8-byte integer.

Handing a CX function to a C callee (v3.270.0)

#include <stdlib.h>
int cmp_dbl(const void *a, const void *b) {
    double x = *(const double *)a;
    double y = *(const double *)b;
    return (x > y) - (x < y);                    // the C idiom; CX compiles it as written
}

qsort(values, n, sizeof(double), cmp_dbl);      // just works
bsearch(key, values, n, sizeof(double), cmp_dbl);

A CX function's C signature returns long long, and qsort wants a comparator returning int with const-qualified parameters. CX emits a small adapter with the callee's exact declared shape and passes that — your function is untouched and stays callable directly from CX in the same program.

This covers qsort and bsearch. It is keyed on the callee, and a function of your own with that name shadows it, so a program that defines its own bsearch is unaffected. Declaring a C function-pointer type (int (*f)(int, int);) is still CX-E0005 — use fp (§7) for CX's own function values.


11. Collections — Lists, Maps, and Queues

list todo.s;             // grows
map scores.i;            // hash map, .i value type

listAdd(todo, "buy milk");
mapPut(scores, "alice", 100);

Byref by default, same rule as arrays/structs (§7): a container variable is a heap handle, so passing it to a function shares the same container, and reassigning one container variable to another is dropped with a warning rather than silently aliasing two owners.

A DECLARED CONTAINER IS LIVE (v3.285.0)

Defining a container initialises it. A declaration is a constructor: after list xs.i; the list exists and is empty — xs->count answers 0, a verb works, a foreach runs zero times. Think of it as an array of 0 elements. The same holds for map, sortndx and array, at module scope and inside a function, on both backends. You never need a separate create call, and there is no window in which a declared container is a handle to nothing:

function demo.v() {
    list xs.i;
    map m.i;
    array a.i[3];
    printf("%d %d %d\n", xs->count, m->count, a->count);   // 0 0 3
}

json is the family this does not describe, and on purpose: a declared json is live, but its root kind is undecided until the first subscript picks it — a string key makes an object, an int index makes an array. That is the next section. (An XML document is a json document too, so it follows the same rule.)

Two families qualify that, and both for reasons of their own:

Initialising a container at its declaration (v3.285.0)

Every container family takes a brace initialiser in the shape its own elements have. The declaration is the constructor, so the literal is part of the construction rather than a series of statements after it:

list  xs.i    = {1, 2, 3};              // positional elements
list  names.s = {"ada", "grace"};       // .i / .f / .s
map   m.i     = {alpha: 1, beta: 2};    // KEYED pairs
sortndx s.i   = {3, 1, 2};              // sorted AS IT BUILDS -> reads 1 2 3
array nums.i[5] = {10, 20, 30};         // padded with the element type's zero

Three things worth knowing:

A subscript store BUILDS the slot it needs (v3.279.0). xs[n] = v grows the list to n + 1 if it has to, padding anything in between with the element type's zero — 0, 0.0 or "". It is the list's version of what a sparse json store does with real nulls (§18), and it means a list can be filled by index without a priming loop:

function demo.v() {
    list xs.i;
    xs[0] = 11;               // builds slot 0
    xs[3] = 5;                // grows to 4; slots 1 and 2 are 0
    printf("%d %d %d\n", xs[0], xs[1], xs->count);   // 11 0 4
}

Reads never grow. xs[99] on a four-element list answers the element type's zero and leaves ->count at 4 — probing a list can never mutate it. Before v3.279.0 the store was the silent one: the slot did not exist, the write vanished, and xs[0] answered 0 without a word.

FunctionReturnsDoes
l->countintelement count (the retired listSize(l) — §11)
listAdd(l, v)voidappend
listGet(l) / listGet(l, i)variesvalue at the cursor, or at index i (0-based)
listGetAt(l, i)variesvalue at index i, never the cursor
listSet(l, v)voidset current value
listFirst(l) / listLast(l)intmove cursor
listNext(l)intadvance
listSort(l [, desc])voidsort
listDelete(l)voiddelete current
listClear(l)voidremove all

mapPut/mapGet/mapDelete/mapClear/mapReset/mapNext/mapKey/mapValue mirror the shape you'd expect.

A read that finds nothing answers the element type's zero

This is one rule across every container, and it never mutates. A mapGet for a key that is not there, a listGet past the end, an xs[99] on a four-element list: each answers the element type's zero — 0, 0.0 or "" — and leaves the container exactly as it was. Both backends, and a declared map m.i; answers the same way the explicit-handle mapCreate() form does.

function demo.v() {
    map totals.f;
    mapPut(totals, "services", 45.5);
    printf("%.2f %.2f\n", mapGet(totals, "services"), mapGet(totals, "hardware"));  // 45.50 0.00
    printf("%d\n", totals->count);                    // 1 — the missing read did not insert
    if (!mapHas(totals, "hardware")) { mapPut(totals, "hardware", 0.0); }
    printf("%d %d\n", totals->count, mapHas(totals, "hardware"));   // 2 1
}

The consequence worth naming: a stored zero and a missing key read alike. When that difference matters — a counter that legitimately holds 0, a name whose value is the empty string — ask mapHas(m, k) (mapHasKey/mapContains are the same function), which tests for the KEY and is the only thing that can tell the two apart. If a missing key is simply not expected, the zero is usually the answer you wanted anyway and no guard is needed.

The cursor API survived the port intact. listFirst/listLast position it, listNext/listPrev step it and answer 1 while they land on an element and 0 at either end, listReset rewinds to before the first, and listSelect(l, i) puts it on element i — all measured on both backends. (This paragraph used to say listPrev, listReset and listSelect-by-index were "not currently present". They are, and they were: the claim was a survivor of the port that nobody had re-checked, found by a stranger reading this page in 2026-08.)

sortndx (a separate sortable-index container) also exists for stable multi-key sorts by comparator or by struct field.

Queues — a bounded ring whose bound can be data

queue samples.f;                    // .i / .f / .s -- or .StructName
queueInit(samples, cfg["window"], 1);   // capacity from JSON; 1 = ROLL
queuePush(samples, 120.0);
while (samples->count > 0) { print(queueTake(samples), "\n"); }

An array's dimension must be a literal (§8), so the moment a limit belongs in a config file a fixed array stops being the right shape — and you end up hand-rolling head/tail indices around it. A queue takes its capacity at queueInit from any runtime expression, so the bound can come straight out of JSON, and the wrap arithmetic lives in the container instead of in your loop. The ring is allocated once at init and never grows: no per-push allocation, which is the point over list for a per-frame fact queue.

One type, no mode flag — the ops decide the behaviour:

FunctionReturnsDoes
queueInit(q, cap, policy)voidset capacity + overflow policy. cap is any runtime expression
queuePush(q, v)voidadd an element
queueTake(q)elementremove + return the oldest → FIFO, a work queue
queuePop(q)elementremove + return the newest → LIFO, a stack
q->countintlive element count right now (the retired queueCount(q) — §11)
q->capintcapacity as set at init (the retired queueMax(q) — §11)
queueClear(q)voiddrop the contents, keep capacity + policy

Using both ends of one queue makes it a deque — that's allowed on purpose. You write one name for each verb; the declaration's .suffix picks the typed entry point at compile time, so the runtime never guesses the element type.

A queue has no subscript, and neither does a sortndx (v3.281.0). Position is the container's own business — a queue hands back the oldest or the newest, a sortndx keeps itself in key order — so there is no slot for you to name, and q[0] is refused at your line on both backends:

queue q.i;
queueInit(q, 4, 0);
q[0] = 1;      // CX-E1090: the queue type has no subscript … use queuePush(q, v) to add,
               //           queueTake(q) for the oldest, q->count

The other four families — list, map, array, json — do subscript, read and write. Which is which is one column in the compiler's container-family table, so a family answers this question by existing, not by being remembered.

Every family's key kind

(v3.287.0) That same column says more than yes-or-no: it says what kind of key the subscript takes. There are exactly two kinds, because there are exactly two ways to name a slot — a position (an int) and a key (a string):

You wroteIt takesBecause
xs[i] on a listan int positiona list is a sequence
a[i] on an arrayan int positionsame — and the same for a passed (dynamic) array
m["name"] on a mapa string keya map is a lookup
d[...] on a jsoneitherthe first subscript decides whether the root is an object or an array (§json)
s[i] on a stringan int positionthe i-th character
p[i] on a pointeran int positionplain C: p[i] is *(p + i)
q[...], s[...] on a queue / sortndxnothingthere is no slot to name — above
d[...] on an image, memfile, …nothinga resource handle is read by verb
n[...] on an int or a floatnothingone number has no slots — below

Offer the wrong kind and CX says so at your line, on both backends, before anything reaches the C compiler — it names the family, what that family takes, what you gave it, and where to go instead:

list xs.i;
listAdd(xs, 1);
xs["k"] = 9;   // CX-E1093: list subscripts take an int position, and this one was
               //           given a string key … Use a map
map m.i;
m["a"] = 1;
m[0] = 9;      // CX-E1093: map subscripts take a string key, and this one was
               //           given an int position … Use a list or an array
image d;
d["k"] = 9;    // CX-E1090: the image type has no subscript … an image is a picture,
               //           addressed by COORDINATE — imagegetpixel(i, x, y)

And a plain number is not a container either (v3.290.0). int and float hold one value, so there is no slot to name and n[0] is refused at your line on both backends — exactly like a queue, for exactly the same reason:

int n;
n = 5;
n[0] = 1;      // CX-E1090: the int type has no subscript … an int is ONE number,
               //           so there is no slot to name — declare a `list xs.i`
               //           for positions or a `map m.i` for string keys

Before v3.290.0 this was the one shape that got past the refusal, and what it cost was the whole point of having one: natively it reached gcc, which answered "subscripted value is neither array nor pointer nor vector" at a line in generated C — a file you never wrote; on the register VM it reached CX-E1038, whose advice was "build native", which walked you straight into that gcc line. A pointer is the case to keep separate in your head: p[0] is valid — it means *p, plain C, and CX keeps all of C.

json has no refusal here, and that is the ruling rather than an omission. It is the one family that takes either kind, because the first subscript is what decides whether its root is an object or an array — so neither key can be wrong at compile time. Ask a json root for the other kind after it has been decided and you get the runtime refusal instead, which writes nothing:

json d = parseDoc("[1]");
d["k"] = 7;            // refused at run time, nothing written — a write creates, it never converts
println(exportDoc(d));  // [1]

The key is judged when the compiler can prove its type: a literal, or a name whose declarations agree (k.s = "b"; m[k] = 1; is fine, xs[k] is not). A key hidden behind a call or an expression is left alone — CX would have to guess, and it would rather say nothing than guess wrong.

Overflow is a policy, chosen per queue, and it's data too:

Every failure is loud — there are no sentinel returns anywhere in this API. Use before queueInit, capacity <= 0, a push onto a full REJECT queue, and a take/pop on an empty queue all stop with an error naming the queue or the op. That's why q->count exists as the drain-loop guard: no caller ever has to probe by failing. Re-initialising with the same capacity and policy is a clear-and-reuse; re-initialising with a different one is an error (a bound that changed between calls means the program read a different config key).

Struct elements (queue window.Sample;) queue several fields as one element, so they can't drift apart the way a queue-per-field pushed in lockstep can:

queue window.Sample;
struct Sample s; struct Sample out;
queueInit(window, 3, 1);
s.ms = 120.0; s.code = 200; queuePush(window, s);   // copies the record IN
queueTake(window, out);                             // copies the OLDEST record OUT

Note that take/pop take a second argument here. A struct is bytes, not a value a call can hand back, so it's copied into a variable you supply; writing the 1-arg form is a compile error, not a surprise at runtime. v1 restriction: a queued struct must hold plain numbers — a string field is refused at the declaration, because the element moves as raw bytes and a copied string handle would have no owner. Queue an id and keep the text in a map alongside.

Queues are byref like the other containers, and for the simplest possible reason: a queue is an opaque handle, so passing it to a function passes the queue itself.

When to reach for it: use a queue when a dropped item is a lost fact. Keep an explicit bound plus a loud counter instead when a drop is a degraded result the program is meant to survive.

Worked example: Examples/507 test queues.cx.


Handle metadata — -> asks about the handle, [...] asks about the data

A container variable has two selves. There is the data it holds, which you reach with a subscript, and there is the handle itself — how many items, how big, whether it is still alive. For a long time both were spelled the same way, so the compiler had to guess which you meant. That guess is the shared root of several long-standing footguns: jsonvar == 0 not meaning what it looks like, an int sink silently reading an object as 0, a queue whose capacity you passed in and could never ask for again.

-> gives the second self its own spelling:

json  config;  queue jobs.s;  list scores.i;

config->doc->count       // 3   — members / elements / children
jobs->cap                // 8   — capacity (queue only)
config->doc->valid       // 1   — still live? 0 after a free
config->doc->type        // 5   — json node kind (runtime); element type elsewhere
config->doc->id          // the raw handle integer, if you need it

while (i < ports->count) { ... }          // the everyday use
if (jobs->count * 2 >= jobs->cap) { ... } // capacity is visible now

config["servers"]->doc->count             // a SUBNODE answers too
config["servers"]->doc->type              // 4 if that member is an array
config["a"]["b"]["c"]->doc->count         // chains as far as you like

A DOCUMENT puts its own facts one level down, under doc. Every other family writes them straight on the arrow — jobs->count, scores->count — and that is the difference to hold on to: a list, map, queue or array has no member names of its own, so nothing can collide there. A document does. Its top-level arrow names belong to the program that wrote the document, so a document carrying a member called count reads that member and nothing else, and the number of members is ->doc->count. See A document's own facts below.

A subscript can be the base, not just a variable. Asking about a subnode — "how many entries are under servers?", "is this member an array or a string?" — is one of the most common metadata questions there is, so doc["key"]->count reads directly rather than forcing you to park the node in a temporary variable first. It chains, and it costs the same as the bare-variable form.

FieldMeansAvailable on
->countelements / members held right nowqueue, list, map, array, sortndx — on a document, ->doc->count
->capcapacity chosen at creationqueue
->typeelement type, a compile-time constantqueue, list, map, array, sortndx — on a document, ->doc->type (its node kind, at runtime)
->validhandle livenessqueue, list, map, array, sortndx — on a document, ->doc->valid (see the note below)
->idthe raw handle integerqueue, list, map, array, sortndx — on a document, ->doc->id

It costs nothing. Every one of these resolves at compile time from the variable's declared type. No struct exists at runtime, nothing carries a tag, nothing is dispatched: config->count becomes exactly the call you would have written by hand, and scores->type folds to a constant before the program runs.

A document's own facts — ->doc->, and doc(h, …)

A json document is the one family whose members have names the program chose, and those names are not drawn from a reserved list. speed, first_name, total — and count, type, valid, id, format, key, options, which are also the words a handle answers about itself. Something has to give, and what gives is the handle's side:

_json order { { "count": 7, "total": 19.99 } }

order->count             // 7  — THE MEMBER. The document carries it; it is yours.
order->doc->count        // 2  — the document's own fact: it has two members.
doc(order, count)        // 2  — the same thing, spelled as a call.

Every top-level arrow name on a document belongs to the members. The document's own facts live one level down, under doc, where they can grow later without ever colliding with a name a program might choose.

doc(h, x) is the same route spelled as a call, not a second mechanism: the compiler rewrites it into h->doc->x while it is reading the line, so the two cannot drift apart, cannot disagree between the backends, and cost nothing extra to compile. The second argument is a name, not an expression — doc(d, count) never looks for a variable called count.

doc is a reserved word. You cannot name a variable doc. A document member called doc is still there and still reachable, by bracket: h["doc"].

Porting: d->count reads as d->doc->count, or doc(d, count) — and the same for type, valid, id, format, key, options. The old spelling is not an alias and not a deprecation; it is a compile error (CX-E1135) that names the replacement at the site. If the document has a member of that name, the old spelling was reading the wrong number and the new one reads the member.

Read-only. These are derived facts about a handle, not storage. jobs->count = 0 is a compile error (CX-E1046) rather than a silently dropped write — the only thing that assignment could mean is "discard the elements", and that already has its own name, queueClear. An unknown field is CX-E1045, and it names the field you probably meant.

On a TYPED handle, -> is the only spelling. queueCount(q), queueMax(q), jsonSize(d), jsonType(d), listSize(xs), mapSize(m), arrSize(a), arrCount(a), and the container-typed resolution of len/length/size/count are compile errors (CX-E1110) that name the field replacing them. This is a type rule, not a retirement: the builtins are live, and on a handle held in a plain int — the older idiom, where -> cannot resolve because an int declares nothing — the call form is the only spelling and still answers. Two things this removed, beyond the duplication:

wasnow
listSize(xs) mapSize(m) arrSize(a) arrCount(a)xs->count
len(xs) length(xs) size(xs) count(xs) (container arg)xs->count
queueCount(q) / queueMax(q)q->count / q->cap
jsonSize(d) / jsonCount(d)d->doc->count (see the note below)
jsonType(d) / jsonValid(d) / queueValid(q)d->doc->type / d->doc->valid / q->valid

Two things are deliberately kept, because each answers a question -> cannot:

A third used to be listed here: the call form on a handle held in an untyped or int variable. It is gone as of 3.3.002.0 — a container is declared as what it is, and the call form on an int is now CX-E1028, the same answer a name nobody ever bound gets. One row survives the change for a measured reason: queueValid(n) still answers, because it is the only way to ask whether a raw handle number that no variable names is still live, and ->valid resolves from a declared type.

If you are tempted to "upgrade" old code by changing a declaration, change the HOLDER, not the BASE. Once the base is declared json the compiler knows jsonMember(base, k) yields a node, and a node assigned to an int becomes its value — int gives 30, float gives 1.5, string gives "Alice", json keeps the handle. So a line that used to store a node handle in an int silently starts storing the value instead. The safe edit is one added line: json m = jsonMember(base, k); and then m->doc->type.

One thing worth knowing, because it is not what you might assume:

-> on a struct pointer is unaffected — this is a resolution step keyed on the declared type, and a struct is not a container family.

This holds inside #pragma rules c rule text too, since v3.187.0. Rule text was the one exemption for a single release — the arrow was unwritable there — and it no longer is; see §28.

Removed VERBS: one name per operation

The section above is a type rule — a live builtin, refused only where the declared type offers the field form. This one is not. The same principle (one spelling per operation) removed a set of verbs outright, and the compiler keeps no memory of them.

These spellings were removed. The compiler does not keep a list of them: calling one is an ordinary unknown-function error (CX-E1028), exactly as if the name had never existed. That is deliberate — a language that memorialises every name it ever had teaches its own history instead of itself.

listRemove · mapRemove · mapWalkNext · mapFirst · remove · sortndxRemove · prt · prtl · jsonCreate · jsonCreateArr · jsonStringify · jsonStringifyPretty · insertCode · deleteCode

Write delete for removing one element from any container, mapNext to walk a map (mapReset first if you need to start at the beginning), jsonExport to serialise, prts to print without a newline, and a bare json d; to make a document — its first subscript decides whether the root is an object or an array. To install a marker's code use setCode then replaceCode; to drop one, delMarker.

redim is the one old spelling the compiler still recognises, and it answers CX-E1048 naming array as its replacement. It is kept for a stated reason: BASIC-heritage programmers genuinely arrive typing it, so the pointer has an audience. One statement now declares and redimensions — array d.i[6] on an existing d resizes it in place, keeping the elements that still fit.

Delete is the verb for removing one element, in every container family. listRemove and mapRemove were aliases onto that same operation; mapWalkNext was an alias onto the same operation as mapNext.

The two rows that were never real: prt and prtl. Unlike every other retirement here, these do not rename a working operation onto its canonical spelling -- there was no working operation. prt expanded natively to ((void)(x)), a silent no-op: it compiled, ran, printed nothing and exited 0, while the register VM refused the name outright at every arity. prtl expanded to exactly the same code as prts (one implementation under two names) and was likewise refused by the VM. Neither appears in the v2.0.28 builtin table, neither had a Builtins Reference entry, and no program in the tracked tree called either one. The no-newline print family is, and always was, prts / prti / prtf / prtc.

remove was the sortndx element verb, never CX's file-delete. The bare name did not reach C's remove() before this retirement either — it was a compiler verb gated on the operand being a sortndx, so nothing that deleted a file has changed. Use fdelete(<path>) for a file; C's remove() is still reachable through _C{}. The compiler repeats that note at the site rather than answering a file-delete with advice about a container.

mapFirst is the one that is not a pure rename. It did two things: rewound the cursor and advanced onto the first entry. So it is replaced by the pair; mapNext alone resumes from wherever the cursor already was. The compiler carries that caveat at the site.

Two of these were only ever spellable on one backend, which is how they survived long enough to need retiring: listRemove and mapWalkNext resolved in the native emitter and nowhere else, so the register VM had always refused them.

containerDelete — releasing a container early (v3.239.0)

listClear and mapClear empty a container and keep it. Until v3.239.0 there was no way to say the other thing — "I am done with this container" — at all, in any spelling. containerDelete(c) is that verb, and it is the only one:

list rows.s;
listAdd(rows, "a");
// ... use it ...
containerDelete(rows);      // release now, not at the end of the block

Containers die like strings — by reference count. containerDelete is not a free(); it is an early scope exit for one reference. It releases your reference and nothing else, and the memory is reclaimed when the last holder releases it — which, for the ordinary case of a single holder, is immediately. What it buys you over just letting the block end is timing: a large container released at the point you finish with it rather than dozens of lines later.

It works on every container family — list, map, dynamic array, sortndx, queue, json — and the compiler picks the right release from the declared type, so there is one verb to remember rather than six.

After the call the name is dead. Using it again is a compile error at your own line, on both backends:

containerDelete(rows);
listAdd(rows, "b");         // CX-E1068: 'rows' was deleted on line 4

"Using it" includes SUBSCRIPTING it, on either side of the = (v3.283.0). A verb call, a ->count read and a foreach were caught from the start; a subscript was not, so containerDelete(v); v[0] = 9; compiled — and natively ran, writing through a released reference. Every subscript position now asks the same question and gets the same code, in every family, and wherever the subscript stands:

list v.i;
list w.i;
listAdd(v, 1);
listAdd(w, 5);
containerDelete(v);
v[0] = 9;                   // CX-E1068 -- a store
printf("%d", v[0]);         // CX-E1068 -- a read
printf("%d", v[0] + 1);     //           -- inside an expression
printf("%d", w[v[0]]);      //           -- as someone else's index
v[0] += 1;                  //           -- a compound assign

A fixed array is the one exception, and for a reason that is not an exemption: containerDelete on a fixed array is itself refused (CX-E1069 below), so the name was never validly dead and the subscript after it is fine.

This is deliberately strict, and strict in the direction that cannot bite you at runtime. A delete in one branch of an if kills the name in the other branch too, and a delete anywhere in a loop body kills it for the whole body — including lines above the delete, because the next iteration reaches them with the container already released. Some programs that would have run are refused; the alternative is a container that silently does nothing, which is worse.

Three things are refused outright rather than quietly accepted:

you wrotewhy it is refused
containerDelete(n) on an int, a string, or a fixed arraynone of these holds a reference. A fixed array is a raw C array — it has no header and no refcountCX-E1069
containerDelete(xs) where xs is a parametera container parameter is borrowed: the caller owns the reference. Delete it where it was declaredCX-E1070
using the name after the deletesee aboveCX-E1068

The deleted variable holds null, and access through it is DEFINED (v3.284.0). containerDelete nulls its operand, the same thing delete does to a struct pointer — so what happens if a read or a write does reach that null handle is not left to whatever the freed memory happens to say. By default a read answers the element type's zero (0, 0.00, "") and a write drops, touching nothing; under #pragma checks on the same access is CX-E5044: access through a deleted container instead, on both backends.

You mostly cannot get there, because the compile error above is what a program actually meets. The one position it cannot see is a delete and a use in different function bodies — a file-scope container deleted at file scope and read inside a function, or the reverse — and that is the position this promise is for:

list v.i;
function use.v() { printf("R=%d\n", v[0]); }   // 0 by default; CX-E5044 under checks
listAdd(v, 1);
containerDelete(v);
use();

Two notes, because both are the kind of thing you would otherwise have to find out by experiment. The write drops rather than rebuilding the container: construction on demand is what a declaration promises, and a store through a handle its owner released is not a declaration. And a json handle is the one family where checked mode stays quiet here — a released json reads back as the same handle 0 a never-built one has, so refusing would refuse subscript-birth (d["k"] = 1 on a fresh json, which is ruled to work); it answers 0 on both backends, identically, as before.

There is no force-free spelling, and there will not be one. A verb that frees a shared container regardless of who else is holding it would hand every other holder a dangling pointer — and since container parameters are byref by default (§11), a container passed to a function has more than one holder by construction. Strings do not have such an escape hatch either; containers follow strings, which is the whole design.

12. Strings

UTF-8, small-string-optimized up to 12 bytes inline (widened from 7 bytes), GC-managed handles beyond that. Concatenation in a loop is amortized O(1) (§5).

Value-isolated rather than immutable, and the distinction only started mattering when s[i] = v arrived (see Characters below). Two names never share observable state: assigning copies, and a write through one name can never be seen through another. What changed is that a single character can now be written in place instead of only by rebuilding the whole string — the guarantee held, the wording had to get more precise.

FunctionDoes
s->len / strlen(s)length — the arrow is the noun, strlen() is a macro over it
left(s,n) / right(s,n) / substr(s, start[, len])substrings — substr is 0-indexed, and a negative start counts from the end (substr(s, -3) is the last three)
trim/ltrim/rtrimwhitespace trim
tolower/toupper/capitalizecase
strstr(hay, needle[, start])substring search — 0-based, and -1 when absent, so >= 0 is the test and if (strstr(...)) is the bug
replacestring/removestring/countstring/reversestring/insertstringmutation-by-copy — these return a new string rather than editing in place (s[i] = v is the one that edits)
contains(s, sub)substring test → 0/1 (also polymorphic over containers, §8)
startswith(s, prefix)0/1
stringsplit(s, sep, idx)one delimited field, 0-based — cheap, because nothing is allocated for the fields you did not ask for
split(s, sep, <list>)every field, into a list, returning the count
hex(n) / bin(n)int → hex/binary string

Characters — s[i], and the seam that used to be here

s[i] reads the byte at position i as an int, and s[i] = v writes one. Both are 0-indexed, both are C's answers, and both work identically on either backend.

s.s = "Hello";
printf("%d\n", s[0]);        // 72  -- 'H'
s[0] = 74;                   // s is now "Jello"

s[i] and substr() both count from 0, so they agree wherever they meet. This table used to be a warning — there was a seam here, and it caught everybody once. It is kept as the reassurance it became:

You writeYou get
s[0]the first character
substr(s, 0, 1)the first character
s[1]the second character — the same as substr(s, 1, 1)

This page used to say mid kept its BASIC heritage and that neither convention would change. That promise was withdrawn on 2026-09-05, and it is left named here rather than quietly deleted, because a reader who remembers it deserves the reason. The reason is that the two conventions could not both be right in one language: s[i] counted from 0 and mid counted from 1, so the same string was indexed two ways on adjacent lines, and every mixed expression carried a +1 or a -1 whose only job was to cross between them. mid is gone rather than remapped — substr(s, start[, len]) counts from 0, like every subscript, every array and every container in the language. Nothing in CX counts from 1 any more.

It is a byte, not a character. On ASCII those are the same and nothing surprises. On UTF-8 they are not: strlen("café") is 5, and s[3] is the first byte of a two-byte é rather than é itself. Writing one byte of a multi-byte character leaves invalid text. This is C's answer and CX gives the same one rather than a codepoint model it would then have to hide — if you need codepoints, work in substr(), which counts the same bytes but at least never splits one.

Out of range is quiet by default and loud on request. s[99] reads 0 and a write to it does nothing. Under #pragma checks on both raise CX-E5023 naming the index and the length, on both backends — the same gate the containers use.

A write can never be seen through another name. Strings share their bytes: b = a makes both read one buffer, and two variables set from the same literal share it too. So s[i] = v takes a private copy first if it has to:

a.s = "a long string well past sso";
b.s = "";
b = a;
a[0] = 88;
printf("%s\n", a);        // X long string well past sso
printf("%s\n", b);        // a long string well past sso  -- untouched

The copy is paid once per string, not once per write, so rewriting a string character by character in a loop is linear, not quadratic. Everything else about strings stays value-like: there is no way to observe someone else's write, which is the property the old one-word summary ("immutable") was really promising.

Literal text with no escapes — """…"""

A run of three or more quotes opens a raw literal. Everything up to a run of exactly the same width is the text, byte for byte: no \n, no \", no escape processing of any kind.

path.s = """C:\data\rules.json""";                 // backslashes survive
frag.s = """{"a": 1, "b": "two"}""";               // quotes survive
prompt.s = """
Explain the rule "hp <= 0" in plain English.
Keep it under 20 words.
""";

It yields an ordinary string — every builtin takes it exactly as before, and it concatenates with ordinary literals ("""raw""" "-plain"). There is no new string type, only a new way to write one.

The fence width is how the text contains its own delimiter. Because the closer must match the opener's width, a four-quote fence carries """ as ordinary text:

q.s = """"a 4-fence carries """ inside"""";

So there is no text this form cannot express — unlike a fixed delimiter, which always has one. If a fence is too narrow for its text the compiler refuses with CX-E1118 and names the width to widen; it never truncates silently.

Exactly one newline is stripped after the opening fence, so a block that starts on its own line does not carry a leading blank. Nothing else is touched — no indentation stripping, no trailing rule — because a whitespace rule you cannot run in your head is worse than a leading tab you can see.

Formatted Output — printf, sprintf

printf writes to stdout, sprintf returns the same formatting as a string, and fprintf(path, fmt, ...) appends it to a file. The conversions are libc's, and so is everything between the % and the conversion letter — the flags, the field width and the precision all work, which is what makes columnar output possible without a padding helper. In the table below · stands for one space, because HTML collapses real ones and the padding is the whole point:

SpecDoesExample output
%-12sleft-justify in a 12-column fieldSupport·····
%12sright-justify in 12·····Support
%6dpad a number to 6 columns····10
%06dpad it with zeros instead of spaces000010
%+dforce the sign on a positive number+42
%9.2f9 columns wide, 2 decimals····45.50
%.3struncate a string to 3 bytesSup
%x / %X / %o / %e / %chex, HEX, octal, exponent, character2a 2A 52 3.142e+00 A
function demo.v() {
    printf("|%-12s|%6d|%9.2f|\n", "Support", 10, 45.5);
    printf("|%-12s|%6d|%9.2f|\n", "Training", 2, 300.0);
    printf("%f  %.6f\n", 1.0 / 3.0, 1.0 / 3.0);        // 0.333  0.333333
}

Two things differ from C, and the second one surprises C programmers.

print/println are the format-free form: they take any printable value, render it the way CX renders it, and append a newline. printf appends nothing — put the \n in the format string, as above.

Converting — the cast is the conversion

A number becomes text, and text becomes a number, by a cast — in all six directions, with no conversion function in between. (string)n is the one most people do not know exists:

You writeYou get
(string)n / (string)fthe number as text — str(n) is the accepted macro over the same conversion; strf(f, decimals) when the format matters
(int)s / (float)sthe text as a number; text with no numeric reading is 0
(int)f / (float)nC's numeric casts — (int)3.7 truncates to 3, as in C
n.i = 42;
f.f = 2.5;
label.s = (string)n + " items at " + (string)f;    // 42 items at 2.500
total.i = (int)"17" + 1;                             // 18
ratio.f = (float)"3.75" * 2.0;                       // 7.500

The cast is explicit on purpose. n.i = "17" with no cast is refused (CX-E1074): a string has no lossless numeric reading, and CX never parses one behind your back. The conversion names this replaced are in the migration table below; §14 lists the same six directions beside the numeric ones.

Inserting — insertstring(s, t, pos)

pos is the 0-based position t is inserted at, and it answers exactly as substr's start does: a negative position counts from the end (-1 inserts before the last character), and out of range clamps rather than refusing — past the end appends, further back than the start prepends. One out-of-range rule for the whole family, on purpose.

printf("%s\n", insertstring("Helo", "l", 2));        // Hello
printf("%s\n", insertstring("Hello", "!", 99));       // Hello!   -- past the end clamps

Regular expressions — full strength, and a bad pattern refuses

The engine is PCRE2, and the four names are unchanged: regexMatch(s, pattern), regexExtract(s, pattern), regexReplace(s, pattern, repl), regexCount(s, pattern). What changed is what a pattern may say: grouping, alternation anywhere, backreferences, lookaround, non-greedy quantifiers and named groups all work — (ab)+, ^(cat|dog)$, (ab)\\1, foo(?=bar), <.+?> — where the previous engine silently did not match them. It runs in byte mode: . matches one byte, as it always did, so this is not a Unicode promise. And it is the native and register-VM story: the browser build carries no PCRE2 yet, so in the playground a pattern beyond the old engine does not match and a bad pattern still answers 0 in silence, as before this version — whether that dependency becomes acquirable there is an open ruling, and until it lands this page does not promise the browser what it promises the desktop.

A pattern the engine cannot compile stops the program, at runtime (a pattern is a string, and only known when the call runs), with the engine's own message and the byte offset into the pattern:

CX-E5067: regular expression 'a(' cannot be compiled -- missing closing
parenthesis (at offset 2). This is a fault in the PATTERN, not a subject
that failed to match; fix the pattern.

This matters because every one of the four names has a legitimate "nothing" answer — regexMatch says 0, regexExtract answers an empty array, regexReplace hands the subject back — so before this, a typo in a pattern was indistinguishable from a subject that did not match, on both backends. That was silence, not tolerance.

regexExtract answers a json document: an array of matches, each an array of groups, group 0 the whole match. A pattern that names a group answers an object per match instead, carrying both the numbers and the names — a json node is an array or an object and cannot be both, and dropping the numbers to gain the names would silently break d[0][1] for anyone whose pattern happens to name one group. Which form is in play is decided once, from the pattern, never per match. No match is an empty array, so a loop over the result runs zero times with no test in front of it; an unset group is an empty string, never a missing element, so group n is always at index n.

json d = regexExtract("2026-09-07 and 2027-01-31", "([0-9]+)-([0-9]+)-([0-9]+)");
printf("%d\n", d->count);            // 2
year.s = d[1][1];                     // a group into a declared .s; printf takes d[0][0] directly
printf("%s %s\n", year, d[0][0]);    // 2027 2026-09-07
json n = regexExtract("2026-09-07", "(?<year>[0-9]+)-(?<mon>[0-9]+)-(?<day>[0-9]+)");
printf("%s\n", exportDoc(n));       // [{"0":"2026-09-07","1":"2026","2":"09","3":"07","day":"07","mon":"09","year":"2026"}]

The export lists an object's members sorted, which is jsonExport's habit and not the regex engine's. Bind the result to a declared json before using -> on it: regexExtract(s, p)->count does not compile, because a handle not bound to a name has no type for the arrow to resolve from. A group itself can be used directly: a node reads as its text wherever the parameter's type is declared — printf, a concatenation, strstr — identically on both backends. What cannot see it is overload SELECTION from the argument's C type, which is how the assert* family binds: there, read the group into a declared .s first, or native refuses to compile and the register VM compares "".

regexReplace's replacement is literal. $1 in a replacement is a dollar and a one, not the first capture; PCRE2 would expand it by default, and CX keeps the contract it always had.

Migrating from the 1-based names

The old names are gone, not remapped: each answers CX-E1028, the same error a name that never existed gets, with no replacement text after it. So the arithmetic a migrator needs lives here, in one table, and nowhere else — the cheatsheet points at it rather than repeating it. It is also the one place in the documentation allowed to spell these names, and the doc gate holds it to that: a table whose first header reads exactly as this one does is set aside before any teaching surface is checked for a gone name, and a second such table anywhere is a failure.

old spelling (gone, CX-E1028)reads asnote
jsonExport(d)exportDoc(d)the document still writes its own format; only the name changed. exportDoc(d, CX_XML) is the new half: the same document through another lens
jsonExport(d, 1)exportDoc(d, CX_JSON, 1)the pretty flag moves to the THIRD slot, because the second is now the format. exportDoc(d, 1) compiles and is wrong — it asks for format 1, not for indentation
mid(s, p, n)substr(s, p - 1, n)0-based, and a negative start counts from the end — so where the old mid answered "" in silence for a 0, p - 1 of 0 is -1, the last character; check the arithmetic at every site rather than trusting the rename
mid(s, p)substr(s, p - 1)to the end
left(s, n)left(s, n), which is substr(s, 0, n)stays, as a macro — there was no base to migrate
right(s, n)right(s, n), which is substr(s, -n)stays, as a macro
instr(s, t) / findstring(s, t)strstr(s, t) + 1the returned number changes: absent was 0, and is now -1. A hit at the first character is 0, so if (strstr(s, t)) is the bug and >= 0 is the test. This is the one row where a program that still compiles answers differently
instr(s, t, start)strstr(s, t, start - 1) + 1the optional start is 0-based too
stringfield(s, n, sep)stringsplit(s, sep, n - 1)renamed, re-indexed at 0, and the separator now comes before the field number. An unconverted call to the old stringsplit alias of split is refused, never answered wrongly — with two arguments for its arity (CX-E1082), with a list as the third for its type
len(s) / length(s)s->lenthe length is a noun the handle holds; strlen(s) is the accepted C macro over it. These two answer CX-E1110, which names ->len, rather than the bare CX-E1028 — the arrow rule reaches them first
ucase(s) / lcase(s)toupper(s) / tolower(s)the same bytes back; these take and return a string, not a character code as C's do
val(s) / vali(s) / atoi(s)(int)sthe cast is the conversion; an implicit store without it still refuses (CX-E1074)
valf(s) / atof(s)(float)s
stri(n) / itoa(n) / ltoa(n)(string)nstr(n) stays as the accepted macro; strf(f, decimals) stays, because formatting is its own operation
insertstring(s, t, pos)insertstring(s, t, pos - 1)0-based; a negative position counts from the end; out of range clamps, as substr does

The family's position arguments, enumerated (every prototype carrying an integer was classified — the rest are counts, widths, values or modes): substr's start and strstr's start moved to 0-based; insertstring's position moved with them (negative from the end, clamped); stringsplit's field number counts from 0; and getc/setc's byte index was 0-based already and did not move. replacestring, removestring, countstring and split carry no position at all. left(s, n) and right(s, n) take a count, not a position, and are unchanged.


13. Math Functions

Unchanged from a C programmer's expectations: abs/fabs, min/max/fmin/fmax, sqrt, pow, mod, sign, the full trig set (sin/cos/tan/asin/acos/atan/atan2), hyperbolic (sinh/cosh/tanh), log/log10/exp, rounding (floor/ceil/round), clamp/lerp/remap, and random/randomseed.


14. Type Conversion

FunctionDoes
(string)n / (string)fnumber → string — the cast is the conversion; str(n) is the accepted macro over it, and strf(f, decimals) formats
(int)s / (float)sstring → int/float; text with no numeric reading is 0
(int)f / (float)nC's numeric casts: (int)3.7 truncates to 3

Int-to-float promotion in mixed expressions is automatic, same as always. The string↔number rows are the same casts §12 shows in use; the names they replaced are in §12's migration table, once.


15. Files, Filesystem, and Memfiles

Basic File I/O

fread(path) / fwrite(path, content[, mode]) — unchanged.

OS-Independent Filesystem Builtins (new)

list files.s;
n = dirlist("data", "*.cx", files);   // glob match, case-insensitive, skips . and ..
listSort(files);
foreach files { f = listGet(files); print(f, "  ", filesize("data/"+f), " bytes\n"); }

fileexists/direxists/filesize/makedir/removedir/renamefile/copyfile/movefile/dirlist round out a full cross-platform (dirent.h/sys/stat.h) filesystem surface that didn't exist in the PureBasic era.

Memfile — an In-Memory Byte-File

A memfile is one GC-allocated, page-grown byte buffer with a read/write cursor — a small integer handle, safe on invalid handles (returns 0/empty, never crashes). It's the substrate the compiler itself uses internally for #include expansion, now exposed as a language feature:

mf = mfnew();
mfputs(mf, "Hello World");
mfseek(mf, 5);
mfinsert(mf, " there");        // "Hello there World" — shifts the tail right
print(mftostr(mf), "\n");
mfsave(mf, "out.txt");
mffree(mf);

mf2 = mfload("out.txt");       // read a real file into a memfile
mfseek(mf2, mfsize(mf2));
mfinsertfile(mf2, "footer.txt"); // splice another file's bytes in — the in-memory #include

A document latches directly onto a memfile's bytes with no copy: parseDoc(h). It is the same door a string goes through — the compiler picks the row from the argument's type — so json, xml, csv and plain text all read from a memfile without a name of their own.

The third argument: an options document

parseDoc(src [, format [, options]]). The options are a JSON document, not a list of flags — nothing positional, no order to remember, and the same document the exporter reads back, which is what makes a sheet come out written the way it went in.

json a = parseDoc(s, CX_CSV, "{\"delim\":\"\\t\"}");        // a tab-separated sheet
json b = parseDoc(s, CX_CSV, "{\"header\":false}");           // rows are arrays
json c = parseDoc(s, CX_CSV, "{\"widths\":[8,12,2],\"names\":[\"first\",\"last\",\"age\"]}");
keymeans
delimone character. Absent, the header line decides: comma if it has one, else ;, else TAB.
quotethe quoting character; RFC 4180 doubling follows it. Default ".
headerfalse: there is no header line to read. Rows become arrays — unless widths names the columns.
widthsfixed-length columns. There is no delimiter at all; each row is cut at these widths and a row wider than they account for is refused by row number.
nameskeys for the widths columns. Without it they are c1…cN.
quotedhow the EXPORT quotes: "needed" (default), "all", "none".

An unknown key is refused by name, and the refusal reaches the caller — the whole parse answers an invalid handle, so ->valid catches a misspelt option. Silently ignoring {"delimiter":"\t"} would leave a program reading a comma sheet believing it had asked for tabs. Five keys are known-but-not-built and say so in those words rather than reading as typos: collapse, skip, columns, encoding, trim.

The lens that takes no options says so. json, xml and text accept none, so handing them any is refused rather than ignored.

->options — what the document was read with

d->options answers the options as text, and it reports what the parse settled on rather than what the caller wrote. A ; sheet read with no options at all still answers {"delim":";"} — which is exactly why it comes back a ; sheet when you export it. Read-only: reporting how a document was read is a smaller claim than changing how it exports.


16. Archives, Compression, and Networking

Compression

compress(s) -> s / decompress(s) -> s — raw zstd, binary-safe via length-prefixed strings.

Archives (libarchive: 7z/zip/tar + gz/bz2/zst/xz, read and write)

archiveCreate("game.dat", "zip", "assets");   // pack a directory
archiveExtract("game.dat", "out/");            // auto-detects format

Convenience one-liners: tarGz/tarBz2/tarXz/zipCreate/sevenZip(out, src).

Networking — the client side

HTTP/FTP/SFTP/email/SSH are shipped, built on a vendored static libcurl + libssh2 + TLS bundle (MinGW/MSVC; auto-detected against system libcurl on Linux). httpGet / httpPost / httpStatus / netError / netTimeout, ftpGet / ftpPut / ftpList / sftpGet / sftpPut, emailSend, sshExec / sshExecKey.

Async variants, email attachments/MIME, and IMAP/POP3 receive are open refinements rather than gaps in the base surface.

Networking — a CX program can HOST

Everything above is a client. Since v3.194.0 a CX program can also be the service. Eighteen builtins, one specification, both backends.

Sockets — netListen(port) binds loopback only; netListenAny(port) is reachable from other machines. Two names rather than one call with a flag, so exposing a port is visible at a glance. Then netAccept(srv), netConnect(host, port), netRead(conn), netWrite(conn, s), netClose(h). All return a handle or a negative error; netError() names the failure in words ("the port is already in use", not 10048).

netPoll(handle, timeout_ms) is the whole blocking model — 1 ready, 0 timed out, negative on error:

HTTP — on an accepted connection: httpRead(conn) parses the request, then httpMethod(conn), httpPath(conn), httpHeader(conn, name), httpBody(conn) read it, and httpRespond(conn, status, contentType, body) answers. httpStream and httpEvent cover chunked and server-sent events.

httpServeDir(conn, mount, dir) answers requests under mount from dir. It returns 0 having written nothing for a request outside the mount, which is what lets it sit in an ordinary if ladder beside routes you write by hand instead of taking the server over:

int srv = netListen(7440);
while (running == 1) {
    if (netPoll(srv, 200) == 1) {
        int conn = netAccept(srv);
        if (httpRead(conn) == 1) {
            if (httpPath(conn) == "/status") {
                httpRespond(conn, 200, "application/json", "{\"ok\":1}");
            } else if (httpServeDir(conn, "/", "www") == 0) {
                httpRespond(conn, 404, "text/plain", "no");
            }
        }
        netClose(conn);
    }
}

/ serves the directory's index.html; a directory without one is a 404, and there is never a directory listing. Containment is one decision, not a ladder of patches: decode once, resolve the path, check the resolved path against the resolved root. .., %2e%2e, ..\, ....//, %252e%252e, an absolute path, an embedded NUL and a symlink out of the tree all get the same 404 an absent file gets — identical on purpose, so a probe cannot learn which guess was right. Same behaviour on all three platforms.

Cross-origin access is off by default. httpAllowOrigin(origin) turns it on for a named origin. A server that always sends * is a liability, so CX does not send one for you — and the best answer is often not to need it: a page served from the same program shares its origin and has nothing to allow.

.wasm is served as application/wasm, which is the row that MIME table exists for — a browser will not stream-compile a module served as anything else.

See examples/network/ — 01_mesh.cx (one program, several machines), 02_chat.cx (hosts, joins, and joins from a browser tab), 03_serve.cx (serves a directory, including a page CX itself compiled to wasm).


17. Graphics: the scene is a document

A CX program does not call drawing functions. It writes fields on a json document, and the renderer draws what the document says.

media scene = parseDoc("{
  \"sky\":  { \"type\": \"rect\",   \"x\": 0,  \"y\": 0, \"w\": 100, \"h\": 70,
            \"color\": 2247253 },
  \"sun\":  { \"type\": \"circle\", \"x\": 72, \"y\": 6, \"w\": 16,
            \"color\": 16766720 },
  \"name\": { \"type\": \"text\",   \"x\": 4,  \"y\": 4, \"size\": 24,
            \"text\": \"good morning\", \"color\": 16777215 }
}");

gpu_screen_open(640, 480, "CX+AI");
mediaRun(scene, 0);            // until the window closes
gpu_screen_close();

That is the whole program. There is no frame loop in it, no draw call, and nothing to keep in the right order.

media is a json

A media declaration is a json declaration that also tells the compiler to link the renderer. Everything true of a json document is true of a scene: it is read by parseDoc, written with subscripts, and read back with the arrow.

scene["sun"]["x"] = 40;                    // the sun moves
println(str(scene->doc->count));           // how many elements
string out = exportDoc(scene);             // the whole scene as text

So a scene can be loaded from a file, generated, received over a network, edited while the program runs, and written back out — because it was never anything but a document.

Three verbs, and you usually need one

verbwhat it does
mediaRun(scene, frames)Hands the window to the renderer. 0 means "until it closes". Answers how many frames it ran.
mediaDraw(scene)Draws one frame. This is the step mediaRun loops over; a program with its own loop uses this.
mediaReport()What the last frame did, as a document: drawn, hidden, unknown, acted, stale, redrew, and types — the list of element types this renderer knows.

They are dual: they work natively and under the register VM (-P pcode=risc) from one implementation.

What an element says

Every element has a ->type and then whatever that type reads. Positions and sizes are percent of the window, per axis — x and w of its width, y and h of its height — so a scene written for one window works in another. Say "units": "absolute" for pixels.

types
flatrect circle text line triangle pixel image emitter
in the worldcube sphere cylinder plane line3d point3d grid model billboard emitter3d
acted on, not drawnwindow camera

The full field list is in the Builtins Reference under mediaDraw. The shape of it is worth knowing in advance, because one field usually replaces a family of function names:

Behaviour is text on the element

An element can carry ->rules: ordinary CX source, as a string, compiled when the renderer first meets it, with the element bound to e.

media scene = parseDoc("{
  \"ball\": { \"type\": \"circle\", \"x\": 18, \"y\": 18, \"w\": 8,
            \"color\": 16434758, \"vx\": 0.7, \"vy\": 0.66, \"tick\": true,
            \"rules\": \"{ float x = e[\\\"x\\\"]; float vx = e[\\\"vx\\\"];
                          x = x + vx;
                          if (x < 1)  { x = 1;  vx = 0 - vx; }
                          if (x > 91) { x = 91; vx = 0 - vx; }
                          e[\\\"x\\\"] = x; e[\\\"vx\\\"] = vx; return 1; }\" }
}");

The ball bounces. Nothing calls it: ->tick asks for the rule every frame, and without it the rule runs only when the element changed. Because the rule is text inside the document, a program can change how the ball bounces while it is bouncing.

Input is a field too

There is no event loop and no callback. The renderer writes what it saw onto the elements, and their own rules read it.

"go": { "type": "rect", "x": 10, "y": 10, "w": 20, "h": 8, "tick": true,
        "rules": "{ if (e[\"clicked\"]) { e[\"color\"] = 16711680; } return 1; }" }

⚠ When a scene has a window element the renderer reads the key queue, and a program calling gpu_key_pressed() beside it will find it already empty.

The world, and the camera that is an element

3D elements carry ->pos and ->dim in world units — never percent, because a world has no edges to be a percentage of. A camera element carries ->pos, ->target and ->fov, and the renderer draws the world inside it and everything flat on top, deciding which is which from the element's type.

The camera is an ordinary element, so a rule can fly it — the same six-line shape that bounces a ball.

"eye": { "type": "camera", "pos": [7, 6, 9], "target": [0, 0.5, 0], "tick": true,
         "rules": "{ float a = e[\"ang\"]; a = a + 0.01; e[\"ang\"] = a;
                     e[\"pos\"][0] = 11 * cos(a); e[\"pos\"][2] = 11 * sin(a);
                     return 1; }" }

Particles

An emitter is an element. The effect templates stay in their own json file — they were always a document — and the element names one:

"jet": { "type": "emitter", "source": "effects.json", "effect": "spark",
         "x": 25, "y": 60, "rot": 270, "on": true }

->on runs it; ->burst emits once and the renderer spends the flag. The update and the render are the renderer's, not the program's.

Underneath: the gpu_* verbs

The record is drawn by an older, larger surface that is still present and still callable: gpu_<category>_<action> — gpu_screen_*, gpu_frame_*, gpu_draw_*, gpu_image_*, gpu_camera_*, gpu_model_*, gpu_shader_*, gpu_light_*, gpu_collide_*, gpu_ray_*, gpu_mouse_*, gpu_key_*, gpu_grid_* (a data-grid widget), gpu_gui_* (an immediate-mode GUI) and gpu_pfx_*.

Every one of them runs on both backends — the register VM calls the same C function the native build does, so a program that draws natively draws under -P pcode=risc too, and so do the three media verbs, which is why a scene draws on both backends from one implementation.

Write new programs as scenes. The verbs are what a scene is made of, and they are there when a program needs something the schema does not yet say.


18. JSON

json is a first-class type with subscript syntax, on both native and the register VM.

json j = parseDoc("{\"user\":{\"name\":\"Alice\",\"age\":30},\"tags\":[\"x\",\"y\"]}");
name = j["user"]["name"];      // chained subscript, coerces on read -> "Alice"
age  = j["user"]["age"];       // -> 30

The arrow is the same read: j->user->name is j["user"]["name"], on a parsed document exactly as on a _json literal, and a missing member answers what the bracket answers. The one word the arrow keeps for itself is doc -- the document's own facts are j->doc->count, j->doc->type and the rest, so j->count is refused naming j->doc->count rather than read as a member called "count" (write j["count"] for that).

Writes, and auto-vivification

ship["hull"] = 90;                 // creates the object on first write
ship["hull"] -= 10;                // compound forms: -= += *= /= %=
cfg["screen"]["w"] = 800;          // "screen" vivifies as an object
cfg["servers"][0]["host"] = "a";   // "servers" vivifies as an ARRAY, [0] as an object
json c;  c[0] = 10;                // an int index at the root: c is an array

A write updates an existing key rather than appending a duplicate. Nested chained writes mint any missing intermediate level as they go, to any depth, including on a still-null base (json d; d["a"]["b"] = 1;).

What kind each level becomes is decided by the subscript that indexes it: a string key means that level is an object, an int index means it is an array. That one rule is read off the chain from left to right and has no depth limit, so doc["servers"][0]["host"] = "alpha" builds {"servers":[{"host":"alpha"}]} — the nesting you wrote is the nesting you get, without a save-and-reload round trip.

Vivification creates, but never converts. An existing level of the right kind is reused; a json null is promoted in place; but an int index into a level that already holds an object — or a string key into one that already holds an array, or anything written through a scalar — is refused loudly, with nothing written, because silently changing a level's kind would throw away data you already stored. A read never vivifies.

Storing past the end of an array grows it, padding the gap with real json nulls (d["a"][5] = 1 on an empty array gives five nulls and then the value). A negative index is refused.

A name whose first use is a string-keyed subscript write needs no declaration at all — inv["gold"] = 100; on a fresh identifier births a json object exactly as if json inv; had preceded it. A read and a computed key still decline and error as before, so this can't surprise you on an already-typed variable.

THE ROOT BELONGS TO THE SUBSCRIPT (v3.286.0)

json d; creates a real document immediately — like every other container declaration — but with no root kind yet. The first subscript decides it, in whichever direction that subscript asks for, and this is the same rule every nested level has followed since v3.264: a string key means an object, an int index means an array.

json a;  a["k"] = 1;      // {"k":1}   -- a string key made the root an object
json b;  b[0]   = 10;     // [10]      -- an int index made it an array
json c;  println(exportDoc(c));   // null -- nothing has decided yet, and it says so

Three consequences worth knowing:

It works wherever the deciding subscript is written. Inside a callee, through a computed index, at file scope — the decision happens when the write runs, so nothing has to see it in advance:

function fill.v(json d) { d["servers"][0]["host"] = "alpha"; }
function demo.v() {
    json d;
    fill(d);
    println(exportDoc(d));      // {"servers":[{"host":"alpha"}]}
}

A decided root stays decided. Once a subscript (or a parse) has given the root a kind, a subscript of the other kind is a loud refusal that writes nothing — creates never convert, exactly as at every nested level:

json d = parseDoc("[1]");
d["k"] = 7;                  // refused, loudly; the document is still [1]

The root's kind is the first subscript's decision, not the constructor's. A string key makes the root an object, an int index makes it an array — the same rule that governs every nested level. Committing the kind at construction is precisely what made a mismatched first store silent: the value went nowhere, or was filed at an index with its key thrown away, and nothing said so. So there is no constructor that decides it. Declare with a bare json d;; where a call is what you need — reassigning an existing variable to a fresh document — jsonUndecided() is the same thing the declaration mints.

A name whose first use is a subscript write still needs no declaration at all — inv["gold"] = 100; on a fresh identifier births the document exactly as if json inv; had preceded it.

Fused read-modify-write

ship["hull"] = ship["hull"] - dmg and its compound form ship["hull"] -= dmg compile to one key walk (a single native call does the read, the op, and the write) instead of two — measured roughly 1.7–1.9× faster on a hot loop. This kicks in automatically for a literal string key, a + - * / % op, and a pure right-hand side; anything else (computed key, a call in the expression) falls back to the two-step path, identically on both backends. jsonRmwNum/jsonRmwStr are the same fusion exposed as callable builtins for manual use.

Exact int64

A json number keeps an exact int64 lane when it was born integral (an int-typed write, an integral parsed token, or an int RMW), so full-range 64-bit integers round-trip through set/get/RMW/export/parse without rounding through a double past 253. A float write to the same key returns it to float.

Coercion on read

The value's type resolves from context: a typed sink (int n = j["k"];, a printf %d/%s/%f) coerces to that type; an untyped sink (print(j["k"])) renders as text; a binary op takes its concrete sibling's type (total + j["price"] reads as float if total is float). json op json with nothing else to infer from defaults to integer and warns — use jsonAsStr/jsonAsInt/jsonAsFloat to be explicit.

Truth in a boolean context

A json variable is a pointer: if (z), z == 0 and z != 0 all test whether there is a node, and agree with each other. The value is reached by a dereference — d["z"], a cast (int)z, or an assignment into a typed sink.

json d = parseDoc("{\"z\": 0}");
json z = d["z"];
if (z)      { }      // TRUE  — the handle
if (d["z"]) { }      // FALSE — the member
int n = z;           // 0     — the member

For a dereference, a scalar is its value; anything that is not a scalar is its validity — {} included, which is true. A string is true, empty or not ("same as C would do"), so the falsy set is exactly null, 0 and false. Emptiness is a different question: ask doc[k] == "", or assign to a string and read ->len — a json node itself has no ->len.

if, while, a for condition, a ternary condition, !, &&, || and the one-argument assert all ask this one question, on both backends.

Building nested JSON

json d;                         // root kind decided by the first write below
els = jsonArr();                  // array
json e;
jsonAddNum(e, "x", 10); jsonAddNum(e, "y", 20);
jsonArrAdd(els, e);                // append object to array
jsonAddNode(d, "els", els);      // nest array under a key
jsonSave("layout.json", d);

exportDoc(j) serializes to a string — compact by default, 2-space indented as exportDoc(j, CX_JSON, 1), because indentation is an argument and not a second verb; jsonSave/jsonSavePretty write the same two forms straight to a file. jsonLen(j) dispatches by node kind (string length / child count / text-form length).

jsonStringify and jsonStringifyPretty were removed (they answer CX-E1028). They were pure aliases -- each bound the very same code as its Export twin -- so one operation answered to two names and no program could tell them apart. Both went together: leaving stringify alive without its pretty twin would have been a surface where one spelling works and its obvious sibling is a hard error.

Literal blocks — _json h { ... }

Instead of building a document call-by-call, or escaping every quote in a parseDoc("{\"...\"}") string, write the JSON as itself in a braced block. The compiler parses it at build time and lowers it to the builder calls above — so a malformed literal is an error at that .cx line, never an invalid handle at runtime.

string sName; int nAge; float fHp;
sName = "Ada"; nAge = 30; fHp = 99.5;

_json cfg {
  {
    "name": sName,
    "age":  nAge,
    "hp":   fHp,
    "tags": ["a", "b", 3],
    "meta": { "level": 7 }
  }
}
name = cfg["name"];                // an ordinary json handle from here on

sName/nAge/fHp are unquoted, so each is a CX variable spliced in — the float enters as a number rather than round-tripping through text — and "meta" shows that objects and arrays nest to any depth. Note what the block does not carry: no // comments inside the braces. It is strict JSON in there, so a comment is a compile error at that line, exactly like a comment in a .json file.

Literal blocks — _rules s { { ... } }

A CX rule is a short C body over the entity json e, and it is data: a string the program can swap at runtime, from a config or from an AI, with no recompile. The price used to be writing that C inside a string literal — every quote escaped, and, worse, unchecked: a typo is not a compile error, so the rule fails to compile the first time it runs, prints a complaint, and does nothing while the program carries on as though it had fired.

_rules takes the same text in braces and runs it through the rules-C frontend at build time:

#pragma rules c

_rules COMBAT {
  {
    float dmg = 0;
    json hits = e["hits"];
    foreach hits { json it = jsonGet(hits); dmg = dmg + it; }
    e["hull"] = e["hull"] - dmg;
    if (e["hull"] < 30) raise(50, e);
  }
}

ruleExec(ship, COMBAT);        // an ordinary string argument -- nothing new here

19. XML and Regex

XML did not keep its original shape. An XML document is a json document that knows how it should be written down, so the thirteen-verb tree API is gone and every question about a document is a json question — d["name"] for an element, d["@id"] for an attribute, exportDoc(d) for the text back out. No name keeps an xml prefix either: reading text is the one moment a format mattered, and parseDoc(s) answers it without being told — a document that begins with < says so. Write parseDoc(s, CX_XML) when you want the strict read. Regex is new since the original reference: regexmatch/regexextract/regexreplace/regexcount.


20. SQL and SQLite

CX carries a real SQL engine, and it is not a lookalike: the engine IS SQLite. The amalgamation (3.53.4) is vendored in third_party/sqlite/ and compiled in, so a database CX writes is an ordinary SQLite file — opened by the sqlite3 shell, by DB Browser, or by any language's SQLite binding, with nothing in between. There is no CX database format and there is no exporter, because there is nothing to export from.

Two surfaces meet in one statement, and the difference between them is most of this section.

THE ENGINE RUNS IN MEMORY, AND WHAT IT CAN SEE IS WHAT THE PROGRAM OFFERED. sqlTable(name, doc) registers a json array of objects under a table name; sqlQuery runs one statement over the registered names and answers a json ARRAY OF OBJECTS keyed by the statement's own columns, so a result is self-describing and printable. Registration is the BOUNDARY: a query — including one written by a rule, or by a model — can reach only what the program deliberately offered, so the set of registered names is the whole of what SQL can see. Registering costs nothing until SQL actually runs, because the engine is opened lazily; a duplicate name is refused rather than silently rebound.

A FILE IS AN ATTACH, AND THAT IS THE ONLY WAY ONE EXISTS. sqlOpen(path) attaches a database beside the in-memory engine under the schema disk — sqlOpen(path, alias) for any other name — creating the file if it is not already there, and opening one written by something else in exactly the same way. sqlClose(alias) detaches it. Closing is rarely necessary, since the connection is torn down with the program, but ten is the engine's limit on attached files, so a program that walks many databases must let go of each one.

The path is BOUND, never spliced into the statement, and ? parameters take their values in order from a json array passed as sqlQuery's second argument. The alias cannot be bound — SQL has no parameter in that position — so it is validated as a plain identifier instead, because it becomes part of the query text the program then writes; a quoted alias would be safe and unusable, and both are refused.

⚠ A REGISTERED CONTAINER IS A LIVE VIEW, NOT ROWS IN THE FILE. This is the one thing that surprises people, and it follows from the two surfaces being different things: sqlTable offers the container as it stands when the query runs, so attaching a file does not persist it, and detaching does not lose it. To persist one, write it; to read it back, read it. No save verb and no load verb exist, because both are already spelled in SQL — and once the attach succeeds, everything else IS already SQL: a registered container and a disk table join in one statement.

_json orders {
  [ { "customer": "ada", "spend": 250 },
    { "customer": "bob", "spend":  90 },
    { "customer": "ada", "spend": 410 } ]
}

sqlTable("orders", orders);                    // the boundary: this, and nothing else

if (sqlOpen("orders.db") == 0) { println(sqlError()); }

sqlQuery("create table if not exists disk.ledger (customer text, spend integer)");
sqlQuery("insert into disk.ledger select customer, spend from orders");   // the view -> the file

json args = jsonArr();
jsonArrInt(args, 100);
json big = sqlQuery("select customer, spend from disk.ledger where spend > ? order by spend desc", args);
foreach big { printf("%-6s %d\n", jsonGet(big)["customer"], jsonGet(big)["spend"]); }

sqlClose("disk");

sqlError() is the one place a failure is explained, whatever refused it — the engine, the table module, or registration — and it answers "" when the last call was happy. Handing that text back to a model is what turns a refused statement into a corrected one.

sqlIndex(table, column) builds a sortndx index over one column of a registered table, and the planner then answers an equality on that column by binary search instead of walking every row. Whether the index was actually used is read off the plan, never asserted.

⚠ IN A BROWSER THERE IS NO ENGINE AT ALL. The wasm module carries no SQLite, so sqlOpen refuses there BY NAME rather than pretending — a file in a tab would live in the tab's memory and die with it. A .db is not a document format either: parseDoc takes text, and a database is a binary meant not to fit in memory, so that spelling refuses and points here.

Reading the rows a query returns is a json question, and §18 answers it — subscript the cursor, and the value coerces at the sink.


21. Testing

CX has two different things called testing, and they are not rivals: #test blocks, which are part of the language and travel inside the program, and the repository's own numbered fixture suite, which is a convention of this tree.

21.1 #test blocks

A #test block is a named piece of code that lives beside what it tests:

function add.i(a.i, b.i)  { return a + b; }
function shout.s(s.s)     { return toupper(s); }
function unused.i()       { return 0; }

#test "add sums two ints" {
   assertEqual(5, add(2, 3));
   assert(add(-2, 2) == 0);
}

#test "shout upcases" {
   assertEqualStr("ABC", shout("abc"));
}

println("the program's own output");

Without a flag asking for them, #test blocks are not compiled at all. They do not reach the scanner, so they cost nothing in a normal build — not a branch, not a symbol, not a byte. A block whose body is not even valid CX will not stop a normal build, which is the honest test of "compiled out" and is how the compiler's own pin checks it.

Two flags carry them, and they do different jobs:

flagwhat it does
--testruns only the blocks, in source order, and answers with one json document written to <program>.test.json beside the program. The program's own top level does not run, and stdout stays the program's own.
--test-out FILEwrites that document to FILE instead of the default. It redirects the answer; it does not duplicate it.
--debugruns the program normally, carrying its blocks, and serves the level-two channel (see --debug-channel).

The exit code is a banded verdict -- one mapping, the same bytes on all three OSes. A POSIX shell shows the low 8 bits unsigned, so "negative" is defined as the two's-complement band; a reader recovers the signed answer as code > 127 ? code - 256 : code:

exitsigned readmeaning
00every block passed, nothing to warn -- and a block-less run, which is information, not a failure
1..127+nwarnings: n functions no block reaches (counted only when the program has blocks)
129..255-nerrors: n failed blocks (255 = one failed block, capped at 127)
128-128no verdict at all: the compile failed under --test

Inside a block, use the assert family the language already has — assert, assertEqual, assertNotEqual, assertEqualStr / assertStringEqual, assertFloatEqual. Nothing new was invented for #test: under --test those same calls are recorded into the document instead of printing their usual [PASS]/[FAIL] line, so an assert that prints is an assert that was not recorded.

The answer is one json document (cx --test demo.cx, abridged here: the program and file paths are shortened, and compiler carries whatever version ran it):

{
  "kind": "cx.test.result", "v": 1,
  "program": "demo.cx", "compiler": "3.3.007.0", "backend": "risc",
  "totals": { "blocks": 2, "passed": 2, "failed": 0,
              "asserts_passed": 3, "asserts_failed": 0 },
  "blocks": [
    { "name": "add sums two ints", "file": "demo.cx", "line": 5, "result": "pass",
      "asserts": [
        { "kind": "assertEqual",    "expected": 5,      "actual": 5,     "result": "pass", "line": 6 },
        { "kind": "assert",         "expected": "true", "actual": "1",   "result": "pass", "line": 7 }
      ] },
    { "name": "shout upcases", "file": "demo.cx", "line": 10, "result": "pass",
      "asserts": [
        { "kind": "assertEqualStr", "expected": "ABC",  "actual": "ABC", "result": "pass", "line": 11 }
      ] }
  ],
  "uncovered": {
    "definition": "a function defined in this program that no #test block reaches -- decided at compile time from the call graph, never instrumented at run time",
    "count": 1,
    "functions": [ { "name": "unused", "file": "demo.cx", "line": 3 } ]
  }
}

Read it with parseDoc like any other document — the point of answering as json is that a program, or an AI, can act on it without parsing prose. Each assert row carries the pair, expected and actual, not only a sentence: a message is for a person, the pair is what a fix is made from. Each row names its own line, not the block's.

uncovered is decided at compile time. It lists every function of the program that no #test block reaches, walking the call graph — so a function reached only through another counts as covered, and a function reached only from a block inside an inactive #if does not. Nothing is instrumented and nothing is measured while the program runs: untested code stays exactly as it was. A program with no #test blocks at all reports every function uncovered and exits 0 — that is information, not a failure.

The exit code follows the document: 0 when no block failed, 1 when one did, on both backends. A runner can gate on it.

Both backends answer the same document. Native and -P pcode=risc produce byte-identical json apart from the backend field, and the compiler fills that field in because it knows which one it emitted — it is never sniffed at run time.

Two shapes worth knowing: a block may be written on one line (#test "x" { assertEqual(2, inc(1)); }), and #test is a directive among directives, so a block inside a false #if is not compiled and does not run.

The level-two half of the protocol — a channel that lets a running program and the engine talk while it runs — is what --debug and --debug-channel serve; it is documented with the anvil rather than here.

21.2 The repository's own test suite

Each test is a .cx file with a numeric prefix, #include "default.cxi" for assertion macros, and assert/assertEqual/assertFloatEqual/assertStringEqual calls:

#include "default.cxi"
x.i = 41;
assertEqual(x + 1, 42);

Numbering convention: 0xx basics/control flow, 1xx functions/typed decls, 2xx arrays, 3xx pointers, 4xx structs, 5xx collections, 6xx algorithms/perf, 7xx perf benchmarks (often assertion-free), 14x the AI suite. The canonical test tree currently runs to at least #720 (well past 300 numbered tests, not counting a retired/ subfolder) — a big jump from the ~120–170 of the PB era; treat any specific pass-count you see quoted elsewhere as a snapshot, since this suite grows with nearly every commit.

The sweep runner (tests/sweep.ps1 / tests/cport_cxc_sweep.ps1) checks a golden expected-output file per test on both the native and register-VM builds — matching each other isn't sufficient, since both could be wrong the same way; each must independently match the golden.


22. Cryptography

md5/sha1/sha256/sha512/crc32/base64enc/base64dec — unchanged.


23. Date/Time

date()/time()/year()/month()/day()/hour()/minute()/second()/milliseconds()/elapsed()/delay(ms)/monthname(n) — unchanged.


24. System

getenv works on both backends. Where a call does not, the compiler says so at your own line rather than at link time:

CallNativeRegister VMIf you write it anyway
getenv(name)yesyes—
exit(code)yesnoCX-E1013: native-only; use the native backend or _C{}

exit has no row in the Builtins Reference and that is not an omission — it is not a CX builtin at all. It is libc's exit, reached because native CX is C; the reference lists what the compiler BINDS, so a libc function that needs no binding cannot appear there. It does set the process exit code (exit(2) exits 2). An absence from that page is evidence about the binding table, never about the language.

The old realptr/realaccess raw-memory-address functions assumed the PureBasic VM's flat memory model; given the register VM's slot-handle representation of .p (§10), don't assume these still mean what they used to without checking.


25. Pragma Directives

Confirmed current:

PragmaEffect
#pragma named "a, b, c"Declare global names AI-generated bytecode may FETCH/STORE
#pragma checks on (aliases check, checkbounds)Turn on the runtime value checks — hard-error grown-array bounds violations instead of the lenient default, refuse a field access through a null struct pointer instead of dereferencing it, and refuse access through a container containerDelete released instead of answering the element type's zero (CX-E5044). Checks are off by default on both backends; this pragma turns them on for both; a program behaves identically compiled native or to the register VM either way. See §"Checked mode" and §containerDelete and §32 (every one of these refusals routes through an installed onerror handler)
#pragma parse strictGoverns the TYPE question only (his ruling 2026-09-09: "strict is about the type not about the error it may encounter"). A document that declares a known type — {/[ for JSON, < for XML — and then fails to parse as it is REFUSED whether this pragma is on or off: an error is an error either way, and handing such a document back as CX_TEXT with none of its keys is what that refusal exists to prevent. Prose that declares nothing is text under both settings and is never reported as an error — a failed JSON probe on a non-declaring document is how TEXT is detected
#pragma risc allForce whole-program register-VM codegen
#pragma decimals NHow many decimals a bare %f prints, in printf/sprintf/print (default 3, not C's 6). An explicit precision such as %.6f always wins; last one in the file wins — see §12
#pragma temp_ring_size NPer-function _temp ring slot count (default 8)
#pragma gc default=N max=N min=NArena sizing, and the two numbers a program may set. default= is the arena's STARTING size; growth from there is automatic. max= is the CEILING (default 4G) — reaching it is a loud refusal that names this pragma and prints the number, never a silent failure. Sizes take a K/M/G suffix. realloc= is default= under its prior spelling and still works. A document too large for the ceiling is refused as a CAPACITY limit and says so — it is never reported as a malformed document, and parseDoc states that the document is well-formed so you are not sent hunting for damage that is not there
#pragma randomseed NPin the RNG stream so the program replays identically — see below
#pragma localdeclares on (default off)An undeclared name inside a function becomes a local that shadows any same-named module-global, instead of binding it — C's discipline, reported with CX-W1011. _local on one function does the same thing at that grain (_forcelocal is its heritage alias). Never an error; see §4 "Global vs. Local"
#pragma BuildLTO yes (same as --lto)File-wide link-time optimization — a measured ~4× win on gcc for container/string-heavy code
#pragma nowarn W1010,W1012Silence those specific warning codes for this program — see below
#pragma onerrordefault offStop CX writing its own structured record for a failure to the side channel. Your registered onerror handler still runs — see §32

#pragma nowarn — silence a warning you can name (v3.262.0)

#pragma nowarn W1010          // one code
#pragma nowarn W1010, W1012   // several; commas or spaces, and the lines accumulate
cx game.cx --build -P nowarn=W1010

The two spellings are two doors onto one judgement, and they union — -P nowarn=W1010 alongside #pragma nowarn W1012 suppresses both. That makes nowarn the one pragma key where the command line does not override the source: a set has no last writer, and overriding would mean that asking for one more suppression removed one.

The warnings are silenced, never uncounted. The per-site lines go; one line per requested code stays, naming the code, the number withheld, and which door asked:

cx: note: CX-W1015: CX-W1010: 566 occurrence(s) suppressed by #pragma nowarn -- remove the suppression to see them
[cx] done     0 warnings (566 suppressed) in 7.56s

So a suppressed program still proves the warning fires, and a count that moves is still information. A requested code that suppressed nothing reports 0, which is how you find a nowarn line that has outlived its reason.

Specific codes only. There is no all or * form, and E codes are refused (CX-E0045) — an error is a refusal, not an opinion. A warning that carries no code cannot be named and therefore always prints.

#pragma cxc_fallback tagged is an opt-in codegen fallback, and the only fallback pragma the compiler reads. When an assignment's right-hand side is a call cxc cannot type — a runtime shim it never parsed — the emitter normally leaves it alone; with this on, it emits a best-effort cast and a warning, so the uncertainty is visible in the generated C rather than silent. Off by default: the auto-trigger also fires on legitimate calls into runtime macros.

Seeding the RNG — three layers (v3.215.0)

Random is random by default. A program that never seeds anything gets a fresh stream on every run, because the runtime seeds from real entropy at startup. When you need the opposite — a golden test, a bug report, a level you can regenerate — you pin it, and there are two ways to do that. The layers override each other in the order they run:

LayerSpellingScopeUse it for
1 — default(nothing)whole programgames, demos, anything that should feel unpredictable
2 — pragma#pragma randomseed Nwhole program, from before the first statementgolden-compared tests, replaying a run without touching the code
3 — callrandomseed(n)from that point onpinning one section (a world seed) while the rest stays unpredictable

#pragma randomseed N is randomseed(N) placed as the program's first statement — not a second, parallel way of seeding. Last one wins if the file carries more than one, the same as #pragma decimals. N must be a plain integer; anything else is refused with CX-E1052 rather than silently ignored, because a program that says it is pinned and is not still looks right until the day its output matters.

Two details worth knowing:

Gone: #pragma optimize (the post-codegen bytecode optimizer it controlled was removed at the 2026-06-04 CISC retirement). The old stack-size pragmas (GlobalStack/FunctionStack/EvalStack/LocalStack) belonged to the retired stack-based VM — the register VM has no eval stack, so these don't carry meaning in the current model; don't rely on them.


26. Markers and Runtime Code Modification

Marker syntax now uses a leading backslash — \{N: / \N:}, not the old {N:/N:}:

\{5: "compute something"
   result = x * 2 + 1;
\5:}
generatecode("compute result as x squared plus y cubed", 5);   // AI writes replacement bytecode
replacecode(5);                                                  // swaps it in
setcode("...");                                                  // hand-supply bytecode, bypassing the LLM (deterministic offline tests)
clonemarker(srcId, dstId);                                       // copy one marker's active body to another

Bytecode Introspection

FunctionReads
peek_code(pc)opcode at PC
peek_i(pc) / peek_j(pc) / peek_n(pc)operand fields
peek_flags(pc)flags field (renamed from peek_ndx — the register VM's instruction shape changed)
peek_funcid(pc)function ID
poke_*write counterparts
code_size()total bytecode length

Running Marker Bytecode as Native (Mode 4)

Marker and codeswap bytecode doesn't have to stay interpreted. cx --run-lift decompiles a program's own register-VM bytecode back to CX source, recompiles that through the normal native path, and runs the result in-process — recovering most of the VM's interpretation cost for whatever the decompiler can faithfully reconstruct. Coverage today is the decompiler's coverage: scalar arithmetic, strings, and simple control flow lift cleanly; structs, arrays, maps, and multi-argument user-function calls don't yet, and a program using them declines the lift with a clear error rather than running it wrong. Outside that subset, the register VM remains the real, current execution path for marker-generated and codeswapped code — this isn't a claim that interpretation is gone, only that it's no longer the only way such code can run.


27. AI Integration

Providers

ai_set_provider("anthropic");   // openai / google / mistral / cohere / xai / deepseek / groq / ollama / kwaai / custom
json cfg;                        // a key is optional — this provider's own env
cfg = parseDoc("{\"models\":[{\"provider\":\"anthropic\",\"key\":\"sk-...\"}]}");
ai_set_options(cfg);             //   var is read otherwise
ai_init();                       // 1 if a key resolves

Each cloud provider reads its conventional env var (ANTHROPIC_API_KEY, etc.). Ollama is local-but-HTTP.

Fallback, Cache, Async

// The ORDER of `models` IS the chain: groq first, ollama behind it. Each
// section carries its own endpoint, model and key, so falling through changes
// nothing about how the first one was configured.
cfg = parseDoc("{\"models\":[{\"provider\":\"groq\"},{\"provider\":\"ollama\"}]}");
ai_set_options(cfg);

hit = ai_cache_get(key);
if (hit == "") { result = ai_call(prompt); ai_cache_put(key, result, 3600); } else { result = hit; }

req = ai_call_async(prompt);
while (ai_async_ready(req) == 0) { }
result = ai_async_result(req);

aifunc — a function whose body the LLM writes

aifunc function add.i(a.i, b.i) { "Return a + b." }
aifunc function clamp.i(x.i, lo.i, hi.i) { "Return x clamped to [lo, hi]." }

aifunc resolves when the function is CALLED, so the key is a run-time requirement and two calls can differ. An aifunc with no return suffix returns a variant — the answer is runtime-shaped, and coerces at the call site (int n = ask(q); parses it, string s = ask(q); keeps the text). Pin a concrete return type to skip that.

Codeswap Markers

Covered fully in §26 — this is the runtime-mutable half of the AI story; aifunc is the compile/startup half.

Legacy Surface

ct_createaifunc/rt_createaifunc/ai_call_func/ai_parse_bytecode/ai_dump_asm still function; new code should prefer aifunc + markers. (ai_set_named/ai_get_named were retired 2026-09-03 -- a #pragma named slot is the program's own variable, so assign to it and read it by name.)

Under the Hood

All AI-generated bytecode — whether from an aifunc or a codeswap marker — now executes on the same register VM as everything else (§ARCHITECTURE), not a separate AI-only interpreter. Named variables bind to real storage in place (a VM slot, or the bound native global) rather than being copied in and out.

Deferred, not shipped: a local-model provider (CxLlama, a llama.cpp daemon) and a peer-shared library of vetted generations are both designed but not implemented — don't budget on either being present.


28. Rules and Automation

A newer subsystem, absent from the original reference entirely: a small embedded rule language that runs against a json entity, aimed at "the data carries its own behavior" game/simulation logic.

ruleExec(ship, "{ float d = e[\"dmg\"]; e[\"hull\"] = e[\"hull\"] - d; if (e[\"hull\"] < 30) raise(51); }");
OnEvent(51, gate, &handler);

The rule source is dialect-sniffed: a leading { is a C-rule (plain C over a json e parameter, compiled at runtime by the shared front-end to register-VM bytecode); anything else is the legacy imperative DSL (bare keys, $temps, for (key)). Both run on the same scratch VM, so behavior is identical native vs. -P pcode=risc. A C-rule needs no pragma: calling ruleExec (or any other rule entry point) is itself the instruction, and the compiler links the rule engine because it can see the call. #pragma rules c remains, and is not deprecated — it is how a program says so when the compiler cannot see it, which is any program whose rule text arrives at run time from a file, a socket or an AI. A program that neither calls a rule builtin nor says the pragma pays nothing for either.

A rule written as a literal is checked when you BUILD, not when it fires. OnEvent(51, "{ ... }", &handler) takes rule text in two of its three slots, and until v3.311.0 neither was looked at until the event happened — so a typo reached you the day the key was pressed, or, on a handler nobody exercised, never. Both slots now go through the same build-time check _rules blocks already used, and a malformed one is an error at the .cx line that wrote it. The two gate forms that are not rules — "1" (always) and a plain number on a timer, which is an interval — are left alone, and a gate held in a variable, or loaded with eventSetTable, is still the runtime's to judge: that registry is json and two-way on purpose, and a rule nobody has written down yet cannot be checked before it exists.

The sandbox is deliberately narrow: scalar locals, full operators, if/switch/while/for, e["key"] access, and a whitelist of math/string/json/print builtins. Pointers, _C{}, non-whitelisted calls, and container declarations are declined loudly rather than silently ignored. Loops are budget-capped (default 1,000,000 iterations/tick; #pragma rules budget N) so a runaway rule can't hang the host program — it dies with a clear error and ruleExec returns, letting the caller's tick survive.

Naming the entity, and keeping the rules you learn

Two pragmas, both optional, both doing nothing at all to a program that omits them.

#pragma rules entity ship            // `e` is now spelled `ship`
#pragma rules persistence yes        // rules survive the run that learned them
#pragma rulesfile "fleet.rules"      // ...in this file, rather than <programname>.rules

#pragma rules entity <name> renames the one parameter every rule is written over. The default is e and always will be; nothing about a program that says nothing changes, down to the emitted C. The name is checked by writing it — CX parses the wrapper the name would produce, so whatever the language reserves is refused, at the pragma, with the reason. That is stricter than it sounds: arr declares fine and cannot be read (arr["hp"] is CX-E0007), and a rule whose entity nobody can touch is worse than a refused pragma.

The honest edge is worth stating, because rule text is data that travels: a rule authored over ship runs only where the entity is called ship, and it does not say so when it lands somewhere else. An undeclared subscript births a rule-local json object (§subscript birth), so a rule written over ship and run where the entity is e writes into a local ship and leaves the real document alone — quietly, and that is true of a build-checked _rules block too. A program that renames its entity owns its own rule text.

#pragma rules persistence yes gives the program a rule store — a json file, <programname>.rules unless #pragma rulesfile "name" says otherwise, read at startup before the first statement runs.

#pragma rules persistence yes
rulePromote("{ e[\"hull\"] = clamp(e[\"hull\"], 0, 100); }");   // durable AND live

rulePromote(rule) writes the rule to the store and compiles it into this run, in one call. A rule that does not compile is written too, marked failed with the error beside it — the store records what was tried, which is the point: it is where you see what an AI reached for and what the sandbox refused. A failed entry is reported once at load and never applied.

The file is text and is meant to be edited. ruleStore() hands the program the same array as data, so a program can walk its own vocabulary, show it, or prune it. There is no file-watching: a store is read at startup, so what you read is what runs.

Handle metadata inside rule text: ->, same as everywhere (v3.187.0)

Rule text reads handle metadata with ->, exactly like program text — int n = e["skus"]->count; inside a rule is the same surface, the same five fields (->count ->cap ->type ->valid ->id), and the same lowering. The §11 type rule therefore applies here too: on a typed handle the call spellings are CX-E1110 in rule text as well, and CX no longer has any place where one idea has two spellings.

Until v3.187.0 this was the single exemption, and the reason was mechanical rather than a preference: -> parses as a field read on a dereference, and the rule sandbox declines that entire node class (the same rule that rejects pointers), so the arrow could not be written in a rule at all and the call forms had to stay legal there. The sandbox now admits that one shape — a metadata field on a handle base — and nothing more.

What is still refused inside a rule, so the sandbox is no wider than before:

in rule textresult
e["skus"]->count, e->validthe metadata surface, allowed
e->hulldeclined — a non-metadata field name is still struct field access, and never was a spelling of e["hull"]
*p, &x, struct/pointer localsdeclined, unchanged
n->count where n is an intan error, not a silent 0
jsonSize(e["skus"]) and the other call spellings, on a typed handleCX-E1110, naming the field that replaces them

This matters most for AI-authored rules: an aifunc writing a rule from documentation older than v3.187.0 will emit the retired call form, and a rule that fails to compile makes arm() return 0 — the rule silently never fires. Regenerate the prompt rather than translating it (see the AI Manual).

This is the mechanism behind "an entity's state and its behavior both live in the same JSON container" — and, notably, the same primitive an aifunc uses to write a rule at runtime means an AI can author new entity behavior as data, with no recompile. It's also the reason the older CGI2D engine is being phased out (§30).


29. Embedding Files and the Asset Resolver

embed(path) bakes a file's bytes into the executable at compile time, through C23 #embed. The toolchain minimum is gcc 15 or clang 19. An older compiler — Ubuntu 24.04's gcc 13, Apple clang 16 — gets a clear compile error naming those versions, never a silent stub, and on such a box asset(disk) / asset(datafile) do the job without baking. Everything else in the language builds on the compiler the installer checks for:

data = embed("levels.json");
levels = parseDoc(data);
atlas = gpu_image_loadmem(embed("sprites.png"), ".png");

asset(path) goes further: one call site, three deployment modes, switched by a single pragma and a recompile, no source change:

#pragma assetmode disk        // dev: load from filesystem
// #pragma assetmode embed    // release: bytes baked into the exe
// #pragma assetmode datafile // packaged: loaded from a game.dat archive at runtime

levels = asset("data/levels.json");   // .json  -> a json doc
sprites = asset("art/sprites.png");    // image  -> a texture handle

Resolution happens entirely at compile time (zero runtime branch) — asset() just lowers to the matching disk/embed/datafile builtin for that file's extension and mode.


30. The 2D UI Layer (CGI2D) and Where It's Going

CGI2D — a JSON-driven 2D HUD/UI engine (describe widgets in JSON, bind them to CX variables, render every frame) — still exists and still works: cgiInit/cgiLoadJSON/cgiRender/cgiUpdate/cgiLinkVar/cgiHitTest/cgiGet/cgiSet, same shape as the original reference described.

It's explicitly being replaced, not extended. The engine's JSON-binding model is exactly the boilerplate the newer rules/automation system (§28) deletes: instead of describing a widget in JSON and hand-wiring it to CX variables, a rule is the binding — an entity is data + rules + an invisible per-tick loop. The plan is a piecewise cannibalization: the pure drawing primitives (text layout, color/font lookup, hit-testing) get lifted into a small standalone draw module with no JSON dependency at render time; the ~20 widget-specific renderers get ported one at a time as draw-rules; and the JSON element-loading/data-binding layer — the part with the real complexity, and a since-patched string-ownership bug — gets dropped in favor of rule state. If you're building new UI, treat CGI2D as usable-but-legacy and look at whether a rule-driven entity fits your case first.


31. Libraries — Status

The original library mechanism (lib_save/lib_load/lib_call, backed by .ocx bytecode serialization) is not currently functional in the C port — the compiler doesn't load .ocx at runtime today (each register-VM-eligible function ships its bytecode as a string baked into the executable instead), so there's no consumer for a saved library file yet. This is open, designed-but-not-built work, not a small gap.

A related but different and also not started effort is exporting CX-compiled code as a consumable DLL for other languages (C#/Python/Rust via P/Invoke or an embedded-engine model) — most of the native↔VM marshaling machinery the register VM already needed turns out to cover much of what this needs too, but it hasn't been wired up as a public feature yet.



32. Error Handling — onerror and #error-check

CX has ONE error door. Install a handler with onerror, and every runtime error that would stop the program passes through it on the way out — the checked-mode guards, the per-verb refusals, and your own checkpoints, all with the same payload.

The handler

function myErrors.v(json e) {
    printf("error %d in %s: %s\n", e["code"], e["function"], e["message"]);
}

onerror(&myErrors);

The handler takes one json parameter and returns nothing. What arrives is a subscriptable handle:

FieldWhat
e["code"]the error number — your checkpoint (1–999) or a CX-E####
e["message"]the full formatted message
e["file"]the source file, where the error site recorded one — native only
e["line"]the source line, likewise — native only
e["function"]the CX function the failure happened in

That is the same payload OnEvent hands its handlers, which is the point: a field can be added later without breaking a signature you already wrote.

e["file"] and e["line"] arrive EMPTY on the register VM, and that is an accepted asymmetry rather than a bug on a list. The VM has reported cx: error: with no position since long before the error door existed: it carries no line table at run time, so its check sites have no position to bake in. Closing the gap means giving the VM one — its own arc, with its own cost on every loaded program — and this door deliberately neither widened the gap nor paid for it. e["function"] is populated on both backends, because the function identity is already there for other reasons; it is the file and the line that are native's alone. Write handlers that read a missing position as missing, not as wrong.
The older two-parameter shape still works. function myErrors.v(n.i, msg.s) compiles and runs exactly as before and warns once (CX-W1017) that it has been superseded. It cannot carry the file, the line or the function, so every field worth having would have been another signature break — which is why the handle form exists.

A function of neither shape is refused where you install it, not where it runs:

function bad.i(n.i) { return n; }
onerror(&bad);

The argument must be &function. Anything else — a variable, a string, an expression — is CX-E1094:

x = 3;
onerror(x);

onerror() with no argument uninstalls. There is one handler at a time: a second onerror(&other) replaces the first rather than stacking.

There is a handler even when you write none

CX writes a structured record for every failure with no registration at all — the code, the message, the source position and the function it happened in — to the same side channel onwatch reports to (§33). One json object per line:

{"kind":"error","code":5023,"message":"CX-E5023: nums: index 9 out of bounds (size 4)","file":"game.cx","line":41,"function":"takedamage"}

It writes nothing at all unless $CX_SIDE_CHANNEL names a destination, so a program's stdout and stderr are exactly what they have always been. Under the screen server the session names one, so every program it runs reports its failures structurally without being changed.

#pragma onerrordefault off turns the record off. Registering your own handler does not: yours runs as well, and the pragma governs only the default.

Your own checkpoints

#error-check <n> <text note> raises error n with that note. It is a directive, so the number and the range are checked while the program compiles.

function myErrors.v(n.i, msg.s) {
    printf("caught %d\n", n);
}
onerror(&myErrors);

#error-check 7 the config file had no [server] section

The note reaches the handler prefixed with the source position, so a log line says where the checkpoint was without you writing the location twice.

Checkpoint numbers are 1 to 999. CX's own error codes are 1000 and up, which is what lets a handler select n on one argument and never confuse "my checkpoint 7" with "CX's bounds error 5023". A number outside the band is refused at compile time (CX-E5045), so a colliding checkpoint cannot ship.

When the message has to be built at run time, call the lowered form directly:

name = "server";
errorRaise(7, "the config file had no [" + name + "] section");

That is the only difference between the two: the directive takes a literal note and gets compile-time checking and a free source position; errorRaise takes any string.

What reaches the handler, and what does not

SourceReaches the handler?n is
#error-check / errorRaiseyesyour number, 1–999
#pragma checks on guards — bounds, divide-by-zero, null pointer, resource cap, stackyesthe CX-E#### number, e.g. 5023
Always-on per-verb refusals — cursor read off a walk (CX-E5041), access through a deleted container (CX-E5044)yesthe CX-E#### number
Compile-time diagnostics — CX-E0xxx, CX-E1xxx, CX-E2xxxno—
Refusals that deliberately do not stop the program — an XML write through a dead node handle reports itself and execution continuesno—

The compile-time boundary is not a gap. Those errors happen before the program runs, so there is no handler to call and never will be for that compile.

The handler does not resume

When the handler returns, the program stops with a non-zero exit — exactly as it would have without one. The handler's job is to see, log and clean up; it decides nothing about control flow, and the statement after a raise never runs.

A handler that raises while it is handling is not re-entered: CX says so and stops, rather than recursing until the stack ends.

Cost

Nothing. The door is only reached where the program was already about to print an error and exit. A program that never calls onerror compiles to byte-identical output on both backends, and nothing on any hot path knows the feature exists. The function in e["function"] costs nothing either: the compiler already knows which function it is emitting, so the name is a constant baked at the error site rather than something the running program keeps track of.

Everything on this page behaves identically compiled native or to the register VM.


33. Watching a Variable — onwatch

A watch is a debugger you compile in. You name a variable, say when you want to hear about it, and CX reports every change that matters — with the old value, the new one, and the function the write happened in.

int hull = 100;

function takeDamage.v(int n) {
    hull = hull - n;
}

onwatch(hull, onchange, "hull <= 70", &takeDamage);

takeDamage(10);
takeDamage(25);

Six arguments, and only the first is required:

PositionWhatValues
1the variablea name, never a string
2the triggeronchange (or 0), or a number of milliseconds
3the conditiona string holding a CX expression, or 0 for "always"
4the scope0 for every function, or &someFunction
5the handleromitted for the side channel alone, or &myHandler
6the relay0 to report every trigger, or a count of triggers per report

The variable is a name, and that is the whole cost model

CX instruments the writes to that name, in that scope, and emits nothing anywhere else. A program with no onwatch in it compiles to byte-identical output — the feature leaves no trace at all.

That only works because the compiler can see which name you meant. A watch named by a string would have to be looked up at run time, which means a table check on every store in the program, watched or not. So the subject is an identifier:

int a = 0;
a = 1;
onwatch("a", onchange, 0, 0);

Watchable variables are int, float and string. A container is refused (CX-E1098): its writes do not go through a store of the name, so there is nothing at that spelling to instrument.

The two triggers, and what the second one costs you

onchange reports every change that satisfies the condition. It is complete, and it pays at every write. With no condition on a hot variable that is exactly what you asked for: one line per change, however many that is. Give it a condition, a scope, or a millisecond trigger.

A number is milliseconds, and it rate-limits the reports:

int ticks = 0;
int i = 0;

onwatch(ticks, 500, 0, 0);

for (i = 1; i <= 100000; i = i + 1) { ticks = i; }

This is a sampler, not a tracer, and it is lossy by construction: writes inside the quiet window update the watch's memory of the value but are never reported, so a value that changes and changes back between samples is invisible.

The trigger limits reports, not instrumentation. Both triggers cost the same at the write — measured at about 1.2 ns per instrumented write, against a bare scalar store of about 1.3 ns. What a millisecond trigger buys is quiet: no line written, no handler called, no clock read except for a write that would have been reported. (An earlier build guarded each time-triggered write with "has the window elapsed?" and it cost twelve times more, because answering that means reading the clock and a clock read is ~18 ns. CX is single-threaded and has no ambient timer to sample from, so a time-based watch cannot be free per store — it can only be quiet.)

The relay: a ceiling that interrupts, not a cap that goes silent

onwatch(v, onchange, 0, 0, 0, 1000) on a hot variable reports once every 1000 changes, and the report carries how many it stands for:

{"name":"ticks","site":"main","old":999,"new":1000,"seq":1,"ms":4,"count":1000,"trigger":0}

Every trigger is still counted — the relay throttles delivery, not the watch. So nothing is lost as a count, including the partial batch when the program ends: at exit CX delivers what is pending, with old set to null because there is no transition to report, only a value and the number of triggers behind it. That tail goes to the side channel only — your handler is not called, on either backend, because by then there is no running program left to call it into. (A millisecond trigger gets a tail too: its last window is delivered with the count of changes it stood for.)

What a relay does cost is the values in between: a relayed line carries the latest change plus a count, not a history of the ones it stood in for. That is the trade, and it is the whole trade.

It composes with the millisecond trigger rather than competing with it — the trigger throttles by time, the relay by volume, and a report waiting on either keeps accumulating its count instead of resetting.

Like the trigger, the relay must be written literally (CX-E1105).

The two real cost controls are still the condition and the scope.

The trigger must be written literally. The compiler picks the instrumentation from it, and it cannot pick from a value that will not exist until run time (CX-E1099).

The condition is written as text and judged as code

You write it in quotes, and CX parses it at that line and compiles the result into each instrumented site. So it is an ordinary CX expression in every way that matters — it reads your variables, it costs what the expression costs, and a typo in it is a compile error on the line that wrote it (CX-E1104), never a watch that silently never fires:

int hull = 100;
function takeDamage.v(n.i) { hull = hull - n; }
onwatch(hull, onchange, "hull <= = 70", &takeDamage);

It reads the variable after the write, which is what a watch means.

Why quotes at all, when an expression needed no parsing? Because a condition is data about the program, and quoting it makes the whole of it legible as one unit — to a reader, to a tool, to a model. What quoting must never cost is the check, and here it costs nothing: the compiler does the same work it did when the condition was unquoted, at the same moment, with the same result. Writing the expression without quotes is refused (CX-E1103) rather than accepted quietly, so nothing silently changes meaning between the two spellings.

A condition is one expression, not a sequence of statements. Text that tries to be more than one is refused by the same rule.

The scope is the other cost control

&someFunction instruments the writes in that body and nowhere else. 0 covers every function. Naming one function is how a watch on a busy global stays cheap.

The record CX keeps is per (variable, site) — so watching one variable in two functions gives you two independent stories, each with its own memory of the last value and its own sequence numbers, rather than one interleaved one.

Where the reports go

Never to stdout. A report on stdout would be indistinguishable from your program's own output. They go to the side channel: the file named by $CX_SIDE_CHANNEL, or a per-process file in the temp directory. One json object per line:

{"name":"hull","site":"takedamage","old":90,"new":65,"seq":1,"ms":12,"count":1,"trigger":0}

count is how many triggers this line stands for — always 1 unless a relay is set (see above).

When CX picks the file itself it says so once, on stderr — but only when stderr is a terminal. A report nobody can find would be the silent failure a watch exists to prevent; a line that varies per run (the name carries the process id) would break every captured comparison there is, which it did to this feature's own corpus fixture before the rule was added. A human gets the hint; a pipe does not, and anything reading a pipe can name the file itself.

Under the screen server, the session sets CX_SIDE_CHANNEL for the programs it runs, so a watched program's reports arrive beside the streams the screen already captured — no change to the program.

A handler, if you want one

string label = "";

function onLabel.v(json e) {
    println(e["name"] + " in " + e["site"] + ": " + str(e["old"]) + " -> " + str(e["new"]));
}

onwatch(label, onchange, 0, 0, &onLabel);

label = "ready";

The handler receives a subscriptable handle — e["name"], e["site"], e["old"], e["new"], e["seq"], e["ms"], e["count"], e["trigger"] — the same payload shape OnEvent hands its handlers, so fields can be added later without breaking your signature. A handler that writes the variable it watches is not re-entered.

A watch that can never fire says so

If nothing in scope ever writes the variable, CX warns (CX-W1016) rather than compiling a watch that would sit silent forever. That silence is the exact failure a watch exists to prevent.

Everything on this page behaves identically compiled native or to the register VM, and the reports are byte-for-byte the same.


Appendix: Complete Example

#pragma decimals 2

struct Point { x.f; y.f; }

function distance.f(a.Point, b.Point) {
   dx.f = b.x - a.x;
   dy.f = b.y - a.y;
   return sqrt(dx*dx + dy*dy);
}

function describe.s(p.Point) {
   return sprintf("(%.1f, %.1f)", p.x, p.y);
}

array path.Point[4];
path[0].x = 0.0;  path[0].y = 0.0;
path[1].x = 3.0;  path[1].y = 4.0;
path[2].x = 6.0;  path[2].y = 1.0;
path[3].x = 10.0; path[3].y = 10.0;

total.f = 0.0;
for (i = 0; i < 3; i = i + 1) {
   d = distance(path[i], path[i+1]);
   printf("  %s -> %s = %.2f\n", describe(path[i]), describe(path[i+1]), d);
   total += d;
}
printf("\nTotal path length: %.2f\n", total);

json d;
jsonAddNum(d, "distance", total);
jsonAddNum(d, "segments", 3);
out = exportDoc(d);
printf("JSON: %s\n", out);
jsonfree(d);

printf("SHA-256: %s\n", sha256(out));

CX+AI Language Reference — v3.137.1 — July 2026. Compiled directly against the live DOCS/ source tree (REFERENCE.md, AI.md, UNIFIED_REGISTER_ISA.md, CLAUDE.md, GPU_NAMING.md, CGI2D_SALVAGE.md, TODO.md, BUILD.md) rather than reconstructed from summary. A handful of items (exact networking builtin names, the full current pointer/peek/poke surface, and the array-param-byref detail in §7) are flagged in place as worth a direct source check if they're load-bearing for your use — everything else here was cross-confirmed across at least two independent live documents.