Contents

Stdlib reference (mere)

221 builtins are always available via initial_env. Check a name's type with mere -te NAME, and the count with sh scripts/host_matrix.sh, which enumerates the environment rather than reading this document — a hand-maintained tally rots, and this one had (it read 202 while the environment held 226).

Legend:

Sugar / prelude added in Phase 36 (2026-06-22)

Syntactic sugar (13): lexer / parser-level changes only; preserves 4-backend compatibility:

Prelude additions (16 of 34 entries added in Phase 36):


I/O (12 + 2 prelude)

NameTypeDescription
printstr -> unitWrite to stdout with newline
print_no_nlstr -> unitWithout newline, and in one `write(2)` rather than through stdio (v0.1.480). main sets stdout line buffered, so an fwrite here flushed at every newline and the recommended idiom — accumulate into a StrBuf, print once — still cost one syscall per line, about 1.6 µs each. Stdio is flushed first, so a print issued earlier still arrives earlier
print_intint -> unitPrint integer with newline
print_boolbool -> unitPrint bool with newline
print_errstr -> unitWrite to stderr with newline. ⚠ Two backends had this wrong until v0.1.503: the Node host for Wasm never provided it (so any program calling it failed to instantiate), and RISC-V wrote the bytes without the trailing newline the other four write
echo'a -> 'aPrelude, not a builtin. Prints its argument to stderr with the line it was written on, and answers the argument, so it drops into the middle of an expression and comes out again without moving anything: let m = echo (n + 1); prints line 7: 21 and binds 21. The front end rewrites echo into echo_at "<where>" — a scope-aware pass, so a binding of your own named echo means what you said. Same bytes on all five backends (scripts/echo_check.sh), and stderr rather than stdout, because a debug print in the answer changes what it is watching (v0.1.503)
echo_atstr -> 'a -> 'aWhat echo becomes. Useful directly when the position you want is not the one you are writing at
read_lineunit -> strOne line from stdin; empty string on EOF
read_file ⚡str -> strRead the whole file as text; raises on failure. On the C backend the str is NUL-terminated, so binary data silently truncates at the first 0x00 byte (the interpreter's strings carry NULs) — use read_file_bytes for binary (v0.1.43)
read_file_bytes ⚡str -> Vec[R, int]Read the whole file as raw bytes — one int (0..255) per byte, binary-safe on every supported backend. interp + C only (v0.1.43, CRC-32 probe). Costs eight bytes per byte; prefer read_bytes
read_bytes ⚡str -> bytesThe whole file as a bytes: one byte per byte, and no NUL hazard. interp + C (v0.1.216, mpng dogfood)
write_bytes ⚡str -> bytes -> unitWrite a bytes to a file. interp + C
bytebuf_newint -> ByteBuf[R]n zeroed bytes, region-bound and mutable. One byte per byte, with random access — which bytes (immutable) and StrBuf (append-only text) leave uncovered (v0.1.218, mpng dogfood)
bytebuf_lenByteBuf[R] -> int
bytebuf_getByteBuf[R] -> int -> intThe byte at an index; out of bounds is an error
bytebuf_setByteBuf[R] -> int -> int -> unitWrite a byte (masked to 0..255)
bytebuf_pushByteBuf[R] -> int -> unitAppend, growing the buffer
bytes_of_bytebufByteBuf[R] -> bytesFreeze a copy, which can then leave the region
bytebuf_of_bytesbytes -> ByteBuf[R]The other way, for editing
print_bytes ⚡bytes -> unitWrite a bytes to stdout in one write(2) (v0.1.480) and with no newline. This is what print_no_nl cannot be: a str is NUL-terminated in the compiled backends, so a zero byte ended the output there and did not on the interpreter. All four backends (v0.1.216, Wasm in v0.1.219)
write_file ⚡str -> str -> unitWrite content to path (overwrite); raises on failure

The FFI byte arena's `bytes` bridge (v0.1.282). tcp_read and friends write into an integer-addressed arena, and mem_to_str cannot bring binary back out — it stops at the first zero byte, which every binary protocol has. Declare these as extern fn alongside the rest of the mem_* family:


extern fn mem_to_bytes: int -> int -> bytes;          // arena ptr, len -> bytes
extern fn mem_copy_bytes: int -> int -> bytes -> int; // arena ptr, offset, bytes -> written

Native (C) and a Wasm component both provide them; scripts/socket_parity.sh requires the two backends to agree. In a plain Wasm build they become an env host import like any other extern, so a host that does not provide them fails at instantiation with a missing-import error rather than a wrong answer.

write_file_bytes ⚡str -> Vec[R, int] -> unitWrite an int vec as raw bytes (each element 0..255) — the write half of the binary path; PPM P6 etc. interp + C only (v0.1.44, Mandelbrot probe)
read_lines ⚡ ★str -> str listRead line by line, returns str list (Phase 19.6; depends on prelude)
file_existsstr -> boolWhether path exists (Phase 19.6; on C native since v0.1.15)
file_deletestr -> boolRemove a file (unlink); true when this call removed it, false for a missing path, a directory, or a refusal. No errno: ask file_exists afterwards for the reason. All five backends; RISC-V through unlinkat (v0.1.622)
file_mtimestr -> floatModification time in seconds; raises if the path is missing (interp + C native)
file_sizestr -> intFile size in bytes (stat); binary-safe length where str_len (strlen) stops at a NUL. interp + C native (v0.1.21)
file_openrwstr -> FileOpen a read/write handle, creating the file if absent and not truncating it. The handle for everything below (v0.1.115, mbtree dogfood)
file_preadFile -> int -> int -> Vec[R, int]Read at most len bytes starting at an offset; a read past the end comes back short rather than padded
file_pwriteFile -> int -> Vec[R, int] -> intWrite a byte vec (each element 0..255) at an offset, extending the file if it writes past the end; returns the count written
file_pwrite_bytesFile -> int -> bytes -> intThe same over bytes, without exploding a byte string into one boxed int per byte (v0.1.222, mraft dogfood)
file_pread_bytesFile -> int -> int -> bytesThe read half of file_pwrite_bytes: len bytes at an offset, straight into a byte string. A read past the end comes back short, like file_pread. This is the one to page a large file with — see below (v0.1.475, medit2 dogfood)
file_fsyncFile -> unitForce the OS to commit this handle's writes to stable storage. The difference between "written" and "durable", and what a store calls at a commit point
file_closeFile -> unitClose the handle
env_var ★str -> str optionFetch env var; None if unset (Phase 19.6; depends on prelude)
args ★unit -> str listThe program's own args (after the script path / binary name); consistent interp ↔ native since v0.1.12. On -rv there is no host to ask: the loader leaves the arguments in RAM and this reads them, so a program given none — or run by a loader that leaves the block alone — sees an empty list rather than an error
runstr -> intRun a command line via the shell, inherit stdio, return its exit code (interp + C native; v0.1.13)
stdin_byteunit -> intOne byte from stdin without blocking; -1 when nothing is ready. read_key blocks, which a device emulator polling a line-status register cannot afford (interp + C native)
tty_rawunit -> unitPut stdin in raw mode: no echo, no line buffering, no software flow control. A no-op off a tty (interp + C native; v0.1.18)
tty_no_signal_keysunit -> unitAlso deliver Ctrl-C / Ctrl-Z / Ctrl-\ as bytes instead of signals — see below (interp + C native; v0.1.476)
tty_restoreunit -> unitPut back the termios the first tty_raw saved
read_keyunit -> strOne byte, blocking; "" at end of input

Why the signal keys are a second call (v0.1.476). tty_raw clears IXON as of that version, and that was a FIX rather than a choice: no full-screen program wants software flow control. With it on, Ctrl-S is XOFF and Ctrl-Q is XON — the line discipline consumes both and the program never sees either byte. The medit dogfood documents Ctrl-S as save and Ctrl-Q as quit; driven under a real pty it drew zero bytes after each and never wrote its file. It shipped that way for two months, because a pipe has no line discipline and every tty test in this project was a pipe.

ISIG is different, and that is why it is tty_no_signal_keys and not part of tty_raw. Clearing it takes Ctrl-C away, so a program that calls it must have a working quit key of its own. An editor needs it — with ISIG set, Ctrl-Z is SUSP and an undo bound to it silently does nothing — and a game that quits on q does not, and should keep the escape hatch. Folding it into tty_raw would take Ctrl-C from every existing TUI to serve the one that asked.

Neither is visible to a piped test, in either direction: through a pipe 0x13 and 0x1a both arrive and everything looks finished. scripts/tty_raw_check.sh drives a Mere program through a real pty and asks whether the bytes it was sent reached it — including the leg that pins Ctrl-C still interrupting under plain tty_raw, so that trade cannot be quietly reversed later.


file_exists "/etc/hosts"            // → true
env_var "PATH"                      // → Some "..."
env_var "BOGUS"                     // → None
read_lines "data.txt"               // → ["line1", "line2", ...]
args ()                             // → ["foo", "bar"] (mere prog foo bar)
run "clang -O2 main.c -o app"       // → 0 on success, nonzero exit code otherwise

Sockets are extern fn declarations rather than builtins — the C backend defines them when a program declares them (native_ffi_names in codegen_c.ml), which is why they take flat-arena offsets rather than bytes. They are native-only in practice.

nametypenotes
tcp_listenint -> intBind a listener on a port (IPv4, all interfaces), SO_REUSEADDR; the fd, or -1. For a chosen address see tcp_listen_at below
tcp_acceptint -> intAccept one connection; the fd, or -1
tcp_connectstr -> int -> intDial host:port; the fd, or -1
tcp_writeint -> int -> int -> intWrite len bytes from an arena offset
tcp_readint -> int -> int -> intRead into an arena offset. See the codes below
tcp_set_timeoutint -> int -> intSO_RCVTIMEO / SO_SNDTIMEO in milliseconds
tcp_closeint -> unitClose the fd

What `tcp_read` returns (v0.1.226, mraft dogfood). A count when it read something, and 0 at end of stream — a peer that closed cleanly, which is information rather than a failure. A negative result says which failure, because with a timeout set these are opposite events for the caller:

meaningwhat a caller does
-1nothing arrived before the deadlinewait again
-2the connection is gonereconnect
-3any other errorusually give up

Before v0.1.226 every failure was -1, and a program that needed the difference had to time the call and ask whether it had failed slowly enough to have been a timeout — inferring a cause from a duration. Every existing < 0 check is unaffected. scripts/tcp_read_codes.sh produces all three rather than describing them.

TLS, both halves (v0.1.338). Also extern fn declarations, and declaring any of them is what makes the C backend link OpenSSL — a program that never mentions TLS does not need it installed.

nametypenotes
tcp_starttlsint -> str -> intClient: upgrade an established fd, SNI, no certificate check
tcp_starttls_verifiedint -> str -> str -> intClient: peer verification against a CA + hostname match
tls_server_initstr -> str -> intServer: load a certificate chain and key, once per process. 0, or negative
tcp_accept_tlsint -> intServer: perform the handshake on an accepted fd. 0, or negative

After either call succeeds, tcp_read / tcp_write / tcp_close on that fd go through TLS with no further change — which is why a plaintext handler becomes a TLS handler by inserting one line.

The server side is two calls rather than one because a certificate is read once and a handshake happens per connection. Folding them together would re-read the files on every accept, and would make "your certificate is unusable" and "this particular client failed" the same result — so a server with a bad certificate would look healthy until a user arrived. tls_server_init returning non-zero is the whole difference.

Until v0.1.338 only the client half existed, so a Mere program could dial a TLS connection but not answer one. Nothing announced this, and nothing worked around it either -- no document in the project mentioned TLS for serving, so there was no workaround to notice. grep tcp_starttls finds TLS and TLS is there. scripts/tls_server_check.sh drives test/tls/https_server.mere with curl and openssl s_client — two TLS implementations that are not ours.

Not on every backend. TLS is C-backend only, client half included: mere -ll and mere -w have no lowering for any of the four, so a program that terminates TLS is a program you build with mere -c.

On Wasm the whole family works — scripts/socket_parity.sh runs the same round trip natively and under wasmtime -S inherit-network=y and compares — with two differences:

Telling them apart means decoding WASI's stream-error variant rather than its is-error bit.

no-op that returned success, so a program that set a deadline blocked forever on the next read. A bounded wait needs the native backend.

Readiness (v0.1.313), also extern fn declarations, C backend. A registered interest set waited on with poll(2); registration persists across waits, so the API survives a later kqueue/epoll implementation unchanged. Documented here since v0.1.475 — it had lived only in the changelog, which is how the medit2 dogfood came to plan a busy-polling event loop before finding it.

nametypenotes
io_poll_newint -> intA pollset id, or -1 when eight are already live. The argument is ignored
io_poll_addint -> int -> int -> intset fd interest; interest is 1 read, 2 write, 3 both. -1 if that fd is already in the set — use io_poll_mod
io_poll_modint -> int -> int -> intChange one fd's interest
io_poll_delint -> int -> intDrop one fd
io_poll_waitint -> int -> intset timeout_ms; how many are ready, 0 on timeout, -1 on error
io_poll_getint -> int -> intThe i-th ready event, packed as fd * 8 + bits (1 read, 2 write, 4 err/hup) — the convention midi_read set
io_set_nonblockingint -> intO_NONBLOCK on one fd

Any fd, not only the ones this runtime handed out. It is poll(2), so fd 0 is a legal member and so is a pipe or socketpair from a C shim of your own. That is what lets one loop wait on the keyboard and a subprocess at the same time, with no busy-wait and no second thread:


let ps = io_poll_new 0;
let _ = io_poll_add ps 0 1;        // stdin
let _ = io_poll_add ps child_fd 1; // a socketpair to a child process
let n = io_poll_wait ps 16;
let ev = io_poll_get ps 0;
let fd = ev / 8;

tcp_read / tcp_write are read(2) / write(2) and are likewise not socket-specific, so the same pair reads whichever fd came back ready.

A listener on a chosen address, `bind(2)`, and the address a socket has (v0.1.555), also extern fn declarations, C backend only. bind(2), getsockname(2) and getpeername(2) take a struct sockaddr *, which no extern can spell, so tcp_listen binds INADDR_ANY and nothing could ask which address or port a socket had. Here an address goes in as the text getaddrinfo(3) reads and comes out as the text getnameinfo(3) writes, and the family comes out as a name, because AF_INET6 is 30 on macOS and 10 on Linux.

nametypenotes
tcp_listen_atstr -> int -> int -> inthost port backlog. getaddrinfo with AI_PASSIVE; each answer in order gets socket, SO_REUSEADDR and bind, and the first that binds is the listener (ruby's TCPServer.new). "" is the wildcard (:: before 0.0.0.0 on a host with IPv6). The fd; -1 with errno kept (the last answer's, when none binds); -2 when the name does not resolve. Ignores SIGPIPE, as tcp_listen does
sock_bindint -> str -> int -> intfd host port: bind(2) an existing socket to a numeric address (no lookup; a name is -2). "" is the wildcard of the socket's own family. No SO_REUSEADDR. 0, or -1
sock_pairint -> inttype (1 stream, 2 datagram, 5 seqpacket): an AF_UNIX socketpair(2) (v0.1.589). Both descriptors in one int, first * 2^20 + second; -1 with errno kept
sock_local_addrint -> strgetsockname(2) as "<family> <port> <address>" — "inet 5000 127.0.0.1", "inet6 5000 ::1", "unix 0 <path>" (the path last: it may hold spaces), "other 0 ". "" on failure
sock_peer_addrint -> strgetpeername(2), the same text; "" with ENOTCONN for a socket with no peer
fd_last_errnounit -> intThe errno the last of these and of the `fd_*` family left, 0 when it succeeded. Per thread

The address text is what ruby's Addrinfo#ip_address says: a scoped IPv6 address keeps its %zone, a v4-mapped one its ::ffff: prefix.

`fd_*` keep errno too. fd_read and fd_write answer -1 on failure as they always have; since v0.1.555 the cause is kept beside the answer. A read of a connection the peer reset is -1 with ECONNRESET — before, a caller could not tell it from a bad descriptor, and a ruby on top of this reported every refused socket write as EPIPE. A refusal the runtime decides itself (a negative fd, an unknown mode or whence, a read of size 0) is EBADF or EINVAL, as the kernel would say. The number is the platform's: EADDRINUSE is 48 on macOS and 98 on Linux.


extern fn tcp_listen_at: str -> int -> int -> int;
extern fn sock_local_addr: int -> str;
extern fn fd_last_errno: unit -> int;

let srv = tcp_listen_at "localhost" 0 128;   // ::1 first on macOS
let at = sock_local_addr srv;                // "inet6 53375 ::1"
let again = tcp_listen_at "::1" 53375 128;   // -1
let e = fd_last_errno ();                    // EADDRINUSE

scripts/sockaddr_check.sh is the gate: every row is a port the kernel gave the probe, compared with what the other end of a connection reports, or a refusal the probe provoked — a port in use, an address the host does not have, a reset provoked by closing a socket with an unread byte in it. The expected errno numbers are read from the host's own <errno.h>, so the transcript holds on macOS and Linux alike.

Resource limits, scheduling priority and advisory locks (v0.1.550), also extern fn declarations, C backend only. None of the five syscalls can be declared by hand: getrlimit / setrlimit move the limits through a struct rlimit *, getpriority / setpriority take an id_t (the emitted int prototype is a "conflicting types" error), and flock is in <sys/file.h>, which an emitted program does not include. The platform's numbers differ exactly here — RLIMIT_NOFILE is 8 on macOS and 7 on Linux, RLIM_INFINITY is 2^63-1 on one and 2^64-1 on the other — so the runtime keeps them and defines its own contracts: a resource is asked for by name, a limit of `-1` is `RLIM_INFINITY` in both directions, the priority selector is 0 process / 1 process group / 2 user, and the lock bits are 1 shared / 2 exclusive / 4 non-blocking / 8 unlock.

nametypenotes
proc_getrlimitstr -> intTake a snapshot of one resource's limits ("CPU", "NOFILE", ...; capitals). 0, or -1 — a name this platform lacks is -1 with EINVAL
proc_rlimit_fieldint -> intRead 0 soft / 1 hard out of the snapshot; no syscall. The limit, -1 for unlimited, -2 when there is no snapshot or the field is not 0 or 1
proc_setrlimitstr -> int -> int -> intname soft hard; -1 is unlimited, any other negative is EINVAL. 0, or -1
proc_rlimit_resourcestr -> intThe platform's RLIMIT_<name> number, or -1 when it has no such resource
proc_rlimit_namesunit -> strEvery resource name this platform has, space-separated, in a fixed order
proc_rlim_conststr -> str"INFINITY" / "SAVED_MAX" / "SAVED_CUR" as the platform defines them, in decimal (2^64-1 does not fit an int); "" when undefined. For showing a value — limits themselves cross as -1
proc_getpriorityint -> int -> intwhich who. The priority — which may be `-1`, so failure is only in proc_last_errno
proc_setpriorityint -> int -> int -> intwhich who prio. 0, or -1
file_flockint -> int -> intfd op. 0 done, `1` refused because another open holds the lock and the non-blocking bit was set, -1 any other failure. A bit outside the four is EINVAL
proc_last_errnounit -> intThe errno the last of these calls left, 0 when it succeeded

Why errno is kept here when `file_*` does not keep it (and fd_* did not until v0.1.555, when it got fd_last_errno). file_* answers -1 and leaves errno alone. getpriority answers -1 for a process whose priority is -1, so its failure is only visible in errno — and errno itself is gone by the time a Mere program can read it, since the next allocation may call into libc. Each call above stores the errno of its own syscall in a per-thread slot; proc_last_errno reads the slot. The number is the platform's (EWOULDBLOCK is 35 on macOS and 11 on Linux), which is why file_flock decides the refusal itself and answers 1 instead of leaving the comparison to the caller. The limits snapshot is also per thread.


extern fn proc_getrlimit: str -> int;
extern fn proc_rlimit_field: int -> int;
extern fn proc_setrlimit: str -> int -> int -> int;

let _ = proc_getrlimit "NOFILE";
let soft = proc_rlimit_field 0;
let hard = proc_rlimit_field 1;     // -1: unlimited
let _ = proc_setrlimit "NOFILE" 256 hard;

scripts/proclimit_check.sh is the gate: every row is a limit the probe set and read back, a priority it raised and read, a lock it took through one open and was refused through another, or a refusal it provoked. The same transcript holds on macOS and Linux, as root and as an unprivileged user.

A signal's disposition, and a refused write to stdout (v0.1.589). signal(2) and sigaction(2) move a disposition through a function pointer and a struct, so no extern can name them; and the runtime's own write path (print_no_nl, print_bytes) dropped a refusal before a program could ask. Same family, same errno slot (proc_last_errno).

nametypenotes
proc_sig_noopint -> intCatch the signal and do nothing — what ruby does to SIGPIPE. Unlike `SIG_IGN`, a handler goes back to the default at `exec(2)`, so a child starts with the default. What was inherited and is not the default (ignored, or a handler) is put back and kept, ruby's rule. 0 installed, 1 kept, -1 refused
proc_sig_defaultint -> intBack to SIG_DFL. 0, or -1
proc_sig_raiseint -> intraise(3) the signal at this thread. 0, or -1
proc_out_errnounit -> intThe errno of the last write print_no_nl or print_bytes could not make, 0 if none since the last ask. Asking clears it
proc_sig_catchint -> int -> int(v0.1.625) proc_sig_catch sig keep: install a handler that only marks the signal; the program takes the marks where it can act on them. SA_RESTART, so a blocking call resumes and the mark waits. keep = 1 is ruby's rule for the signals it handles by default: an inherited SIG_IGN (nohup's SIGHUP) is put back and kept. 0 installed, 1 kept, -1 refused
proc_sig_takeunit -> int(v0.1.625) The lowest marked signal, its mark taken down; 0 when none is marked. Two arrivals before a take are one mark
proc_sig_ignoreint -> int(v0.1.625) SIG_IGN, which a child inherits across exec(2) — ruby's trap(sig, "IGNORE"). 0, or -1

The pair a program that writes to pipes wants: proc_sig_noop 13 (SIGPIPE does not end it, and its children are not left ignoring it), then proc_out_errno () after a write — 32 is EPIPE on every POSIX host — and, to end the way an unhandled SIGPIPE ends a process, proc_sig_default 13 followed by proc_sig_raise 13. print and print_err still go through stdio and are not covered. scripts/procsig_check.sh is the gate, each row in the scene it is for: nothing inherited, SIG_IGN inherited, a pipe whose reader has exited, and a run that must end of the signal (exit status 141).

Positioned file I/O (file_openrw through file_close) works on all four backends: interp and C natively, Wasm over host imports since v0.1.153 (bytes cross in the mere_bytes layout rather than one call per byte), LLVM since v0.1.163. It is the group a paged store or a write-ahead log needs, and it was documented only in the changelog until v0.1.222 — which is how the mraft dogfood came to write its log through the Vec-taking call for a whole slice before noticing.

`file_pread` or `file_pread_bytes` (v0.1.475). Both read at an offset; the difference is what they build. file_pread returns Vec[R, int], which is what mbtree wants — a page it is about to index into as numbers. file_pread_bytes returns a bytes, which is what a reader streaming a large file wants, and the difference is not small. Measured on 208 MB in 256 KiB pages, C backend:

timepeak RSS
file_pread + region3.64 s10.1 MB
file_pread_bytes + region0.36 s1.8 MB
read_bytes (whole file)0.03 s210 MB

Ten times faster and a fifth of the memory, and the third row is the choice this removes: before it, paging a file meant picking between memory and speed. Two things make the difference. file_pread costs eight bytes per byte and builds the Vec one fgetc at a time. And a `Vec` whose data array outgrows a `region` block's arena is not reclaimed by that block at all — around 4-5 MiB of data on the C backend — so a loop reading pages into Vecs retains every page it has read, inside a region block that is otherwise doing its job. A bytes is one flat allocation and does not hit that. Reach for file_pread when the page is numbers; reach for file_pread_bytes when it is a file.

★ Codegen status (v0.1.246 for the first two): print_no_nl and print_err lower on interp + C + LLVM; both were refused by LLVM until it grew a write(fd, ...) for its own panic diagnostic, at which point they were three lines each. On Wasm print_err is refused: it used to write to the same host sink as print, so a diagnostic landed in the program's own output and nothing said so, and the JS host ABI has no second sink to give it. print / print_int / print_bool / read_file / write_file work in all 3 backends (Wasm goes through host imports; scripts/run_wasm.js provides puts / read_file / write_file). print_int / print_bool were the exception until v0.1.190 — this line claimed them for years while only the interpreter had them; C emitted a call to an undefined symbol and LLVM / Wasm refused outright. They now lower on all four (C through printf, LLVM and Wasm through the str_of_int they already had), locked by test/parity/print_int_bool.mere. read_lines / env_var are interpreter-only (codegen would need 'a list / 'a option construction + systematic outside-world access; not yet covered by Phases 22-31). args works on all four backends: C and LLVM read the argc/argv their main was handed, Wasm folds the host's arg_count / arg_get (v0.1.159 for Wasm, v0.1.169 for LLVM). The native-CLI / dogfood builtins run / print_err / file_exists / file_mtime / file_size / tty_raw / tty_restore / read_key / random_int also work on the C native backend (added for the mk / mrog / mwasm dogfoods, v0.1.13-v0.1.21).


let _ = print "Hello";
let _ = print_no_nl "Name: ";
let name = read_line () in print ("Hi, " ++ name);

// File round-trip
let _ = write_file "/tmp/out.txt" "hello lang";
let content = read_file "/tmp/out.txt" in print content;

Value conversion (3)

NameTypeDescription
str_of_intint -> strInteger to string
int_of_str ⚡str -> intParse after trim; raises on bad input
bool_of_str ⚡str -> boolTrim then "true"/"false" only; raises otherwise
float_of_intint -> floatint → float (no precision loss)
int_of_floatfloat -> intfloat → int (truncation)
float_bits_hi ★float -> intThe top 32 bits of a double's IEEE-754 pattern (v0.1.281)
float_bits_lo ★float -> intThe bottom 32 bits (v0.1.281)
float_of_bits ★int -> int -> floatfloat_of_bits hi lo — the inverse of the two above (v0.1.281)
f32_bits ★float -> intThe 32-bit IEEE-754 pattern of the float32 nearest this double. The narrowing is the rounding (round-to-nearest-even); a value with no float32 becomes ±inf rather than wrapping (v0.1.281)
float_of_f32_bits ★int -> floatA float32 pattern back to a double (v0.1.281)

Why the bits come out in two halves. A double's pattern read as one signed int64 does not fit the interpreter's native int — which is OCaml's, and 63-bit — for a large share of ordinary values: -1.5, 1e308, inf and nan all exceed it. A single 64-bit accessor would therefore answer differently on the interpreter than on every compiled backend, for a literal as plain as 1e308. Each 32-bit half is always below 2^32, so there is nothing left to diverge about. Pinned by test/parity/float_bits.mere on all four backends.

These are the primitive the rest is built from: contrib/proto/wire.mere writes a protobuf double as put_double and a float as put_float on top of them, and neither the caller nor the generated codec learns about the split.

str_of_floatfloat -> strFloat to string (OCaml semantics)
float_of_str ⚡str -> floatParse after trim; raises on bad input

str_of_int 42        // "42"
int_of_str "  -7  "  // -7
bool_of_str "true"   // true

String operations (23)

NameTypeDescription
str_lenstr -> intByte length
str_containsstr -> str -> boolSubstring containment
str_starts_withstr -> str -> boolPrefix test
str_ends_withstr -> str -> boolSuffix test
str_countstr -> str -> intNon-overlapping occurrence count
str_index_of ★str -> str -> intFirst position of needle; -1 if not found. Empty needle returns 0 (Phase 19.1)
str_last_index_of ★str -> str -> intLast position of needle; -1 if not found. Empty needle returns the haystack length — it occurs one past the final byte too, and that is the last such position (v0.1.302)
str_split ★str -> str -> str listSplit by delimiter; returns str list. Requires type 'a list = ... declared. Empty delimiter returns a single-element list (Phase 19.1)
utf8_len ★str -> intCodepoint count (a str is bytes; str_len is the byte length). Invalid bytes count as single units (v0.1.38)
utf8_chars ★str -> str listSplit into codepoints — the building block for text processing (v0.1.38)
utf8_atstr -> int -> stri-th codepoint (prelude, on utf8_chars)
utf8_substr -> int -> int -> strCodepoint-indexed substring (prelude)
utf8_revstr -> strCodepoint-wise reverse — str_rev is byte-wise and scrambles multibyte text (prelude)
utf8_widthstr -> intDisplay width (East Asian Width, wcwidth-lite): CJK / fullwidth / emoji = 2 columns, combining marks = 0, halfwidth katakana = 1. utf8_len counts codepoints; terminals draw columns — use this for alignment (prelude, v0.1.45). Fourteen hand-written ranges, and contrib/unicode/width.mere is the generated one — measured over 17,661 code points they disagree on 2,083 (11.8%), so reach for the contrib when a CURSOR has to land where the glyph ends rather than when a table column has to line up (v0.1.477)
pad_rightstr -> int -> strPad with spaces to a display width (table columns, left-aligned); no-op if already wide enough (prelude, v0.1.45)
pad_leftstr -> int -> strRight-align to a display width — numbers in table columns (prelude, v0.1.45)
str_join ★str -> str list -> strJoin with separator. Empty list → empty string (Phase 19.1)
str_compare 🌐str -> str -> intLexicographic -1 / 0 / 1 (Phase 31.0 ported to 3 backends; sign-normalized)
str_repeat ⚡str -> int -> strRepeat N times; raises on N<0
str_replacestr -> str -> str -> strReplace all; empty needle = no change
str_revstr -> strReverse string
str_trimstr -> strStrip leading/trailing whitespace
str_unescape ⚡str -> strDecode \n \t \r \\ \" \/; raises on unknown escape
substring ⚡str -> int -> int -> strs[start:end_excl]; raises on out of range
char_at ⚡str -> int -> strIndex access (length-1 str); raises on OOB
chr ⚡int -> strint in 0..255 to single-char str; raises out of range
ord ⚡str -> intSingle-char str to int code point; raises if length != 1
to_upperstr -> strASCII uppercase
to_lowerstr -> strASCII lowercase
is_digitstr -> boolTrue for single char in '0'..'9'; otherwise false
is_alphastr -> boolTrue for single char that's a letter
is_spacestr -> boolTrue for single char that's space/tab/\n/\r

type 'a list = Nil | Cons of 'a * 'a list;
str_split "a,b,c" ","                          // ["a", "b", "c"]
str_join "-" ["alpha", "beta", "gamma"]        // "alpha-beta-gamma"
str_index_of "hello world" "world"             // 6
str_index_of "hello" "xyz"                     // -1

★ Codegen status: str_index_of / str_last_index_of / str_split / str_join / str_count / str_compare / str_trim / str_starts_with / str_ends_with / str_contains / str_replace / str_repeat / str_rev all work across all 4 backends (Phase 19.1.1 added str_index_of; Phase 22 added str_split / str_join; Phase 26.5 added all Wasm str ops; Phase 31.0 added str_compare; Phase 36 added str_trim / starts_with / ends_with / contains / replace / repeat / rev). not / abs / min / max / clamp / chr / ord / to_upper / to_lower / even / odd / gcd / bool_of_str also reached the 3 backends in Phase 36. The fn (_: unit) -> body wildcard parameter was also parser-fixed in Phase 36.


str_replace "foo bar foo" "foo" "X"           // "X bar X"
substring "hello world" 6 11                  // "world"
char_at "abcdef" 2                            // "c"
"world" |> str_contains "hello world"         // true (pipe + curry)
str_unescape "a\\nb"                          // a + newline + b (3 chars)

Numeric operations (23)

NameTypeDescription
minint -> int -> intSmaller
maxint -> int -> intLarger
absint -> intAbsolute value
signint -> int-1 / 0 / 1
clampint -> int -> int -> intclamp lo hi x restricts to [lo, hi]
pow ⚡int -> int -> intbase^exp by square-and-multiply; raises on negative exp
squareint -> intx x
cubeint -> intx x x
incrint -> int+1
decrint -> int-1
evenint -> booln mod 2 == 0
oddint -> booln mod 2 != 0
gcdint -> int -> intEuclid (handles negatives and 0 correctly)
lcmint -> int -> inta/gcd b; 0 in input → 0
divmod ⚡int -> int -> (int * int)(quotient, remainder); raises on 0 div — the check is its own, because bare / by zero raises on interp and returns 0 on C and LLVM
sum_rangeint -> int -> intSum over lo..hi (Gauss formula, O(1)); halves inside the product, so it is portable over the whole range its result can hold
notbool -> boolLogical negation
bit_andint -> int -> intBitwise AND on the backend's native int width (v0.1.42)
bit_orint -> int -> intBitwise OR (v0.1.42)
bit_xorint -> int -> intBitwise XOR (v0.1.42)
bit_notint -> intBitwise complement; numerically -x - 1 on every backend (v0.1.42)
bit_shlint -> int -> intShift left. A count outside 0..63 gives 0 on every backend (v0.1.512)
bit_shrint -> int -> intArithmetic (sign-propagating) shift right; equals floor division by 2^n (v0.1.42). A count outside 0..63 gives the sign: 0 or -1 (v0.1.512)

★ The shift count, and the four answers it used to have (v0.1.512, Q-039). A count at or beyond the width had no shared meaning: the interpreter defined one, the C backend's << on a signed long long was undefined behaviour (UBSan: left shift of 4611686018427387904 by 1 places cannot be represented in type 'long long'), LLVM IR calls a shift by 64 or more poison, and Wasm masks the count mod 64 by spec — so bit_shl x 70 quietly meant x << 6 there. The contract had already been written down in the RISC-V backend, which chose zero for a left shift and the sign bit for a right one to match what the other backends give; three backends had never implemented it. They do now, a negative count reads as huge-unsigned (out of range) in all four, and test/parity/shift_counts.mere holds them to it.

★ Where the interpreter's integer stops being the compiled one's. The lexer refuses a literal above 2^62-1 on every backend and says why — literals are held in a 63-bit int. Nothing enforces that on a computed value, and 2^62 is exactly where the two widths part company: bit_shl 1 61 is one number everywhere, bit_shl 1 62 is 4611686018427387904 on C / LLVM / Wasm and -4611686018427387904 on the interpreter, whose int is OCaml's 63-bit one. test/parity/int_width_boundary.mere pins that difference rather than hiding it, so moving the boundary is a failure and not a surprise.

This is why there is no `bit_ushr`. A logical right shift is a statement about the top bit, and the top bit is the thing the two widths disagree about: bit_ushr (-8) 1 would be 2^63-4 on the compiled backends, a value the interpreter cannot hold at all. Code that needs unsigned semantics masks explicitly and stays under 2^62 — which is what the varint work here already does with its own lshr7 / lshr8.

★ A transcendental is not correctly rounded by anybody (v0.1.248). exp and log reach C through libm, LLVM through @llvm.exp.f64, and Wasm through the host's Math.exp — and measured, exp -10 is 4.5399929762484854e-05 on three of those and 4.539992976248485e-05 through JavaScript. That is not a bug in any of them. test/parity/exp_log.mere therefore prints exact values only at the points that are exact in binary floating point (exp 0, log 1) and asserts everything else as an identity within a tolerance — which still fails an exp that returns its argument or a log wired to log10, and does not report the C library's build options as a difference.

RISC-V is the one place they are correctly rounded (v0.1.609). RV32IM and RV64IM have no libm, so exp, log, f_pow, sin, cos, tan and atan2 are prelude Mere there, computed in double-double and rounded once — within about 0.52 ulp of the true value, measured against a 50-digit reference; sin / cos / tan reduce by Payne–Hanek past 2^20. The same prelude also answers twenty-one libm functions that a program declares as extern fn with libm's own signature (atan asin acos sinh cosh tanh asinh acosh atanh cbrt log2 log10 log1p expm1 erf erfc tgamma lgamma as float -> float, hypot and fmod as float -> float -> float, ldexp as float -> int -> float): on a host the declaration links libm, on RISC-V it is bound to the prelude instead of refused. Because they are correctly rounded, these can differ in the last bit from a host libm that is not — macOS's is a ulp away on 3% to 49% of points for most of them. test/float/rv_libm_ext.mere holds them to the correctly rounded values.

★ Integer `/` and `%` by zero raise (v0.1.247): division by zero and modulo by zero, catchable with try_or, on the interpreter and the C, LLVM and Wasm backends. It cost a branch per division to make that true, and it was worth it because the alternative was not one behaviour but four: the interpreter raised, the C backend emitted a bare a / b — undefined behaviour in C, which an arm64 build answers with 0 and an x86-64 build answers with SIGFPE — LLVM emitted sdiv, which is undefined in IR and licenses the optimizer to assume it cannot happen, and Wasm trapped with no message at all. INT_MIN / -1 is the other undefined case and wraps now, which is what the interpreter already did.

The `-rv` backend is the exception, and it is measured rather than assumed: under QEMU's virt board, 17 / 0 is -1 and 17 % 0 is 17 there — the RISC-V specification's non-trapping answer. That backend targets bare metal, where there is no stream to write a diagnostic to and no process to exit: the platform's answer is the answer. Float division is IEEE on every backend and keeps giving inf / nan.

★ On the width these are actually computed at (v0.1.245): a builtin can be present on every backend, answer every small question correctly, and still be implemented at a narrower width than the language's int. gcd was a static int __lang_gcd(int, int) in the generated C — gcd 3037000493 3037000493 came back as 1257966803 — and int_of_str on LLVM parsed with strtoll and then truncated the result to i32, so the largest int read back as -1. Both were invisible to every existing test and to host-matrix.md, because the arguments used to probe a builtin were all one or two digits.

test/parity/int_width.mere is the gate for this: every deterministic int builtin, with arguments above 2^31, held to one answer on all four backends. It found the second bug while being written for the first. Values there stay inside ±(2^62 − 1) so that the interpreter's 63-bit int is not itself the difference — and note that an intermediate counts: sum_range used to form a product twice the size of its own answer, which made it portable over only half the range its result could hold.

Float arithmetic (4)

Note (v0.1.44): the infix operators + - * /, all comparisons, and unary - are numeric-overloaded and work directly on floats, on every backend — prefer them. The f_ functions below remain as ordinary function values (useful for passing to higher-order functions). The overload resolves to float only when an operand is concretely float; annotate fn params (fn (x: float) -> ...) in float-heavy code.
NameTypeDescription
f_addfloat -> float -> floatAddition
f_subfloat -> float -> floatSubtraction
f_mulfloat -> float -> floatMultiplication
f_divfloat -> float -> floatDivision (IEEE 754: 0 div is inf/nan)
f_ltfloat -> float -> boolLess than
f_lefloat -> float -> boolLess than or equal
f_gtfloat -> float -> boolGreater than
f_gefloat -> float -> boolGreater than or equal
f_negfloat -> floatUnary minus (Neg is int-only, so use this for float)
f_absfloat -> floatAbsolute value
sqrtfloat -> floatSquare root (NaN for negatives)
floorfloat -> floatFloor
ceilfloat -> floatCeiling
roundfloat -> floatRound
f_min ★float -> float -> floatSmaller (Phase 19.7)
f_max ★float -> float -> floatLarger (Phase 19.7)
f_pow ★float -> float -> floatPower base ^ exp (Phase 19.7)
log ★float -> floatNatural log (Phase 19.7; all 4 backends in v0.1.248)
exp ★float -> floate^x (Phase 19.7; all 4 backends in v0.1.248)
sin ★float -> floatSine (radians; Phase 19.7)
cos ★float -> floatCosine (Phase 19.7)
tan ★float -> floatTangent (Phase 19.7)
atan2 ★float -> float -> floatatan2 y x for angle (Phase 19.7)
fmafloat -> float -> float -> floatfma a b c is a * b + c rounded once -- IEEE-754's fusedMultiplyAdd, correctly rounded like + - * / and sqrt, so every backend gives the same bits (v0.1.534, Q-176). Not what a * b + c means: that rounds twice on every backend at every optimization level (v0.1.315). fma 0.1 10.0 (0.0 - 1.0) is 5.551115123125783e-17 where the unfused expression is 0.0. C: fma(3); LLVM: llvm.fma; Wasm has no fma instruction, so it is computed in software (in integers, about 12 ns a call on node, four times the unfused expression); refused on RV32IM / RV64IM
random_int ★ ⚡int -> intrandom_int n returns int in 0..n-1; raises if n<=0 (Phase 19.7)
random_float ★unit -> floatFloat in [0.0, 1.0) (Phase 19.7)
pifloatπ ≈ 3.14159265 (constant builtin)
efloate ≈ 2.71828183 (constant builtin)

★ Codegen status: the 11 entries added in Phase 19.7 are interpreter-only. Codegen support requires libm linking or per-backend wiring of built-in math functions, planned for a follow-up slice (19.7.1).


f_add 1.5 2.5                    // 4.0
f_div 10.0 4.0                   // 2.5
3.14 |> f_mul 2.0                // 6.28

clamp 0 100 150                  // 100
pow 2 10                         // 1024
gcd 12 18                        // 6
sum_range 1 100                  // 5050
fst (divmod 100 7) + snd (divmod 100 7)   // 14 + 2

Control / error (3)

NameTypeDescription
fail ⚡ ★str -> 'aPanic that unifies with any type
assert ⚡bool -> str -> unitOn false, raises "assertion failed: MSG"
try_or ★(unit -> 'a) -> 'a -> 'aEvaluate the thunk; catch Eval_error and return default
try_or_msg ★(unit -> 'a) -> (str -> 'a) -> 'aThe same catch, with the reason handed to the handler

let safe = fn s -> try_or (fn () -> int_of_str s) (- 1);
safe "42"      // 42
safe "abc"     // -1

if x < 0 then fail "negative" else x

fail is polymorphic, so type inference works at branch merges (if c then fail msg else int_val → int).

★ What an uncaught failure does (v0.1.246): the program writes one line to stderr and exits 1, on every backend. The line is the message raised, tagged fail: when it came from the fail builtin — the tag belongs to the builtin, so a backend's own failures (int_of_str on junk, an out-of-range index) are not tagged, which is what the interpreter has always done. The interpreter additionally prefixes the source file it is running; a compiled binary has none.

None of that was true before. The same program exited 1 on two backends and 134 (SIGABRT) on two others, wrote its diagnostic to stderr on two and stdout on two, tagged the message on three and not on the fourth, and int_of_str on junk named the offending input on two backends and not on the other two. It went unnoticed because the parity harness compared stdout and nothing else, so no parity test used `fail` — none could have passed. test/parity/fail/*.mere is the gate now: exit status, the output written before the failure, and the message, on all four.

Two limitations used to be pinned here. Both are gone, and this entry outlived them by more releases than either took to fix:

ABI provided env.puts and nothing else, so the line landed in stdout there. Adding echo put print_err on the host (scripts/run_wasm.js), and scripts/echo_check.sh now compares that output across all five backends. Under --component the same backend writes through WASI.

on the way out, so statements after it in the same body still ran; the check goes in at emit_instr now, which every call passes through, and test/parity/region_fail_unwind.mere holds all four backends to one answer. The .wasm.expected pin this entry named came down with the fix — which is exactly when the entry should have.

try_or catches all of these on all four backends, including the ones raised inside the backend rather than by fail. What it hands back is the default; `try_or_msg` hands back the reason (v0.1.510, Q-165):


try_or_msg (fn () -> int_of_str s) (fn (m: str) -> let _ = print m in 0)
// int_of_str: "abc" is not a valid int

What the handler receives is the diagnostic line, the bytes the failure would have written to stderr had nobody caught it — so fail "boom" arrives as fail: boom, tag included, and a failure raised inside the backend arrives under its own name. The tag belongs to fail rather than to the printer, which is why that is also the string every backend already had in hand at the catch; taking it off would be four different subtractions and one more thing to keep equal. test/parity/fail/reason_*.mere writes the same failing expression twice, caught and uncaught, and scripts/fail_reason_check.sh compares the two on every backend that can be built.

Two things this does not give you. The message is a str, not a value: a library that wants callers to branch on the kind of failure still has to parse it or return ?t instead. And it is capped at 255 bytes — the C backend has copied into a 256-byte buffer since v0.1.67 and the other backends now copy into one the same size, so a longer message is cut at the same place everywhere rather than in one place and not another.

The copy is the part worth knowing about. The string the raiser built can live in a region the catch jumps out of, so what the handler gets is a copy made at the failure and re-allocated in the catcher's region — not a pointer into the buffer, which a second failure (including one raised by the handler) would overwrite.

try_or_msg is lowered on RV32IM and RV64IM too (v0.1.599), with the same message as the other backends.


Polymorphic helpers (8)

NameTypeWhereDescription
show ★'a -> strbuiltinStringify any value via to_string
fst ★('a * 'b) -> 'abuiltinTuple first
snd ★('a * 'b) -> 'bbuiltinTuple second
id'a -> 'apreludeIdentity function
pair'a -> 'b -> ('a * 'b)preludeTuple constructor (curried)
swap('a * 'b) -> ('b * 'a)preludeTuple swap
const'a -> 'b -> 'apreludeDrop second arg, return first (b is still evaluated — Mere is call-by-value)
flip('a -> 'b -> 'c) -> ('b -> 'a -> 'c)preludeReverse arg order of a curried fn (higher-order)

All eight work on all four backends (v0.1.318). The bottom five were builtins with polymorphic schemes in the typer and implementations in the interpreter, and nothing in any code generator — so they worked under mere file.mere and nowhere else. LLVM and Wasm refused them by name; the C backend emitted a reference to an undeclared mu_pair and left the diagnosis to the C compiler, which reported it in terms of a symbol the author never wrote. As prelude definitions they are ordinary closures: every backend compiles them, partial application and use in value position work without a special case in any emitter, and there is one implementation rather than one per backend to keep in step. (The same migration pow, lcm, divmod and assert already made.)

They are ordinary bindings, so a program may shadow them — test/parity/prelude_shadowing.mere pins that. One Wasm limit is worth knowing: a polymorphic function used at several types and passed as a value is refused there by name ("no single value form on Wasm yet"), which applies to these exactly as it does to any let-polymorphic function of your own.


show 42                          // "42"
show (Some 5)                    // "Some 5"
show [1, 2, 3]                   // "[1, 2, 3]"   (Cons/Nil chains shown as [..])
show [Some 1, None, Some 3]      // "[Some 1, None, Some 3]"

fst (pair "hi" 42)               // "hi"
let always_7 = const 7 in always_7 "anything"   // 7
let sub = fn a -> fn b -> a - b in (flip sub) 3 10   // 7 (= sub 10 3)

JSON, derive-style (5 ★)

Structural JSON, compile-time-specialized per type (no trait machinery), like show. to_json works on all four backends — on LLVM it shares the emitter with show, since the two differ only in literals (v0.1.184). of_json and its siblings are interp / C / Wasm: decoding needs a JSON parser in the target language, and LLVM has no hand-written one.

NameTypeDescription
to_json ★'a -> strSerialize any value to JSON structurally — including containers, which of_json cannot read back (see below)
of_json ★str -> 'aParse JSON into a typed value; fails fast on error (trusted input). Not containers
of_json_opt ★str -> 'a optionSame, but returns None on any error (safe for untrusted input)
of_json_like ★'a -> str -> 'aTarget type from a witness value instead of an annotation (v0.1.183)
of_json_opt_like ★'a -> str -> 'a optionThe non-crashing witness form

The pair is not symmetric: containers write but do not read

to_json serialises a Vec as an array and a Map as an object, and of_json decodes neither:

valueto_jsonof_json round-trip
int / float / str / tuple / option / list / variant / recordyesyes
Vec, Map, StrBuf, ByteBuf, ListBuf, OwnedVec[1], {"1":"a"}, "x"refused when the program is type-checked

Since v0.1.584 an of_json / of_json_opt / of_json_like whose target is a container, or holds one (a record with a Vec field, a tuple with a Map), is a type error at the call -- before, it type-checked and failed when the program read its checkpoint back (of_json: expected a variant value for Vec):


type error: of_json cannot build a Vec: a Vec has a region and an identity, and
decoding makes values, not containers -- `to_json` writes one, but nothing can
read it back as one
help: decode the contents as a list (or a record / tuple of lists) and rebuild
the Vec from it

The witness form says why in its own words: "the witness must be a record, a constructor, a tuple or a scalar — a closure or a handle cannot say what to decode into". A container is a handle with identity and a region, and a decoder would have to choose both.

This matters most where it is least expected: a program that checkpoints its state to resume after a crash finds that the state it wants to save is a Map, writes it happily, and cannot read it back. The way through is to keep the checkpointed state in the shapes that round-trip -- lists of tuples, records, variants -- and rebuild the containers from them on resume. test/durable/fold.mere does exactly that.

Decoding inside a polymorphic function

of_json reads the target type off the call node, which is fine at a use site with an annotation and impossible inside a generic helper: there the node's type is a variable, and the interpreter has no runtime types to resolve it with. So a generic "decode it back" had to name the record type, and every record needed its own copy.

A witness supplies the type instead. The interpreter reads it off the value's runtime shape — a record carries its type's name — and the compiled backends read it off the witness's static type, which is the same variable the result unifies with:


let with_field = fn (rec_) -> fn (name: str) -> fn (v: str) ->
  ... of_json_opt_like rec_ (rebuilt_json) ...

The witness is a value the caller already has whenever this comes up: replacing one field of a record means holding the record. contrib/schema is this, and examples/claims generates its whole form from it.

A polymorphic record still needs the annotation — a value carries its type's name but not its type arguments, so a witness cannot describe Box[int].

The of_json result type comes from the use site — annotate the expression: (of_json s : T). A JSON object maps to a record's fields (by name), an array to a list or tuple, null/value to option (None / Some), and a string / {"Ctor": payload} to a variant. to_json uses the same mapping in reverse, so (of_json (to_json x) : T) == x.

A repeated object key is refused (v0.1.303). {"id":1,"id":2} does not decode: of_json fails and of_json_opt answers None, at any nesting depth, because the check belongs to the parser rather than to the decoder generated for a particular type. Every decoder used to accept it and keep the first value, which was not a decision anyone made — it fell out of looking up an assoc list built in document order. Go's encoding/json v1 kept the last for an equally accidental reason, and Go 1.27's v2 stopped picking. Two implementations resolving the same bytes to different values is the argument: there is no right one to choose, so the input is rejected.

Invalid UTF-8 inside a string is refused too (v0.1.306). The three hand-written parsers process no \uXXXX escapes, so a decoded string is exactly the raw bytes between the quotes — and until v0.1.306 nobody looked at them. The validator (shortest form, no surrogates, max U+10FFFF) runs in the parsers' string path, so keys and nested strings are covered without a per-type rule. Go 1.27's encoding/json/v2 made the same call; v1 silently rewrote bad bytes to U+FFFD, which is a lossy edit nobody asked for. This is the parser's rule, not str's — utf8_len still counts an invalid byte as one unit on purpose, because a str already in memory has no better answer.


type User = { id: int, name: str, bio: str option };
to_json (User { id = 1, name = "ada", bio = None })
                                 // {"id":1,"name":"ada","bio":null}
let u = (of_json body : User);   // fails fast if body is malformed
match (of_json_opt body : User option) with
| Some u -> u.name               // decoded
| None   -> "bad request"        // malformed / missing field — no crash

Comparison, derive-style (v0.1.11)

== / != (structural equality) and < <= > >= (structural ordering) are compile-time-specialized per operand type — the same no-trait mechanism as show / to_json. Both work on interp / C / Wasm.

lexicographically).

order; lists compare element-wise (a shorter prefix is smaller); variants order by declaration order (the constructor listed first is smallest), then by payload. All backends agree byte-for-byte, so a value sorts the same under the interpreter, a native binary, and Wasm.


(1, 2) < (1, 3)                        // true  (tuple, lexicographic)
[1,2] < [1,2,3]                        // true  (prefix is smaller)
type C = Red | Green | Blue; Red < Blue // true  (declaration order)
list_sort_by (fn (a: float) -> fn (b: float) -> a < b) [3.1, 1.2]  // [1.2, 3.1]

Honest edges. float uses a total order where NaN sorts as least. Comparing two functions is defined but meaningless (they order as equal). The bare default list_sort still bakes in an int comparison — its comparator's type variables default to int, the same rule that keeps fn a -> fn b -> a < b monomorphic — so sorting a non-int list needs list_sort_by with an annotated comparator (as above). A fully-polymorphic list_sort over any orderable element would need ad-hoc-polymorphism resolution (deferred).


Loop helper (1 ★)

NameTypeDescription
iter_n ★int -> (unit -> unit) -> unitApply thunk N times (side-effect loop); no-op when N≤0

Capability (2 + 2 builtin record types)

Used by the effect system (see effects.mere). The Logger and Metrics cap types are pre-registered as builtins. Users can also override with their own type Logger = ....


type Logger  = { info: str -> unit, warn: str -> unit, error: str -> unit };
type Metrics = { inc: str -> unit, record: str -> int -> unit };
NameTypeDescription
mk_loggerstr -> LoggerCreate a prefixed Logger. Each field prints as prefix [LEVEL] msg
mk_metricsunit -> MetricsCreate a Metrics. inc / record print as [METRIC] ...

let lg = mk_logger "app" in
{ lg.info "started";
  lg.warn "slow query";
  lg.error "abort" }

let m = mk_metrics () in
{ m.inc "users";
  m.record "latency_ms" 23 }

For a complete cap-passing example see examples/effects.mere.

Raw memory, CSRs, traps and tasks (14, RV32I bare-metal only)

A Raw is a window onto physical memory — the one capability that is not a record of functions, because its operations lower to load and store instructions. It is the escape hatch a device driver needs, and it is a value rather than an ambient builtin so that "this function cannot touch raw memory" is something you read off a signature.

Raw is opaque: nothing constructs one, and there is no function that mints one. The only source is the argument mere -rv --bare hands to the program's top-level main, and raw_window can only narrow it. Offsets are relative to the window, so a driver holding a UART window cannot express an address outside it; every access bounds-checks the offset, and widening faults.

NameTypeDescription
raw_windowRaw -> int -> int -> RawA window over [off, off+len) of another. Faults if that is not inside it
csr_readint -> intA machine CSR by number — the number must be a literal (it is an immediate field of the instruction)
csr_writeint -> int -> unitWrite a machine CSR. Not behind a capability: a CSR has no base and length to narrow, and the hardware's privilege modes are what separate a kernel from a user process
raw_lenRaw -> intIts length — so a kernel can partition a window it was handed without hardcoding the runtime's geometry
raw_baseRaw -> intA window's base as a number. Not authority — touching anything still needs a window — but a stack pointer is an address and hardware wants the number
trap_saveRaw -> RawThe trap trampoline's 31-word register save area. A context switch is a copy through this: outgoing registers to a TCB, incoming registers back
machine_scratchRaw -> RawReserved RAM the runtime is not using — where task stacks come from. A bare program owns no fixed address of its own: the heap grows up from 2MB and the stack down from the top
closure_code(unit -> unit) -> intA closure's entry point. A task IS a closure, so starting one means building a context whose PC is this
closure_env(unit -> unit) -> intIts environment — the value the first argument register must hold when that PC is entered. ABI knowledge, which a kernel has
set_trap_handler(int -> int) -> unitInstall a trap handler. The argument is mcause; the result is the PC to resume at. Anything else (mepc 0x341, mtval 0x343) is a csr_read away. A closure, not a named function: a handler needs the machine capability to do anything useful and an interrupt has no caller to hand it one, so it captures instead. Codegen emits the trampoline that saves the register set and returns with mret
raw_peek8Raw -> int -> intThe byte at that offset
raw_peek32Raw -> int -> intThe 32-bit word at that offset (unsigned; a device register is 32 bits whatever the CPU width)
raw_peekwRaw -> int -> intThe i-th machine WORD (cell-indexed, xlen-wide) — for walking a register save area at either width
raw_poke8Raw -> int -> int -> unitStore a byte
raw_poke32Raw -> int -> int -> unitStore a 32-bit word (32 bits by name, either width)
raw_pokewRaw -> int -> int -> unitStore the i-th machine word (cell-indexed, xlen-wide)

let putc = fn (uart: Raw) -> fn (c: int) -> raw_poke8 uart 0 c;

let main = fn (mach: Raw) ->
  let uart = raw_window mach 0x10000000 256 in    // the UART, and nothing else
  putc uart 65;

A context switch needs no new mechanism: the trampoline saves the interrupted register set to the area trap_save hands back and restores from it before mret, so a handler swaps tasks by copying through it and returning the incoming task's PC. Switch every register, gp included, and give each task a heap arena of its own (carve it from machine_scratch — heap up from the bottom, stack down from the top). Sharing one heap looks workable until a region R { } in one task rolls the bump pointer back and frees what another task allocated meanwhile; the rule that survives is that contexts share gp only if they genuinely share a heap, and a context that uses regions must not. See examples/riscv_bare_sched.mere.

Device MMIO sits above any RAM (the UART data register is at 0x10000000, the address QEMU's virt machine uses), so a device address does not move when --ram does. On every other backend these refuse: there is no honest physical address in a hosted process. See examples/riscv_bare_uart.mere.


Channel receive: which of the three blocks

All three block. They differ in what ends the wait.

returns whenuse it for
channel_recva message arrivesa worker that runs until the process does
channel_recv_opta message arrives, or the channel is closed and drained (None)a worker loop that must terminate: channel_close ends it
channel_recv_timeouta message arrives, or the deadline passesthe only one that answers "is there something now"

All three, and channel_close, work on the interpreter, C, LLVM and Wasm (LLVM and Wasm since v0.1.600). On every one a send on a closed channel and a channel_recv on a closed, drained one fail -- catchably, with the same message.

channel_recv_opt is the one that gets misread, because a name ending in _opt reads like a try-receive. It is not: the implementation is while (len == 0 && !closed) cond_wait, and None means closed, never empty. A thread polling two channels with it parks in the first one and never reaches the second — which is a deadlock that looks like a delivery bug, since the first channel's traffic keeps waking it up just often enough to seem alive. Use channel_recv_timeout to poll, or put both kinds of message on one channel. contrib/http/sse_native.mere took the second route and says why.


Coroutines (v0.1.543; values v0.1.561)

A coroutine is a second stack on the thread that made it. Nothing runs in parallel: a transfer hands the thread over, with a value, and comes back when something hands it back.

type
coro_new(Coro['m] -> 'm -> CoroExit) -> Coro['m]a suspended coroutine; nothing runs yet. The body is handed its own handle and its first message
coro_new_sizedint -> (Coro['m] -> 'm -> CoroExit) -> Coro['m]the same, on a stack of at least that many bytes (v0.1.566): rounded up to a power of two, 64 KiB at least, 1 GiB at most
coro_transferCoro['a] -> 'a -> Coro['b] -> 'bcoro_transfer c v me: suspend me (the one running), resume c with v; returns what me is handed next
coro_exitCoro['a] -> 'a -> CoroExithow a body ends: coro_exit c v gives c the value v, and c runs when the body returns
coro_rootunit -> Coro[unit]the thread's own stack
coro_switchCoro[unit] -> unitcoro_transfer c () <running>, dropping what it is handed later
coro_scan_intsCoro['m] -> int -> int -> (int -> unit) -> unitthe integers in a range that a suspended stack can still reach (see Q-182)

'm in Coro['m] is the type of what that coroutine receives -- its first message and everything it is handed after. Transfer is symmetric: there is no implicit parent to return to, and any coroutine may transfer to any other.


let root = coro_root ();
let cell = map_new ();   // where the generator finds its consumer
let rec gen_loop = fn (me: Coro[int]) -> fn (i: int) -> fn (want: int) ->
  if want < 0 then coro_exit root ()
  else gen_loop me (i + 1) (coro_transfer (map_get cell 0) (i * 10 + want) me);
let gen = coro_new (fn me -> fn (first: int) -> gen_loop me 0 first);
let consumer = coro_new (fn me -> fn (_start: int) ->
  let a = coro_transfer gen 1 me in    // 1
  let b = coro_transfer gen 2 me in    // 12
  let _ = print_int (a + b) in
  coro_exit gen (0 - 1));
let _ = map_set cell 0 consumer;
let _ = coro_transfer consumer 0 root;  // prints 13
0

Why the types hold. What a coroutine receives is typed by its own handle: the 'b of coro_transfer is the 'm of me, and the runtime checks that me is the one running (a fail naming it otherwise). Handles come only from coro_new, which types one by its body, and coro_root, so everything that can send to a coroutine was checked against the type it receives. That is also why there is no coro_self (removed in v0.1.561): "the one running" has no type anyone could check it against. A body gets its own handle as its first argument instead. A let bound to coro_new .. is not generalised, as for a channel.

A message is a scalar for now: int, bool, float, unit or a Coro. It travels in one word and the receiver needs no copy; anything boxed would have to be copied into the receiver's region, and a coroutine running in the default region would then allocate there on every resume. Send an index into something both sides can see. The root receives only unit, so a consumer that is handed values runs in a coroutine of its own (as above).

What each stack keeps as its own: the current region, the blocks it has open, its innermost try_or, and (compiled) its stack bounds. So a fail inside a coroutine is caught by a try_or on that coroutine's stack or by nothing -- never by one the switching stack had open, which would be a jump onto a stack that is not running. Uncaught, it ends the program, as on a thread.

What is refused, and where:

or captured by spawn (type error). Transferring to another thread's coroutine is therefore not expressible.

shape before v0.1.561 (fn () -> .. returning the coroutine to run next).

a type error. The body's env is the coroutine's own copy -- the coroutine may run after the block has ended -- and a container's copy is the same handle. A captured value (a str, a record) is copied into the default region with it and is fine.

the one running, and a body that ends by naming itself or a finished one, fail with a message (the first one catchably).

What a finished coroutine keeps: nothing, on C (v0.1.558). A handle is a slot and a generation, not an address: when a finished coroutine is reaped its record, stack, saved state and env are freed and its slot is reused under a new generation, and a handle from before still answers "has finished". A program that makes a coroutine per connection holds its peak of live ones, not every one it made -- measured flat over 100k connections. Two ways to keep it that way:

bits -> serve me fd bits). That lambda's env is made as the coroutine's own. A closure made elsewhere -- coro_new (serve_curried fd) -- is made where the currying happens, in the current region, and outside a region block that is the default one: the coroutine owns a copy, and the original stays.

a saturated call allocates nothing on C.

A handle held across 2^24 reuses of one slot would name a later coroutine. LLVM is the same since v0.1.565: a handle is a slot and a generation there too, and a body written as a lambda at the coro_new has its env as its own (a million finished coroutines: 1.5 MiB on both).

Stacks are pooled (v0.1.566). A finished coroutine's stack is kept for the next one of the same size, up to the most coroutines this thread has had alive at once (at least 16), so making one is no system call; a server with fifty connections in flight no longer maps and unmaps a stack per connection (2000 rounds of 64 coroutines: 1.0 s before, 0.05 s now). The pool never holds more stacks than were alive together, so it does not raise a program's peak memory; it only stops giving it back below that. coro_new_sized is for many small coroutines: only the pages a stack touches are resident (one 16 KiB page for a coroutine that does little, on either backend), so a smaller stack saves address space rather than memory -- and on Linux vm.max_map_count (65530) limits a process to about 32k live coroutines whatever their size, two mappings each. Overflowing a small stack is named like any other.

Backends: the interpreter, C and LLVM (v0.1.544). On both native backends a coroutine's stack is what stack in mere.toml asks for (as for spawn), otherwise 8 MiB, reserved rather than committed, below a guard page -- an overflow is named like any other. C switches with a few lines of assembly (arm64 and x86-64; another target is a C compile error naming the builtin). LLVM IR is not specific to a machine, so there the switch is _setjmp / _longjmp between existing stacks and llvm.stackrestore onto a new one.

RISC-V (v0.1.623, both widths) has no virtual memory to reserve a stack in, so its coroutines share one region and COPY: every coroutine runs with its stack ending at the same address, and a switch writes the stack being left out to a buffer of its own and the one being entered back to the same addresses (nothing that points into a stack ever moves). A suspended coroutine costs the words its stack holds -- ten thousand suspended ones fit in a few megabytes -- and a switch costs a copy of the two stacks' live parts. A program that makes or switches to a coroutine is laid out differently: a sixteenth of RAM (at least 256 KiB) is the main stack, the region below it (a sixteenth too, or --coro-stack <MB>) is the coroutines', coro_new_sized's size is how far below that region's top a coroutine may reach, and every function entry checks the running stack's floor (an overflow is named like any other). A program without coroutines is laid out and checked as before. The runtime is the RV prelude's, in Mere. A region block's rollback cannot reach past a switch (each switch raises the high-water mark), and what a compaction frees while coroutines exist waits until a walk of every stopped stack shows nothing reaches it (v0.1.624; C's pins, checked a megabyte at a time). On one bump heap shared by every stack, a program that switches every few statements reclaims little: no rollback reaches past a switch.

On C, a store compacted (map_compact, vec_compact, map_recycle) while a coroutine is suspended keeps any arena that coroutine's stack still points into, and frees it at a later compaction (v0.1.547). "Points into" is what the stack REACHES (v0.1.589): its own words, and the pointers inside the nodes they point at -- as far as they go in the coroutine's own regions, three hops in anyone else's, the walk coro_scan_ints makes. A retired arena is not tried again while one of the coroutines that pinned it is the one running. And since an arena under a pin cannot be given back anyway, map_compact and vec_compact leave such a container where it is rather than copying it beside the old arena (a recycle still moves, because it must empty the map). (LLVM has no compaction builtins.)

What a stack still holds: coro_scan_ints (v0.1.549)

type
coro_scan_intsCoro -> int -> int -> (int -> unit) -> unitevery integer in [lo, hi) that coroutine c's stack can still reach, handed to f once each

For a program that keeps its own handles -- integers indexing tables of its own, as an interpreter written in Mere does -- and collects them itself. A coroutine that stops in the middle of an expression can hold a value in a C variable of its own stack that none of the program's roots names: something it took out of a table just before it stopped. coro_scan_ints c lo hi f reports the numbers in that range the stack can reach, so the program can keep those entries.

(coro_root () on the root, or the handle its body was given) from the call up, its registers included. A finished or never-started one reaches nothing.

that is allocated) is followed, and the 16 words from there are read the same way: through the coroutine's own regions (the ones open on its stack) as far as the pointers go, and at most three hops into any other block -- a number inside a node inside a node. A word inside no live block is never read through.

that only looks like a handle is reported too, which costs the caller an entry it keeps one collection longer. The contract is a superset.

counters, lengths and loop indices are all integers; if the handles are small the reports are full of numbers that are not handles, and what they keep can keep more (a kept table can hold further handles). Handles numbered from 2^48 are above every user-space address and every small number: ask for that range and the reports are the handles.

ask is answered from them when its range is inside what was read (from lo up with no end when lo is at least 2^47; up to 2^32 when hi is at most 2^32). f is called after the whole stack has been read, so it may allocate.

Backends: C reads the stack. The interpreter and LLVM hand over every integer in the range -- the superset the contract allows, because neither keeps a list of live blocks to read against (an interpreter's suspended body is an OCaml continuation). RISC-V (v0.1.623) reads the stack, as C does -- a stopped coroutine's from the buffer its stack was copied to, the running one's after spilling the saved registers -- and follows the same hops, without a limit above the innermost open region block's mark (every allocation shares one bump pointer, so "its own regions" is that block); it keeps no findings between asks. Wasm refuses it by name, as it refuses coroutines.

One difference: on LLVM a coroutine made inside a region R { } block (lexically, or by a function called from one) fails by name. LLVM closures carry no env copier, so the body's env would be released with the block; C copies it out. Make the coroutine outside the block.

Wasm and RV refuse by name: core Wasm cannot switch stacks, and the bare-metal runtime has one stack. See scripts/coro_check.sh.


Virtual clock (v0.1.305)

MERE_VIRTUAL_CLOCK=1 makes the interpreter advance a virtual clock instead of waiting: when every live thread is parked, time jumps to the earliest pending deadline (Go 1.27 testing/synctest's rule). Timer firing order becomes deterministic and a minute of timers costs milliseconds — this is for tests. Covered: channel_recv / channel_recv_opt / channel_recv_timeout, sleep_ms / sleep, join, and time (reads the virtual clock; only differences are deterministic). Not covered: par_map's internal joins and OS-level waits — a thread in those counts as running, so the clock refuses to advance past it. All-blocked with no timer pending fails naming the deadlock instead of hanging. Off by default; the C backend is untouched. See scripts/virtual_clock_check.sh.


What a spawned thread may touch (v0.1.574)

spawn e is checked twice. What e mentions directly is classified by type: a value that is Sync is shared, one that is only Send is moved into the thread, and anything else -- Map, Vec, StrBuf, ByteBuf, ListBuf, Coro -- is refused. Then every function e mentions whose definition the program text shows is followed to what IT mentions, and the spawn is refused if that walk reaches a value that is not Sync. The error names the path:


type error: this thread reaches `tbl` : Map['a, int, int] (line 1) through
w > fill > put, and a Map is not safe to share between threads

Two things are allowed through on purpose:

program is the first argument of a read (vec_get, vec_len, map_get, map_has, map_len, bytebuf_get, ...), threads may read it together -- a lookup table built at start-up is the case. Hand it to a function, alias it, or write it once anywhere, and it is not read-only any more.

received over a channel, a call's result. A library that spawns the handler it was given (http_serve_mt) cannot see what the handler touches; that is left to the run-time check below.

At run time (v0.1.575), a Map, Vec, StrBuf, ByteBuf or ListBuf remembers the thread that made it, and a WRITE from any other thread stops the program, on the interpreter, C and LLVM alike:


vec_push: a Vec made by one thread was written by another -- a Vec is not safe
to share between threads; give it one owner thread and send it messages over a
Channel

A read by another thread freezes the container (v0.1.582). The first time a thread other than its owner reads one, the container becomes shared, and a shared container is read-only from then on -- the owner's next write fails too:


map_set: a Map another thread has read was written -- once a second thread
reads a Map it is shared, and a shared Map is read-only; give it one owner
thread and send it messages over a Channel

That is what stops the owner rewriting a Map (or compacting a Vec) under a reader, which used to crash or read garbage. A table built once and then only read by many threads still works, and its owner may still read it. The rule refuses one honest shape: threads read a table, are joined, and then the owner updates it. Build a new table for the next phase instead, or keep the table in one thread and ask it over a Channel.

A program that never spawns pays one load of a global per read and per write; one that does pays a thread-id comparison.

`OwnedVec` is moved, not shared: the closure given to spawn takes the OwnedVecs it captures with it, and they belong to the new thread. Any other thread's use fails (owned_vec_push: an OwnedVec belongs to one thread ...) -- one reached through a closure handed to a library that spawns it, say, where nothing moved it.

`sync type` vouches that a type is safe to share, and is refused when the type holds a builtin container (Map, Vec, StrBuf, ByteBuf, ListBuf, OwnedVec, File, ...): those have no lock for the marker to stand for.

`--lib` turns the owner check off: a host may call in from whichever thread it likes. A library whose entry file has a top-level container -- module state -- refuses a call that overlaps another thread's call (MERE_FAIL, with the sentence called while another thread is inside this library), so that state is never written by two threads at once. A library without one still takes concurrent calls.

The usual fix is to give the shared value one owner thread and send it messages over a Channel -- an append is a message, a read sends a channel for the answer to come back on.


When a spawned thread fails (v0.1.586, v0.1.590)

A thread's failure does not end the program. It is one line on stderr when it happens, whoever holds the thread's handle, and the exit status stays the main thread's -- the same on every backend:


mere: thread 1 failed: fail: boom

`join h` raises the failure again, in the joining thread. Uncaught there, it is the program's failure (fail: boom, exit 1); try_or takes it like any other. par_map joins each element's thread before taking its result (v0.1.590), so a failure in its function is raised in the caller.

(v0.1.586 printed the line at exit for a thread nobody had joined or detached; a daemon never exits, so a handler whose handle was dropped failed in silence.)

A worker POOL should catch a handler's failure itself (try_or around the handler), or each failure costs it a worker; contrib/http's pools do.

⚠ A channel does not know who sends on it unless it is told, so a thread that waits on a channel for a message a failed thread was going to send waits for good -- before v0.1.586 the failure ended the whole program instead. Two ways out:

and the failure reaches the waiter.

(a long-lived producer, an actor): it registers thread h as a sender of ch. Once every registered sender has finished and the channel is empty, a receive stops waiting as it does on a closed channel: channel_recv fails with channel_recv: sender thread 2 failed: fail: boom (the first registered sender that failed, with its message) or channel_recv: every sender has finished and the channel is empty, and channel_recv_opt / channel_recv_timeout answer None. A channel with no sender registered waits as it always did. Register every thread that sends, or the receive ends while one of them is still going to.

let results = channel_new (); let h = spawn (fn (u: unit) -> produce results); let _ = channel_sender results h;

The same on the interpreter, C, LLVM and Wasm (test/parity/channel_senders.mere). The compiled backends notice a sender's end within 50 ms of it; a Wasm channel takes at most 256 senders.

Threads are numbered in the order they were spawned. Before v0.1.586 the C and LLVM backends ended the process from the failing thread (an exit status that changed from run to run), Wasm printed the failure, and the interpreter said nothing. See scripts/thread_fail_check.sh.

Thread-leak report (v0.1.304)

MERE_THREAD_REPORT=1 makes a program print, at exit, the threads that were neither joined nor detached, and what each was doing. It goes to stderr and is off unless the variable is set, so it never changes what a program prints.


mere: 1 thread(s) neither joined nor detached at exit
  thread 1: blocked on channel_recv

detach is what marks a thread as meant to block forever (a server's accept loop), so a detached thread is never reported. The interpreter says what a thread was blocked on; the C, LLVM and Wasm builds (v0.1.586) say still running, finished, never joined or died: <message>, never joined. Not covered: the main thread. See scripts/thread_leak_check.sh.


Region stats (v0.1.307)

MERE_REGION_STATS=1 makes a C-backend binary print, at exit, each arena's block count, total block capacity and cumulative bytes handed out. Stderr only, off unless the variable is set.


region-stats default: blocks=2 cap=272629776 alloc_total=272435592
region-stats named region R loop: arenas=42 alloc_total=62376872 peak_cap=3145728
region-stats named-total: sites=1 alloc_total=62376872

The named lines arrived in v0.1.319. Before them the meter saw the default region only — so a program that did its allocating inside region R { } or region R loop, which is to say a program managing its memory deliberately, read as a few hundred bytes. The number that exists to replace peak RSS was blind to the construct that exists to manage memory, and peak RSS was the only figure left for exactly those programs.

Named arenas are created and destroyed during the run, so there is nothing to walk at exit: each is charged to its source name as it is released. Hence arenas= (how many were opened) and two different totals — alloc_total is everything that arena's generations ever asked for, and peak_cap is the most any single one held at once. A region loop raises the first and lowers the second: it pays a copy per generation to bound what is resident. Only the second half of that trade was measurable before.

Two limits worth knowing. An arena still live at exit is never released and so is never charged — in practice a map's private arena between compactions. And the table holds 32 distinct names; past that, releases are counted and reported as a region-stats WARNING line rather than dropped silently, because an undercount looks exactly like an improvement.

Capacity and allocation are functions of the program, not of the machine — unlike peak RSS, which is quantized to powers of two and stops reproducing above a few GB — so a gate can hold their ratio to a bound: scripts/region_slack_check.sh does exactly that. cap far above alloc_total names arena slack (stranded block tails, an inflated doubling base); alloc_total itself is the number a collector has to attack. The report fires through main's epilogue or through exit(), whichever comes first, and only once.


Which lines filled the default region (--region-sites, v0.1.570)

The default region is never given back, and it is where a container goes when nothing decided its region — typically a buffer a function makes and uses but does not return, so it does not appear in the function's type. Compile with --region-sites and the same meter names the source line of every such container, most bytes first:


mere -c --region-sites main.mere > p.c && cc -O2 p.c -o p
MERE_REGION_STATS=1 ./p
region-stats default: blocks=12 cap=17175674880 alloc_total=11554780168
region-stats default-sites: 29 lines, alloc_total=11552355112 of 11554780168
region-stats default-site mgz/inflate.mere:341: alloc_total=8347724352
region-stats default-site mgit/store.mere:210: alloc_total=1829840680

Each line counts what its containers asked for: the struct, the storage, every growth and every copy stored into them. Where the memory lives does not change — each line's region forwards every allocation to the default one — so the default: total is the same with and without the flag, and so is the output. Strings and other values allocated in the default region directly are not lines here, which is the gap between the two totals.

The usual answer to a line at the top is a region block around where the container is made and used. The example above is mgit reading 17,500 git objects: one 64 KiB inflate window per object, 8.3 GB of 11.5. Inside a region W { } of its own the default region was 3.2 GB, with the same output. C backend only; without the flag the emitted calls are exactly what they were.


System / constants (4)

NameTypeDescription
timeunit -> floatUnix epoch seconds (gettimeofday). For benchmarks / timestamps
exit ★int -> 'aExit the process with an exit code (never returns; polymorphic return)
int_maxintMax int value (OCaml runtime dependent; 2^62-1 on 64-bit) — constant builtin
int_minintMin int value — constant builtin

let start = time () in
{ run_heavy_computation ();
  print ("elapsed: " ++ str_of_float (f_sub (time ()) start) ++ " sec") }

if config_invalid then exit 1 else continue ()

iter_n 3 (fn () -> print "===")   // prints === three times

The named builtins, alphabetical (122)

Not the whole 221: the vec_* / owned_vec_* / strbuf_* / map_* families below are registered separately. id / pair / swap / const / flip left this list in v0.1.318 — they are prelude definitions now (see Polymorphic helpers above). The heading said 129 while the list held 127, before that.


abs args assert atan2 bit_and bit_not bit_or bit_shl bit_shr bit_xor
bool_of_str ceil char_at chr clamp
cos cube decr divmod e env_var even exit exp f_abs f_add
f_div f_ge f_gt f_le f_lt f_max f_min f_mul f_neg f_pow
f_sub fail file_delete file_exists float_of_int float_of_str floor
fst gcd incr int_max int_min int_of_float int_of_str
is_alpha is_digit is_space iter_n lcm log max min mk_logger
mk_metrics not odd ord pi pow print print_bool
print_err print_int print_no_nl random_float random_int
closure_code closure_env csr_read csr_write machine_scratch
raw_base raw_len raw_peek32 raw_peek8 raw_poke32 raw_poke8 raw_window trap_save
stdin_byte
read_file read_file_bytes read_line read_lines round show sign sin snd sqrt
square str_compare str_contains str_count str_ends_with
str_index_of str_join str_len str_of_float str_of_int
str_repeat str_replace str_rev str_split str_starts_with
set_trap_handler str_trim str_unescape substring sum_range tan time
to_lower to_upper try_or write_file write_file_bytes

Q-010 collection builtins (vec_* / owned_vec_* / strbuf_* / map_* / len) are registered builtins outside this table; see language-reference / tutorial.

The names in those families, and the bytes ones

Listed because delegating a family to another document only works if something says who is in it. scripts/doc_coverage_check.sh compares this against mere --dump-builtins; before it existed, twenty-five names were reachable from a program and mentioned nowhere a reader would look -- the whole binary type among them, and one Vec operation that appeared in no file under docs/ at all. They are not named in this paragraph on purpose: the gate asks only whether a name is spelled somewhere here, so a sentence about a missing builtin would satisfy it and hide the very row it was describing.

`bytes` — the immutable binary type (a distinct type from str: length-prefixed, NUL-safe):

builtinsignaturenotes
bytes_lenbytes -> intlength in bytes
bytes_getbytes -> int -> intthe byte at an index, 0-255
bytes_slicebytes -> int -> int -> bytesoffset and length
bytes_concatbytes -> bytes -> bytes
bytes_of_str / str_of_bytesstr -> bytes / bytes -> strthe two directions across the text boundary
bytes_of_hex / hex_of_bytesstr -> bytes / bytes -> strlowercase hex, no separators
bytes_of_vec / vec_of_bytesVec[R, int] -> bytes / bytes -> Vec[R, int]the bridge to the integer view

`f64x2` / `f32x4` / `u8x16` — the 128-bit SIMD types (Q-109; two doubles, four floats, or sixteen bytes in one value. C: the compiler's vector extension; LLVM: <2 x double> / <4 x float> / <16 x i8>; Wasm: v128, boxed only when it escapes; RV32IM / RV64IM: u8x16 only, as a 16-byte box through the RISC-V Vector extension (v0.1.426) -- f64x2 and f32x4 are refused there, which is the floating-point unit it does not have. Not comparable with == or < -- compare lanes. show prints f64x2(a, b) / f32x4(a, b, c, d) / u8x16[32 hex digits], to_json [a, b] / [a, b, c, d] / a hex string.

`f32x4` is the one whose boundary converts. Its lanes are single precision and the language's scalar float is a double, so a value is narrowed going in and widened coming out, and both conversions are the backend's own round-to-nearest-even -- the delegation f32_bits makes (Q-038). A double with no float32 becomes an infinity rather than wrapping. The arithmetic between the conversions is f32: f32x4_add (f32x4_splat 1.0) (f32x4_splat 0.1) is not 1.1, it is the float32 nearest it. There is deliberately no `f32x4_load` / `f32x4_store`: for f64x2 a load is a load because Vec[R, float] already holds doubles, but the same spelling on f32 lanes would be a narrowing conversion of four doubles wearing the name of a load, and a caller that wants four real f32 lanes off a buffer wants them from bytes. Two different right answers, so neither is defined until a caller settles it):

builtinsignaturenotes
f64x2_splatfloat -> f64x2both lanes the argument
f64x2_extractf64x2 -> int -> floatlane 0 or 1; another index fails like an out-of-range vec_get
u8x16_splatint -> u8x16the low 8 bits in every lane
u8x16_extractu8x16 -> int -> intlane 0..15, as 0-255
f64x2_makefloat -> float -> f64x2lane 0, lane 1
f64x2_add / f64x2_sub / f64x2_mul / f64x2_divf64x2 -> f64x2 -> f64x2lane-wise
f64x2_reduce_addf64x2 -> floatlane 0 + lane 1, in that order
f64x2_fmaf64x2 -> f64x2 -> f64x2 -> f64x2fma per lane, each rounded once (v0.1.534). One fmla.2d on arm64 through C and LLVM; two software fma calls on Wasm
f64x2_loadVec[R, float] -> int -> f64x2lanes [i, i+2) of the Vec; past the end fails like vec_get
f64x2_storeVec[R, float] -> int -> f64x2 -> unitthe two lanes into [i, i+2)
f32x4_splatfloat -> f32x4all four lanes, narrowed to f32
f32x4_makefloat -> float -> float -> float -> f32x4lanes 0..3, each narrowed
f32x4_extractf32x4 -> int -> floatlane 0..3, widened back to a double; another index fails like an out-of-range vec_get
f32x4_add / f32x4_sub / f32x4_mul / f32x4_divf32x4 -> f32x4 -> f32x4lane-wise, at f32 precision
f32x4_reduce_addf32x4 -> float(((l0 + l1) + l2) + l3), each add at lane precision, widened once at the end. The order is part of the answer -- a pairwise tree rounds differently -- so it is fixed here and every backend does this
u8x16_from_bytes / u8x16_loadbytes -> u8x16 / bytes -> int -> u8x16the first 16 bytes / bytes [i, i+16); short fails like an index
u8x16_and / u8x16_or / u8x16_xoru8x16 -> u8x16 -> u8x16bitwise, lane-wise
u8x16_sub_satu8x16 -> u8x16 -> u8x16unsigned subtract saturating at 0
u8x16_equ8x16 -> u8x16 -> u8x160xFF where the lanes are equal, else 0
u8x16_swizzleu8x16 -> u8x16 -> u8x16table, indices: lane i = table[idx[i]], or 0 when idx[i] > 15 (pshufb / tbl / i8x16.swizzle)
u8x16_shru8x16 -> int -> u8x16each byte shifted right by 0..7
u8x16_shift_inu8x16 -> u8x16 -> int -> u8x16prev cur k: the last k bytes of prev, then the first 16-k of cur
u8x16_any_trueu8x16 -> boolany lane non-zero
u8x16_reduce_addu8x16 -> intthe lanes summed as integers (0..4080)
u8x16_first_trueu8x16 -> intthe lowest lane that is non-zero, or -1. Non-zero, not "high bit set" -- u8x16_first_true (u8x16_eq a b) is the first lane where they match. Not on the RV32IM/RV64IM backends (the two RVV instructions it needs, vmsne.vi and vfirst.m, are outside the emulator's subset); it is refused there by name

`Vec` — the rest of the family (vec_new / push / get / set / len / iter / map / filter / fold / sort / to_list / to_owned are covered in language-reference and the tutorial):

builtinsignaturenotes
vec_concatVec[R, T] -> Vec[R, T] -> Vec[R, T]a new Vec in the same region
vec_reverseVec[R, T] -> unitin place, returns unit
vec_bytesVec[R, T] -> intbytes this Vec currently holds — a measurement, not a capacity
vec_compactVec[R, T] -> unitshrink the buffer to the live length (v0.1.294-300 reclamation arc)

`Map` (map_new / get / set / has / len / delete / iter are covered in language-reference and the tutorial):

builtinsignaturenotes
map_clearMap[R, K, V] -> unitdrop every entry, keep the table
map_compactMap[R, K, V] -> unitshrink the table to fit what is left
map_recycleMap[R, K, V] -> unitreturn the table's storage for reuse
map_bytesMap[R, K, V] -> intbytes this Map currently holds

`OwnedVec`: owned_vec_get : OwnedVec[T] -> int -> T joins owned_vec_new / push / len / to_vec, which are in language-reference.

Not a family, and missing for the same reason:

builtinsignaturenotes
file_openstr -> Fileopen for reading; file_openrw is the read/write one
file_read_lineFile -> str optionNone at end of file
read_stdinunit -> strthe whole of stdin
list_dirstr -> str listentries, order unspecified
mkdir_pstr -> unitcreates parents, succeeds if it exists
channel_newunit -> Channel[T]see "Channel receive: which of the three blocks" above

Stdin on plain Wasm reaches the host (v0.1.528). read_stdin and read_line used to answer the empty string on mere -w, on the ground that a browser has no stdin. A Node host has one, and an empty string is what EOF looks like, so a program that read nothing could not tell. Both go through the host now, like run / env_var / file_exists since v0.1.350; a host with no stdin answers the empty string and that is the host saying so. scripts/wasm_stub_check.sh is the gate, and it states how much of the Wasm host surface its probes reach rather than only how many of them passed.

Allocation shape (v0.1.414). vec_sort sorts in place; its scratch buffer is malloc/free, so a sort leaves nothing in the arena. list_sort_by is the same stable merge sort over an immutable list and allocates about 3n log n cells, plus one closure environment per comparison (the comparator is curried); on 5,000 pairs that measured 8.3 MB. For a large list that is sorted once, build a Vec and use vec_sort -- see patterns §8.4.

`ListBuf[R, T]` — a list built in order (v0.1.416). An immutable list grows only at the front, so code that produces one first-to-last either recurses without a tail call (and overflows the stack in the tens of thousands) or conses in reverse and reverses at the end — and in a bump arena every cell of that reversed accumulator is garbage the moment the reverse returns. On the JSON benchmark that was 23 MB of a 105 MB run. The builder appends by writing the previous cell's tail, which nothing can observe because no cell is visible until the list is handed out:

builtintypenotes
lb_newunit -> ListBuf[R, T]bound to the innermost active region (or the program-lifetime one), like vec_new
lb_pushListBuf[R, T] -> T -> unitappends; the value is referenced, not copied
lb_to_listListBuf[R, T] -> T listhands out the head and freezes the builder; an empty builder gives Nil

Two refusals make the no-copy append sound, and they are the same failure on every backend (catchable with try_or): lb_push on a builder that lb_to_list has frozen — its cells may already be someone's list — and lb_push while a different region is current than the one the builder was created in — the cell would point at a value that dies with that region. Push where the builder was born; to build inside a block, create the builder inside it and let lb_to_list copy the list out. A ListBuf is a container: it cannot leave its region or be stored through another container across one (the same escape rules as Vec). Not available on the RV32I backend yet (refused at emit time).

`vec_sort : Vec[R, T] -> (T -> T -> int) -> unit` carries two guarantees that are worth stating because a Mere program can observe both (v0.1.349). It is stable — equal keys keep insertion order — and all four backends run the same bottom-up merge sort, so the comparator is called the same number of times in the same order on each of them; a comparator that counts or prints is a legitimate thing to write. O(n log n) comparisons. It was an insertion sort on the compiled backends (O(n²), and O(n²) memory too, because each comparison's curried application allocates in the region) against Array.sort on the interpreter (O(n log n), unstable) — one program with two asymptotics and two orderings, which `test/parity/vec_sort_stable.mere` now pins by printing the comparison count.

The per-comparison allocation remains: applying a curried closure to its first argument builds the intermediate closure's environment in the current region. It is O(n log n) of them rather than O(n²), which is a rate rather than a leak, but a sort in a hot loop still wants a region R { … } around it. Phase 19.2 added `map_iter : Map[R, K, V] -> (K -> V -> unit) -> unit` (works in all 4 backends).


See also