Contents
- v0.1.626 — 2026-10-07
- v0.1.625 — 2026-10-07
- v0.1.624 — 2026-10-06
- v0.1.623 — 2026-10-06
- v0.1.622 — 2026-10-06
- v0.1.621 — 2026-10-06
- v0.1.620 — 2026-10-06
- v0.1.619 — 2026-10-06
- v0.1.618 — 2026-10-06
- v0.1.617 — 2026-10-06
- v0.1.616 — 2026-10-06
- v0.1.615 — 2026-10-05
- v0.1.614 — 2026-10-05
- v0.1.613 — 2026-10-05
- v0.1.612 — 2026-10-04
- v0.1.611 — 2026-10-04
- v0.1.610 — 2026-10-04
- v0.1.609 — 2026-10-04
- v0.1.608 — 2026-10-04
- v0.1.607 — 2026-10-04
- v0.1.606 — 2026-10-04
- v0.1.605 — 2026-10-03
- v0.1.604 — 2026-10-03
- v0.1.603 — 2026-10-03
- v0.1.602 — 2026-10-03
- v0.1.601 — 2026-10-03
- v0.1.600 — 2026-10-03
- v0.1.599 — 2026-10-03
- v0.1.598 — 2026-10-03
- v0.1.597 — 2026-10-03
- v0.1.596 — 2026-10-03
- v0.1.595 — 2026-10-03
- v0.1.594 — 2026-10-03
- v0.1.593 — 2026-10-03
- v0.1.592 — 2026-10-03
- v0.1.591 — 2026-10-03
- v0.1.590 — 2026-10-03
- v0.1.589 — 2026-10-02
- v0.1.588 — 2026-10-02
- v0.1.587 — 2026-10-02
- v0.1.586 — 2026-10-02
- v0.1.585 — 2026-10-02
- v0.1.584 — 2026-10-02
- v0.1.583 — 2026-10-02
- v0.1.582 — 2026-10-02
- v0.1.581 — 2026-10-02
- v0.1.580 — 2026-10-02
- v0.1.579 — 2026-10-02
- v0.1.578 — 2026-10-02
- v0.1.577 — 2026-10-02
- v0.1.576 — 2026-10-02
- v0.1.575 — 2026-10-02
- v0.1.574 — 2026-10-02
- v0.1.573 — 2026-10-02
- v0.1.572 — 2026-10-02
- v0.1.571 — 2026-10-02
- v0.1.570 — 2026-10-02
- v0.1.569 — 2026-10-02
- v0.1.568 — 2026-10-01
- v0.1.567 — 2026-10-01
- v0.1.566 — 2026-10-01
- v0.1.565 — 2026-10-01
- v0.1.564 — 2026-10-01
- v0.1.563 — 2026-10-01
- v0.1.562 — 2026-10-01
- v0.1.561 — 2026-10-01
- v0.1.560 — 2026-10-01
- v0.1.559 — 2026-10-01
- v0.1.558 — 2026-10-01
- v0.1.557 — 2026-10-01
- v0.1.556 — 2026-09-30
- v0.1.555 — 2026-09-30
- v0.1.554 — 2026-09-30
- v0.1.553 — 2026-09-30
- v0.1.552 — 2026-09-30
- v0.1.551 — 2026-09-30
- v0.1.550 — 2026-09-30
- v0.1.549 — 2026-09-29
- v0.1.548 — 2026-09-29
- v0.1.547 — 2026-09-29
- v0.1.546 — 2026-09-29
- v0.1.545 — 2026-09-28
- v0.1.544 — 2026-09-28
- v0.1.543 — 2026-09-28
- v0.1.542 — 2026-09-28
- v0.1.541 — 2026-09-28
- v0.1.540 — 2026-09-28
- v0.1.539 — 2026-09-28
- v0.1.538 — 2026-09-27
- v0.1.537 — 2026-09-27
- v0.1.536 — 2026-09-27
- v0.1.535 — 2026-09-27
- v0.1.534 — 2026-09-27
- v0.1.533 — 2026-09-24
- v0.1.532 — 2026-09-24
- v0.1.531 — 2026-09-24
- v0.1.530 — 2026-09-24
- v0.1.529 — 2026-09-24
- v0.1.528 — 2026-09-24
- v0.1.527 — 2026-09-24
- v0.1.526 — 2026-09-24
- v0.1.525 — 2026-09-24
- v0.1.524 — 2026-09-24
- v0.1.523 — 2026-09-24
- v0.1.522 — 2026-09-24
- v0.1.521 — 2026-09-23
- v0.1.520 — 2026-09-23
- v0.1.519 — 2026-09-23
- v0.1.518 — 2026-09-23
- v0.1.517 — 2026-09-23
- v0.1.516 — 2026-09-23
- v0.1.515 — 2026-09-23
- v0.1.514 — 2026-09-23
- v0.1.513 — 2026-09-23
- v0.1.512 — 2026-09-23
- v0.1.511 — 2026-09-23
- v0.1.510 — 2026-09-23
- v0.1.508 — 2026-09-22
- Counts
- v0.1.507 — 2026-09-22
- Counts
- v0.1.506 — 2026-09-22
- Counts
- v0.1.505 — 2026-09-22
- Counts
- v0.1.504 — 2026-09-22
- Measured, and dropped
- Counts
- v0.1.503 — 2026-09-22
- The four bugs that only appeared when it ran
- Counts
- v0.1.502 — 2026-09-22
- v0.1.501 — 2026-09-21
- v0.1.500 — 2026-09-21
- v0.1.499 — 2026-09-21
- v0.1.498 — 2026-09-21
- v0.1.497 — 2026-09-21
- v0.1.496 — 2026-09-21
- v0.1.495 — 2026-09-21
- v0.1.494 — 2026-09-21
- v0.1.493 — 2026-09-21
- v0.1.492 — 2026-09-21
- v0.1.491 — 2026-09-21
- v0.1.490 — 2026-09-21
- v0.1.489 — 2026-09-20
- v0.1.488 — 2026-09-20
- v0.1.487 — 2026-09-20
- v0.1.486 — 2026-09-20
- v0.1.485 — 2026-09-15
- v0.1.484 — 2026-09-14
- v0.1.483 — 2026-09-14
- v0.1.482 — 2026-09-14
- v0.1.481 — 2026-09-14
- v0.1.480 — 2026-09-14
- v0.1.479 — 2026-09-13
- v0.1.478 — 2026-09-13
- v0.1.477 — 2026-09-13
- v0.1.476 — 2026-09-13
- v0.1.475 — 2026-09-12
- v0.1.474 — 2026-09-10
- The count, measured before it became an error
- Which is how it found a real bug of its own
- What this closes
- v0.1.473 — 2026-09-10
- The line, which is now a property rather than a phase
- After both
- v0.1.472 — 2026-09-10
- The count, and a correction to v0.1.470's account of it
- What it can say that the old check could not
- What stayed the same, deliberately
- v0.1.471 — 2026-09-10
- v0.1.470 — 2026-09-10
- One answer, on all five paths
- Two parity cases, because one shape cannot hold it
- What is still open, and why it is not here
- v0.1.469 — 2026-09-10
- The gate is a differential, because the hazard is not a bug
- v0.1.468 — 2026-09-10
- What promoting it found
- Gate
- v0.1.467 — 2026-09-10
- Why it took thirteen versions
- v0.1.466 — 2026-09-09
- It is not the same hook, because this backend has no uncurried twin
- The bug that made it look like the region was not reaching the callee
- v0.1.465 — 2026-09-09
- v0.1.464 — 2026-09-09
- Where it had to be keyed, which is the part that took two attempts
- What was checked, and what did not move
- v0.1.463 — 2026-09-09
- What is checked, since nothing observable changed
- v0.1.462 — 2026-09-09
- The numbers, on programs that exist
- And a counting bug in the gate, found by its own new output
- v0.1.461 — 2026-09-09
- v0.1.460 — 2026-09-09
- What it reports, and what it costs to be wrong about
- The numbers
- Also
- v0.1.459 — 2026-09-09
- The measurement first
- What the poison found
- And the gate's own bug, which the same poison showed
- v0.1.458 — 2026-09-09
- The fix is to stop reclaiming, not to keep refusing
- Five things the gates said, in the order they said them
- v0.1.457 — 2026-09-09
- And Wasm answers -424242 today. That is Q-132.
- Five comments that claimed the withdrawn behaviour
- v0.1.456 — 2026-09-09
- What survives, and it is not nothing
- The lesson, since it cost three versions
- One cell in the host matrix stopped being about something else
- v0.1.455 — 2026-09-09
- Two things fell out that were on m3d's list, not on this one
- The top-level `let`, typed once
- v0.1.454 — 2026-09-09
- Also
- v0.1.453 — 2026-09-09
- Quantified AND allocation, which the first attempt got wrong
- No hidden argument, because a call does not change the current region
- The tag stopped naming the region, in all four backends
- Also
- v0.1.452 — 2026-09-08
- v0.1.451 — 2026-09-08
- v0.1.450 — 2026-09-08
- v0.1.449 — 2026-09-08
- v0.1.448 — 2026-09-08
- v0.1.447 — 2026-09-07
- v0.1.446 — 2026-09-07
- v0.1.445 — 2026-09-07
- v0.1.444 — 2026-09-06
- v0.1.443 — 2026-09-06
- v0.1.442 — 2026-09-06
- v0.1.441 — 2026-09-06
- v0.1.440 — 2026-09-06
- v0.1.439 — 2026-09-06
- v0.1.438 — 2026-09-06
- v0.1.437 — 2026-09-06
- v0.1.436 — 2026-09-06
- v0.1.435 — 2026-09-06
- v0.1.434 — 2026-09-06
- v0.1.433 — 2026-09-06
- v0.1.432 — 2026-09-06
- v0.1.431 — 2026-09-06
- v0.1.430 — 2026-09-05
- v0.1.429 — 2026-09-05
- v0.1.428 — 2026-09-05
- v0.1.427 — 2026-09-05
- v0.1.426 — 2026-09-05
- v0.1.425 — 2026-09-05
- v0.1.424 — 2026-09-05
- v0.1.423 — 2026-09-05
- v0.1.422 — 2026-09-05
- v0.1.421 — 2026-09-05
- v0.1.420 — 2026-09-05
- v0.1.419 — 2026-09-05
- v0.1.418 — 2026-09-05
- v0.1.417 — 2026-09-05
- v0.1.416 — 2026-09-05
- v0.1.415 — 2026-09-05
- v0.1.414 — 2026-09-05
- v0.1.413 — 2026-09-03
- v0.1.412 — 2026-09-03
- v0.1.411 — 2026-09-03
- v0.1.410 — 2026-09-03
- v0.1.409 — 2026-09-03
- v0.1.408 — 2026-09-03
- v0.1.407 — 2026-09-03
- v0.1.406 — 2026-09-03
- v0.1.405 — 2026-09-03
- v0.1.404 — 2026-09-03
- v0.1.403 — 2026-09-03
- v0.1.402 — 2026-09-03
- The bug underneath
- What the corpus says now
- And one of the 38 turned out to be small
- v0.1.401 — 2026-09-02
- v0.1.400 — 2026-09-02
- v0.1.399 — 2026-09-02
- v0.1.398 — 2026-09-02
- v0.1.397 — 2026-09-02
- v0.1.396 — 2026-09-02
- v0.1.395 — 2026-09-02
- v0.1.394 — 2026-09-02
- v0.1.393 — 2026-09-02
- v0.1.392 — 2026-09-02
- v0.1.391 — 2026-09-02
- v0.1.390 — 2026-09-02
- The bug, and why every earlier layer hid it
- What it took to find
- v0.1.389 — 2026-09-02
- What the sweep says now
- v0.1.388 — 2026-09-02
- An unimplemented service is now a failure a program can catch
- The gate that was missing
- mere-ruby on the self-made CPU
- v0.1.387 — 2026-09-02
- The gate: our arithmetic against the machine's
- Also
- v0.1.386 — 2026-09-01
- v0.1.385 — 2026-09-01
- v0.1.384 — 2026-09-01
- v0.1.383 — 2026-09-01
- v0.1.382 — 2026-09-01
- v0.1.381 — 2026-09-01
- v0.1.380 — 2026-09-01
- v0.1.379 — 2026-09-01
- v0.1.378 — 2026-09-01
- v0.1.377 — 2026-09-01
- v0.1.376 — 2026-09-01
- v0.1.375 — 2026-09-01
- v0.1.374 — 2026-09-01
- v0.1.373 — 2026-09-01
- v0.1.372 — 2026-09-01
- v0.1.371 — 2026-09-01
- v0.1.370 — 2026-09-01
- v0.1.369 — 2026-09-01
- v0.1.368 — 2026-09-01
- v0.1.367 — 2026-09-01
- v0.1.366 — 2026-09-01
- v0.1.365 — 2026-08-31
- v0.1.364 — 2026-08-31
- v0.1.363 — 2026-08-31
- v0.1.362 — 2026-08-31
- v0.1.361 — 2026-08-31
- v0.1.360 — 2026-08-30
- v0.1.359 — 2026-08-30
- v0.1.358 — 2026-08-30
- v0.1.357 — 2026-08-30
- v0.1.356 — 2026-08-30
- v0.1.355 — 2026-08-30
- v0.1.354 — 2026-08-30
- v0.1.353 — 2026-08-30
- v0.1.352 — 2026-08-30
- v0.1.351 — 2026-08-30
- v0.1.350 — 2026-08-30
- v0.1.349 — 2026-08-29
- v0.1.348 — 2026-08-28
- v0.1.347 — 2026-08-28
- v0.1.346 — 2026-08-28
- v0.1.345 — 2026-08-28
- v0.1.344 — 2026-08-28
- v0.1.343 — 2026-08-28
- v0.1.342 — 2026-08-28
- v0.1.341 — 2026-08-28
- v0.1.340 — 2026-08-28
- v0.1.339 — 2026-08-28
- v0.1.338 — 2026-08-28
- v0.1.337 — 2026-08-28
- v0.1.336 — 2026-08-27
- v0.1.335 — 2026-08-27
- v0.1.334 — 2026-08-27
- v0.1.333 — 2026-08-27
- v0.1.332 — 2026-08-27
- v0.1.331 — 2026-08-27
- v0.1.330 — 2026-08-27
- v0.1.329 — 2026-08-27
- v0.1.328 — 2026-08-27
- v0.1.327 — 2026-08-27
- v0.1.326 — 2026-08-26
- How it was found, which is the part worth keeping
- v0.1.325 — 2026-08-25
- v0.1.324 — 2026-08-25
- v0.1.323 — 2026-08-25
- v0.1.322 — 2026-08-25
- v0.1.321 — 2026-08-24
- v0.1.320 — 2026-08-24
- v0.1.319 — 2026-08-24
- v0.1.318 — 2026-08-24
- v0.1.317 — 2026-08-24
- v0.1.316 — 2026-08-23
- v0.1.315 — 2026-08-23
- v0.1.314 — 2026-08-23
- v0.1.313 — 2026-08-23
- v0.1.312 — 2026-08-23
- v0.1.311 — 2026-08-23
- v0.1.310 — 2026-08-23
- v0.1.309 — 2026-08-23
- v0.1.308 — 2026-08-23
- v0.1.307 — 2026-08-23
- v0.1.306 — 2026-08-22
- v0.1.305 — 2026-08-22
- v0.1.304 — 2026-08-22
- v0.1.303 — 2026-08-22
- v0.1.302 — 2026-08-22
- The parity gate can now hold a program that blocks
- v0.1.301 — 2026-08-24
- v0.1.300 — 2026-08-23
- v0.1.299 — 2026-08-22
- v0.1.298 — 2026-08-22
- v0.1.297 — 2026-08-22
- v0.1.296 — 2026-08-22
- v0.1.295 — 2026-08-22
- v0.1.294 — 2026-08-22
- v0.1.293 — 2026-08-21
- v0.1.292 — 2026-08-21
- v0.1.291 — 2026-08-21
- v0.1.290 — 2026-08-21
- v0.1.289 — 2026-08-20
- v0.1.288 — 2026-08-20
- v0.1.287 — 2026-08-20
- v0.1.286 — 2026-08-20
- v0.1.285 — 2026-08-20
- v0.1.284 — 2026-08-19
- Five defects on the first run, all one family
- Two recorded rather than fixed
- The axes are kept apart
- v0.1.283 — 2026-08-18
- Q-045: a user binding that shadows a builtin the prelude calls
- Q-046: a parameter named the same as a top-level binding
- Three test assertions were pinning a symbol name
- The line-break gate was comparing against nothing
- contrib — GraphQL introspection and validation — 2026-08-18
- Introspection is answered by the ordinary executor
- Validation, and how a partial validator is gated
- A harness bug worth naming
- contrib — HTTP/2 flow control and gRPC streaming — 2026-08-18
- An export list is not a coverage list
- What poisoning found that valid traffic could not
- Two wrong expectations, both the same slip
- Streaming is the list having more than one element
- v0.1.282 — 2026-08-18
- v0.1.281 — 2026-08-18
- 2026-08-18 — A protobuf code generator, and two branches nothing reached
- 2026-08-18 — A .proto parser, and the bootstrap closing
- 2026-08-18 — RPC statuses, and an expectation the wire corrected
- 2026-08-18 — A GraphQL executor, and the one case out of 67 that mattered
- v0.1.280 — 2026-08-18
- 2026-08-18 — A gRPC server in Mere, and two bugs loopback was hiding
- 2026-08-18 — HPACK, and three ways a gate can pass without checking anything
- 2026-08-18 — GraphQL's type-system half, and HTTP/2 frames
- 2026-08-18 — Two protocols: the protobuf wire format, and GraphQL documents
- v0.1.279 — 2026-08-17
- v0.1.278 — 2026-08-17
- v0.1.277 — 2026-08-17
- v0.1.276 — 2026-08-17
- v0.1.275 — 2026-08-17
- v0.1.274 — 2026-08-16
- v0.1.273 — 2026-08-16
- v0.1.272 — 2026-08-16
- v0.1.271 — 2026-08-16
- v0.1.270 — 2026-08-16
- v0.1.269 — 2026-08-16
- v0.1.268 — 2026-08-16
- v0.1.267 — 2026-08-16
- v0.1.266 — 2026-08-16
- v0.1.265 — 2026-08-16
- v0.1.264 — 2026-08-16
- v0.1.263 — 2026-08-16
- v0.1.262 — 2026-08-15
- v0.1.261 — 2026-08-15
- v0.1.260 — 2026-08-15
- v0.1.257 — 2026-08-14
- v0.1.256 — 2026-08-14
- v0.1.255 — 2026-08-14
- v0.1.254 — 2026-08-14
- v0.1.253 — 2026-08-14
- v0.1.252 — 2026-08-14
- v0.1.251 — 2026-08-14
- v0.1.250 — 2026-08-14
- v0.1.249 — 2026-08-14
- v0.1.248 — 2026-08-14
- v0.1.247 — 2026-08-14
- v0.1.246 — 2026-08-14
- v0.1.245 — 2026-08-14
- v0.1.244 — 2026-08-13
- v0.1.243 — 2026-08-13
- v0.1.242 — 2026-08-13
- v0.1.241 — 2026-08-13
- v0.1.240 — 2026-08-13
- v0.1.239 — 2026-08-13
- v0.1.238 — 2026-08-13
- v0.1.237 — 2026-08-13
- v0.1.236 — 2026-08-13
- v0.1.235 — 2026-08-13
- v0.1.234 — 2026-08-13
- v0.1.233 — 2026-08-13
- v0.1.232 — 2026-08-13
- v0.1.231 — 2026-08-13
- v0.1.230 — 2026-08-13
- v0.1.229 — 2026-08-13
- v0.1.228 — 2026-08-13
- v0.1.227 — 2026-08-13
- v0.1.226 — 2026-08-13
- v0.1.225 — 2026-08-13
- v0.1.224 — 2026-08-13
- v0.1.223 — 2026-08-12
- v0.1.222 — 2026-08-12
- v0.1.221 — 2026-08-12
- v0.1.220 — 2026-08-12
- v0.1.219 — 2026-08-12
- v0.1.218 — 2026-08-12
- v0.1.217 — 2026-08-12
- v0.1.216 — 2026-08-12
- v0.1.215 — 2026-08-12
- v0.1.214 — 2026-08-12
- v0.1.213 — 2026-08-12
- v0.1.212 — 2026-08-12
- v0.1.211 — 2026-08-12
- v0.1.210 — 2026-08-12
- v0.1.209 — 2026-08-12
- v0.1.208 — 2026-08-12
- v0.1.207 — 2026-08-12
- v0.1.206 — 2026-08-12
- v0.1.205 — 2026-08-12
- v0.1.204 — 2026-08-12
- v0.1.203 — 2026-08-12
- v0.1.202 — 2026-08-12
- v0.1.201 — 2026-08-11
- v0.1.200 — 2026-08-11
- v0.1.199 — 2026-08-11
- v0.1.198 — 2026-08-11
- v0.1.197 — 2026-08-11
- v0.1.196 — 2026-08-11
- v0.1.195 — 2026-08-11
- v0.1.194 — 2026-08-11
- v0.1.193 — 2026-08-11
- v0.1.192 — 2026-08-11
- v0.1.191 — 2026-08-11
- v0.1.190 — 2026-08-11
- v0.1.189 — 2026-08-11
- v0.1.188 — 2026-08-11
- v0.1.187 — 2026-08-11
- v0.1.186 — 2026-08-11
- v0.1.185 — 2026-08-11
- v0.1.184 — 2026-08-11
- v0.1.183 — 2026-08-11
- v0.1.182 — 2026-08-11
- v0.1.181 — 2026-08-11
- v0.1.180 — 2026-08-11
- v0.1.179 — 2026-08-11
- v0.1.178 — 2026-08-11
- v0.1.177 — 2026-08-11
- v0.1.176 — 2026-08-11
- v0.1.175 — 2026-08-11
- v0.1.174 — 2026-08-11
- v0.1.173 — 2026-08-10
- v0.1.172 — 2026-08-10
- v0.1.171 — 2026-08-10
- v0.1.170 — 2026-08-10
- v0.1.169 — 2026-08-10
- v0.1.168 — 2026-08-10
- v0.1.167 — 2026-08-10
- v0.1.166 — 2026-08-10
- v0.1.165 — 2026-08-10
- v0.1.164 — 2026-08-10
- v0.1.163 — 2026-08-10
- v0.1.162 — 2026-08-10
- v0.1.161 — 2026-08-10
- v0.1.160 — 2026-08-10
- v0.1.159 — 2026-08-10
- v0.1.158 — 2026-08-10
- v0.1.157 — 2026-08-10
- v0.1.156 — 2026-08-10
- v0.1.155 — 2026-08-10
- v0.1.154 — 2026-08-10
- v0.1.153 — 2026-08-09
- v0.1.152 — 2026-08-09
- v0.1.151 — 2026-08-08
- v0.1.150 — 2026-08-08
- v0.1.149 — 2026-08-08
- v0.1.148 — 2026-08-08
- v0.1.147 — 2026-08-08 — the Mere compiler runs on the Mere CPU
- v0.1.146 — 2026-08-08
- v0.1.145 — 2026-08-08
- v0.1.144 — 2026-08-08
- v0.1.143 — 2026-08-08
- v0.1.142 — 2026-08-08
- v0.1.141 — 2026-08-08
- v0.1.140 — 2026-08-08
- v0.1.139 — 2026-08-08
- v0.1.138 — 2026-08-08
- v0.1.137 — 2026-08-08
- v0.1.136 — 2026-08-08
- v0.1.135 — 2026-08-08
- v0.1.134 — 2026-08-08
- v0.1.129 — 2026-08-07
- v0.1.128 — 2026-08-06
- v0.1.127 — 2026-08-06
- v0.1.115 — 2026-08-04
- v0.1.114 — 2026-08-04
- v0.1.113 — 2026-08-04
- v0.1.112 — 2026-08-04
- v0.1.111 — 2026-08-04
- v0.1.110 — 2026-08-04
- v0.1.109 — 2026-08-04
- v0.1.108 — 2026-08-03
- v0.1.107 — 2026-08-03
- v0.1.106 — 2026-08-03
- v0.1.105 — 2026-08-03
- v0.1.104 — 2026-08-03
- v0.1.103 — 2026-08-03
- v0.1.102 — 2026-08-03
- v0.1.101 — 2026-08-03
- v0.1.100 — 2026-08-03
- v0.1.99 — 2026-08-03
- v0.1.91 — 2026-07-31
- v0.1.90 — 2026-07-30
- v0.1.89 — 2026-07-30
- v0.1.88 — 2026-07-30
- v0.1.87 — 2026-07-29
- v0.1.86 — 2026-07-29
- v0.1.85 — 2026-07-29
- v0.1.84 — 2026-07-29
- v0.1.83 — 2026-07-29
- v0.1.82 — 2026-07-29
- v0.1.81 — 2026-07-29
- v0.1.80 — 2026-07-29
- v0.1.79 — 2026-07-28
- v0.1.78 — 2026-07-28
- v0.1.77 — 2026-07-28
- v0.1.76 — 2026-07-28
- v0.1.75 — 2026-07-28
- v0.1.74 — 2026-07-28
- v0.1.73 — 2026-07-28
- v0.1.72 — 2026-07-28
- v0.1.71 — 2026-07-28
- v0.1.70 — 2026-07-27
- v0.1.69 — 2026-07-27
- v0.1.68 — 2026-07-27
- v0.1.67 — 2026-07-18
- v0.1.66 — 2026-07-18
- v0.1.65 — 2026-07-18
- v0.1.64 — 2026-07-18
- v0.1.63 — 2026-07-18
- v0.1.62 — 2026-07-18
- v0.1.61 — 2026-07-18
- v0.1.60 — 2026-07-18
- v0.1.59 — 2026-07-18
- v0.1.58 — 2026-07-18
- v0.1.57 — 2026-07-18
- v0.1.56 — 2026-07-17
- v0.1.55 — 2026-07-17
- v0.1.54 — 2026-07-17
- v0.1.53 — 2026-07-17
- v0.1.52 — 2026-07-17
- v0.1.51 — 2026-07-17
- v0.1.50 — 2026-07-17
- v0.1.49 — 2026-07-17
- v0.1.48 — 2026-07-17
- v0.1.47 — 2026-07-17
- v0.1.46 — 2026-07-16
- v0.1.45 — 2026-07-16
- v0.1.44 — 2026-07-16
- v0.1.43 — 2026-07-16
- v0.1.42 — 2026-07-16
- v0.1.41 — 2026-07-16
- v0.1.40 — 2026-07-16
- v0.1.39 — 2026-07-16
- v0.1.38 — 2026-07-16
- v0.1.37 — 2026-07-15
- v0.1.36 — 2026-07-15
- v0.1.35 — 2026-07-15
- v0.1.34 — 2026-07-15
- v0.1.33 — 2026-07-15
- v0.1.32 — 2026-07-15
- v0.1.31 — 2026-07-15
- v0.1.30 — 2026-07-15
- v0.1.29 — 2026-07-15
- v0.1.28 — 2026-07-15
- v0.1.27 — 2026-07-14
- v0.1.26 — 2026-07-14
- v0.1.25 — 2026-07-14
- v0.1.24 — 2026-07-14
- v0.1.23 — 2026-07-14
- v0.1.22 — 2026-07-14
- v0.1.21 — 2026-07-14
- v0.1.20 — 2026-07-14
- v0.1.19 — 2026-07-13
- v0.1.18 — 2026-07-13
- v0.1.17 — 2026-07-13
- v0.1.16 — 2026-07-13
- v0.1.15 — 2026-07-13
- v0.1.14 — 2026-07-13
- v0.1.13 — 2026-07-13
- v0.1.12 — 2026-07-13
- v0.1.11 — 2026-07-13
- v0.1.10 — 2026-07-12
- v0.1.9 — 2026-07-12
- v0.1.8 — 2026-07-12
- v0.1.7 — 2026-07-11
- v0.1.6 — 2026-07-11
- v0.1.5 — 2026-07-10
- v0.1.4 — 2026-07-10
- v0.1.3 — 2026-07-10
- v0.1.2 — 2026-07-10
- v0.1.1 — 2026-07-10
- v0.1.0 — 2026-07-09 (first tagged release)
- 2026-07-06 — Tutorial: implement type inference in Mere (roadmap step 4, third of three — series complete)
- 2026-07-06 — Tutorial: build a Redis client in Mere (roadmap step 4, second of three)
- 2026-07-06 — Tutorial: build a REST API in Mere (roadmap step 4, first of three)
- 2026-07-05 — Cloudflare Worker: package registry v0.1 (JSON API)
- 2026-07-05 — Cloudflare Worker: playground snippet share (KV-backed)
- 2026-07-05 — Cloudflare Worker template (roadmap step 2)
- 2026-07-05 — package system v0.1: `.mere_modules/` walk-up resolution
- 2026-07-05 — `contrib/db/redis_ratelimit`: distributed fixed-window limiter
- 2026-07-05 — `contrib/os/parallel_map`: N shell commands in parallel
- 2026-07-05 — `contrib/os/subprocess`: sync shell-out (Q-012 Path A)
- 2026-07-05 — `contrib/http/websocket`: RFC 6455 hub
- 2026-07-05 — `examples/http_admin_dash`: integration dogfood
- 2026-07-05 — `contrib/db/redis_lock`: distributed mutex + `gen_request_id` shared
- 2026-07-05 — `contrib/http/cache`: Cache-Control postures + ETag / 304
- 2026-07-05 — `contrib/db/redis_stream`: consumer groups (XGROUP / XREADGROUP / XACK / XPENDING)
- 2026-07-05 — `contrib/db/redis_stream`: XADD / XREAD / XLEN
- 2026-07-05 — `contrib/http/csrf`: synchronizer-token CSRF middleware
- 2026-07-05 — playground: `wordcount` demo + build tail-call flag
- 2026-07-05 — `contrib/db/redis_hll`: HyperLogLog cardinality estimators
- 2026-07-05 — `contrib/log`: level filtering + field-taking variants + `LOG_LEVEL` env
- 2026-07-05 — `contrib/http/basic_auth`: RFC 7617 Basic Auth middleware
- 2026-07-05 — Blog-engine papercuts: lexer + typer polish
- 2026-07-05 — `sse_bridge_from_redis`: multi-instance SSE fanout
- 2026-07-05 — `contrib/http/session`: consolidate cookie-session pattern
- 2026-07-05 — `contrib/http/metrics`: Prometheus-style metrics + middleware
- 2026-07-05 — `examples/gh_stars`: first CLI demo
- 2026-07-05 — `redis_pubsub_run_forever` + `sleep_ms` extern + tcp_worker `end`-event fix
- 2026-07-05 — `contrib/db/redis_queue`: list-backed work queue
- 2026-07-05 — `http_fetch` shared across both runners
- 2026-07-05 — `contrib/http/client`: request + response headers, per-call timeout
- 2026-07-04 — `contrib/db/redis_pubsub`: dispatch layer
- 2026-07-04 — `contrib/http/router`: `route_prefix` mount points
- 2026-07-04 — `contrib/http/router`: `:capture` path params
- 2026-07-02 — Phase 54.36 runtime codegen bootstrap unblocked
- 2026-07-02 — Phase 54.35 web backend Stage A (contrib/http)
- 2026-06-30 → 2026-07-01 — Phase 54 self-host bootstrap loop closes
- 2026-06-22 (cont. — Phase 38.G-1 OwnedVec auto scope-bound Drop)
- 2026-06-22 (cont. — Phase 38.C multi-arg curried builtin first-class)
- 2026-06-22 (cont. — Phase 37 public-release prep)
- 2026-06-22 (cont. — Phase 36 syntactic sugar + dogfood examples)
- 2026-06-22 (cont. — Phase 32 C1 FFI)
- 2026-06-22
- 2026-06-21
- 2026-06-20
- 2026-06-19
- 2026-06-17
- 2026-06-16
- 2026-06-15 — 06-16 (early week)
- 2026-06-06 (start date)
- Cumulative (as of 2026-06-16)
- Not yet started (future)
Changelog (mere)
Major implementation milestones recorded per-slice (newest first). See git log for detailed commit messages.
v0.1.626 — 2026-10-07
_On -rv / -rv64, two syntax errors -- or two imports that resolve nowhere -- are reported at the program's own lines again, not at the RISC-V prelude's._
mere -rv64 of a tree whose .mere_modules had not been installed answered
parse error: import: cannot resolve path `mgz/inflate.mere` ...
--> <rv-prelude>:312:1
with the prelude's line 312 printed under it; -c said main.mere:312:1. When the parse fails, the report parses the program again on its own, with recovery, to list every error. Those positions are the file's own lines -- no prelude is in front of it, and an import is resolved before the prelude is glued on -- but on the RISC-V paths they were read as if they counted from the top of the prelude, and moved into it. v0.1.604 fixed the same thing for type errors; the parse's report kept it. A lone error was right, because it was reported at the first parse's position, which does count from the prelude; it now takes the re-parse's, which is the same line.
scripts/rv_prelude_check.sh compiles three programs (two syntax errors, one unresolvable import, two of them) with -c and -rv and requires the same --> lines; against v0.1.625 the first and the third fail.
v0.1.625 — 2026-10-07
_A signal a program wants to hear about: proc_sig_catch installs a handler that only marks it, proc_sig_take takes the marks where the program can act on them, and proc_sig_ignore is SIG_IGN. For mere-ruby, whose trap answered only Process.kill to itself and died of a SIGUSR1 sent from outside._
mere-ruby's Signal.trap wrote the handler into a table and installed nothing, so a signal from another process took the default action: trap(:USR1) { ... } and then kill -USR1 from outside ended it (exit 158), where ruby runs the handler on its main thread. CRuby's test_process waits for a child ruby to print from a trap while it sits in system; the child died, its cat kept the pipe, and the whole file hung (mere-ruby note 290).
A handler may do almost nothing, and an interpreter's handler code is anything but nothing. So the handler marks the signal (a sig_atomic_t per number and one that says some mark is up) and returns; the program asks at a point of its own -- an interpreter at a statement boundary -- and does there what the handler could not.
proc_sig_catch sig keep--SA_RESTART, so a blocking call resumes and the
mark waits for the next ask. keep = 1 is ruby's rule for the signals it handles by default: an ignored disposition inherited from the parent (nohup's SIGHUP) is put back and kept. 0 installed, 1 kept, -1 refused (EINVAL for a number out of range, in proc_last_errno).
proc_sig_take ()-- the lowest marked signal, its mark taken down;0when
none. Two arrivals before a take are one mark, as two pending signals are one in the kernel.
proc_sig_ignore sig--SIG_IGN, inherited acrossexec(2), which is what
ruby's trap(sig, "IGNORE") gives a child.
C only, like the rest of the proc_sig_* family; the other backends refuse the extern by name. scripts/procsig_check.sh grows from 22 rows to 50: the marking handler from the default (a mark from raise, from another process, two arrivals as one, the lower number first, a child starting with the default, an unknown number refused) and under an inherited SIG_IGN for SIGHUP (keep = 1 kept, keep = 0 catches). Two poisons (a take that leaves the mark up, a keep that is not honoured) each fail it. Run once on the CI image (Linux x86-64, gcc -O0 and -O2, dash): 50/50.
v0.1.624 — 2026-10-06
_RISC-V gives back what a compaction frees while coroutines exist, keeping only blocks a stopped stack can reach -- C's pins. mere-ruby's real Fibers (its collector runs on a coroutine of its own) now run corpus 165 and 213 on RV64. And sleep_ms exists there._
v0.1.623 reused no arena block while any coroutine existed. mere-ruby collects on a coroutine of its own, so every collection's frees were kept for good, and with its Fiber layer on RV (m_fiber_coro in place of the one that runs a fiber to completion) corpus 165 ran out of memory in its first collection. Now:
- A block handed back while coroutines exist waits on a retired list. After a
compaction builtin, or at a region block's exit, once a megabyte has been retired since the last time, the prelude's rvcoro_release walks every stopped stack -- the walk coro_scan_ints makes, three hops and the stack's own region values -- and gives back the blocks none of them reaches; the rest wait for the next time (not counted again, so a stack that pins a megabyte does not make every block exit a release).
- The walk allocates nothing: a Map and a Vec per walk were garbage inside
mere-ruby's collector. Its visited set is an open-addressing table kept across walks (stamped, so nothing is cleared) and its work list a reused array.
- A stack's "own region values" run from its innermost open block's mark to
where the heap stood when it stopped -- not to the heap's current top, which for mere-ruby's main stack, stopped while the collector ran, took in all of the collector's tables. A release's walk is capped (65,536 addresses); one that reaches the cap keeps every block, as before.
coro_scan_intscallsfafter its walk, as C does: called from inside it,
mere-ruby's collector code exited region blocks, whose release started a second walk over the first one's state.
sleep_mswaits on the clock (there is no sleep to ask the emulator for);
mere-ruby's Fiber scheduler sleeps.
Measured with mere-ruby 528bdd6 on RV64 at 256 MB, its Fiber layer in place: corpus 213, 220, 231, 233 and 237 agree with ruby (they did not with the layer that runs a fiber to completion), and 165 still does (34 s → 44 s). 214, 219 and 235 switch fibers every few statements and still run out of memory: every switch raises the high-water mark, so on one bump heap shared by every stack no region block's rollback reaches past it. Giving each stack an allocation area of its own is the next step.
compact_reuse now churns past a megabyte so that the release runs, and coro_check poisons it (RV 9: the release gives back a block a stopped stack reaches -- the held string is overwritten).
v0.1.623 — 2026-10-06
_RISC-V has coroutines: coro_new, coro_new_sized, coro_transfer, coro_switch, coro_exit, coro_root and coro_scan_ints, on both widths. Every coroutine runs in one region of RAM and a switch copies stacks in and out of it, so ten thousand suspended coroutines fit in a few megabytes._
The RV backend refused coroutines by name: "the bare-metal runtime has one stack and nowhere to map another". It still has no virtual memory, so a stack per coroutine would be a fixed reservation of real RAM -- and test/coro/many.mere keeps ten thousand suspended while sized_reuse recurses twenty thousand deep in one. So the stacks are COPIED instead:
- Every coroutine runs with its stack ending at the same address, in a region of
its own. __rv_cswap saves ra, s0 and s1..s11 on the stack being left, writes that stack's live part out to a buffer kept with the coroutine, writes the one being entered back to the same addresses, and restores. Nothing that points into a stack (a saved fp, a try_or record) ever moves. The main stack is not copied. A coroutine that has never run holds a seeded frame whose return address is __rv_coro_boot.
- The rest of the runtime is the RV prelude's, in Mere (
rvcoro_): a table of
records with a free list, generational handles (a stale handle to a reused slot is told apart, as C's are), the checks and messages the C backend gives (finished, not the running coroutine, a body that hands over to itself), and what each stack carries across a switch -- its try_or record, its region depth and block marks. It is written on a few new internal primitives (raw words, the runtime's words, the switch, the layout).
- Layout. A program that makes or switches to a coroutine gives a sixteenth
of RAM (at least 256 KiB) to the main stack and the region below it to the coroutines -- a sixteenth too, or --coro-stack <MB> (refused if the heap would be left less than an eighth). An eighth each was the first choice; mere-ruby's corpus 213 needed the heap back. The heap's limit is kept in tp, which this backend did not use, so the allocation check is still one instruction; every function entry checks the running stack's floor (coro_new_sized sets how far below the region's top a coroutine may reach) and an overflow prints the C backend's "stack overflow (recursion too deep)". A program without coroutines is laid out, and checks, exactly as before.
- Regions and arenas. Every allocation comes from one bump pointer, so a
region block the main stack opened before a switch could roll back what a coroutine allocated after it: each switch raises the high-water mark, past which no rollback goes. While any coroutine exists, an arena block handed back by a compaction is not reused -- coarser than C, which keeps exactly what a stopped stack reaches.
coro_scan_intsreads a stopped coroutine's buffer (or, for the running one,
its stack after spilling the saved registers) and follows the pointers in it as C does: three hops, without a limit above the stack's innermost open block's mark. (The outermost first: mere-ruby keeps a block open for its whole run, and with one bump pointer the whole heap counted as "its own" and was walked -- corpus 165 ran out of memory in its collector's scan.)
scripts/coro_check.sh runs its fixtures on memu at both widths (when MEMU is set; RV's failure message has no fail: prefix, and the output is compared, since memu exits 0 whatever the guest's exit said), with six RV poisons that edit the prelude (mere --rv-prelude prints it; MERE_RV_PRELUDE_FILE substitutes one). Two fixtures are new, for things the existing ones did not exercise: region_cross (the main stack's region closed after a coroutine allocated in it -- red without the switch's high-water mark) and compact_reuse (a map compacted and its blocks taken by others while a coroutine holds a value from it -- red when blocks are reused). Both run on C and the interpreter too.
The stack growing into the heap is caught too (every RV program, one instruction per call). An allocation checks gp against sp, but a call made after the heap had come close was not checked: its frame landed on heap data. Moving the heap up by the coroutine runtime's words made test/parity/region_growth.mere at 32 MB on RV64 trap in the float helpers instead of reporting the exhaustion (memu's trace stopped at a prologue with sp below gp). Each prologue now checks sp against gp (a program with coroutines checks its stack's floor instead; a --bare program, whose trap handler may run with another process's gp, does not); mere-ruby's corpus 165 run ten times takes 0.6% more instructions. And rv_exec_check no longer counts an RV run that printed NOTHING as agreeing: it compares once more with C's last line dropped (for a final value only a compiled-in main prints), and an empty run matched any one-line program -- which is how region_growth's trap passed for "now agrees" there.
v0.1.622 — 2026-10-06
_file_delete : str -> bool removes a file, on all five backends. Mere had no way to: a program ran rm -f, and a RISC-V binary on memu has no shell to run it with._
mere-ruby's File.delete is rm -f behind run. On RISC-V there is no process to start, so the eight corpus programs that end by cleaning up their temporary files stopped there, after everything else they printed had come to agree (mere-ruby 528bdd6 had just moved File.realpath and friends off the shell).
file_delete p is unlink: true when this call removed the file, false for a path that is not there, for a directory (unlink refuses one on every host; this is not rmdir), and for a refusal. It answers yes or no rather than an errno, because the LLVM backend has no portable way to read errno (__error on Darwin, __errno_location on glibc, and a weak reference to the other one does not link); a caller that needs the reason asks file_exists or the stat devices after a false.
- interpreter:
Unix.unlink; C and LLVM:unlink(2); Wasm: the host's
file_delete import (scripts/run_wasm.js: fs.unlinkSync), as file_exists goes through the host.
- RISC-V:
__rv_unlinkisunlinkat(AT_FDCWD, path, 0), syscall 35, the Linux
number, like the openat / faccessat the hosted target already uses, and the prelude's file_delete is __rv_unlink p == 0. memu answers it (memu needs this Mere or later to build, since it answers with file_delete).
test/parity/file_delete.mere writes a file, removes it, removes it again, removes a path that was never there and tries a directory: true true false false false false true on the interpreter, C, LLVM, Wasm, RV64 and RV32. The name is random so that parity and rv_exec running it at once do not share a file.
v0.1.621 — 2026-10-06
_On LLVM and Wasm a long list carried out of a region block no longer overflows the stack (Q-201): the copy walks the list's spine as a loop, as C's has since v0.1.605._
Leaving a region block deep-copies the value into the enclosing region. For a recursive variant the copier called itself once per node -- @__mcopy_<tag> on LLVM, $__mcopy_<tag> on Wasm -- so a list's length was the copy's recursion depth. Measured: 10,000 and 30,000 elements copied, 100,000 overflowed on both backends (Q-201 had put the limit at ten million, measured only on LLVM; Wasm had not been measured). C has walked the spine since v0.1.605, for the same reason.
Both now do what C does: a constructor whose payload is a tuple ENDING in the type itself (a list's Cons, a tree's last child) copies its other fields, leaves the last one open, and goes round again with the old last field; each new node is stored into the slot the previous one left. Every other constructor is copied as before and ends the walk. LLVM keeps the walk's position in two allocas, Wasm in locals.
scripts/deep_list_check.sh gains test/deep_list/region_out_long_list.mere: 300,000-element int, str and tuple lists and a 100,000-node tree carried out of region blocks, on C, LLVM and Wasm (its expected output is in the script: the interpreter's own recursion depth stops short of these lengths). Its --poison turns each backend's loop back into a call to itself, and each then overflows. v0.1.620 overflowed on LLVM and on Wasm with the same program.
v0.1.620 — 2026-10-06
_On RISC-V a comparison with a string literal is inline, and char_at / chr allocate nothing. mere-ruby's start-up on RV64 takes 10.6% fewer instructions. The new test also found C and LLVM comparing strings with strcmp in four places, which stops at a NUL inside the value._
RISC-V. Counted over mere-ruby's start-up on RV64 (390M instructions, about 4 s on memu, almost all of it lexing and parsing the ~8000 lines of ruby it keeps as its core prelude), 4.6M of the 5.2M string comparisons were against a literal -- let c = char_at s i in ... str_eq c "x" -- at about 18 instructions each through __str_eq. str_eq x "lit", x == "lit" / != and a "lit" pattern now compare in place: the length against the literal's (a string of another length costs one load and one branch), then, for a literal that fits a word, the first data word masked to that length against the literal's bytes. The read is in bounds because every string is its length word and its bytes rounded up to whole words, and it happens only once the lengths agree. A longer literal checks the length inline and calls __str_eq only when it matches.
char_at and chr allocated two words for every character. They now return an entry of a table of the 256 one-byte strings, kept in data like a literal (a string is never written through, so nothing can tell an entry from a fresh copy).
Measured: start-up 390M → 349M instructions; mere-ruby's corpus 165 run ten times 858M → 806M, output unchanged; the binary 1.7% larger.
C and LLVM: a NUL inside a string. test/parity/str_literal_compare.mere compares strings of lengths 0 to 9 (both sides of a word at both widths) and strings with NUL, high and multi-byte characters. On C and LLVM the match arm "a" took "a\u0000c" and "" took "\u0000": a string pattern was compared with strcmp, which stops at the first NUL. == had been by length since v0.1.264; the same strcmp remained in
- C: a string pattern, a str field in structural equality (
eq_*), a str in a
structural ordering, a Map's str keys (so "a" and "a\u0000b" were one key);
- LLVM: the same four.
All now compare by length (__lang_str_cmp in C, @__lang_str_eq / @__lang_str_compare in LLVM). The test has rows for each.
`lsp_smoke` no longer races. v0.1.619's CI went red on it: two publishes where three were expected. Since v0.1.576 the server handles an unbroken run of didChanges for one document as its last one, and the gate pipes every message at once, so whether the second change had already arrived when the first was handled depended on how fast the server started. Starting it a second late reproduces the failure every time. A request now sits between the two changes, which breaks the run whatever the timing.
v0.1.619 — 2026-10-06
_On the C backend a channel message no longer corrupts the heap's block lists. Since v0.1.549 examples/concurrent_loops.mere summed 1..100 to a garbage number about once in two thousand runs._
v0.1.549 put every live region block on one of two lists, so that a conservative scan of a suspended coroutine can ask whether an address may be read: the default region's blocks on a shared list behind a lock, every other region's on its thread's own list ("Map, Vec and region borrows are not Send, so a region belongs to its thread"). A channel message is the exception: the sender builds it in a fresh region and the RECEIVER frees that region. Its blocks went on the sender's list, and the receiver's unlink edited the receiver's list instead -- it set its own head to the sender's next block, and left the sender's head on a block it had just freed. Two ways that went wrong:
- two threads racing on the lists (
concurrent_loops: wrong total about once in
2,000 runs alone, 5 in 4,000 under load; ThreadSanitizer names __lang_blk_link against __lang_blk_unlink);
- deterministically, when a thread sends a batch, waits for the receiver some other
way (a join is not a message) and then builds another message: its link writes into the freed block.
A region now says whether another thread frees it (xthread, set for channel messages), and its blocks go on the shared, locked list. Only the C backend keeps these lists. test/uaf/channel_xthread.mere is the second shape, run by region_uaf_check under AddressSanitizer (heap-use-after-free on every run before the fix; its poison puts the region back on the maker's list). That gate now sets ASAN_OPTIONS=use_sigaltstack=0: a spawned thread's alternate signal stack is the runtime's own (Q-178), and ASan on macOS tried to unmap it when the thread ended.
v0.1.618 — 2026-10-06
_RISC-V frames are sized by the bindings live at once rather than by every binding in the body, and a self tail call loops back past the prologue instead of rebuilding the frame. mere-ruby's corpus 165 on RV64 runs in another 5% fewer instructions; from v0.1.615 it is 47% fewer in all._
Slots are handed back. A function's frame had a slot for every binding its body made (count_lets summed them), and the first ten slots live in s1..s10, which the prologue saves and the epilogue restores. Now a binding's slot goes back when its scope ends: compile_expr restores the slot counter after each expression, and every match arm starts from the same base. The frame is sized by max_lets, the most bindings in scope at any one point, and the check after each body compares that with the counter's high-water mark (it used to compare count_lets with the final counter). The binary shrinks 7%; the count of saved registers hardly moves, because the functions that run most were already under ten.
A self tail call is a loop. Most of the remaining saves and restores were in self-recursive loops (go i acc = ... go (i + 1) ...): every iteration tore the frame down, jumped to the function's entry and ran the prologue again. A saturated self call in tail position now resets sp to the frame base and jumps to a label just after the prologue, where the arguments are copied into their slots; the saved registers stay saved for the whole loop. Calls to other functions in tail position still tear down and jump.
On LLVM, a tail call after a `let` is `musttail` again. The new parity test overflowed the stack there and nowhere else: the LLVM backend's let arms never passed tail position on to their body, so a loop of the commonest shape, let s = f i in go (i - 1) (acc + s), was a plain call that grew the stack once per iteration (at -O0; clang's sibling-call optimisation hid it at -O2). Now every let arm hands tail position on -- except a let whose fresh OwnedVec is freed at scope end, where a ret in the body would skip the free.
Not done: putting a leaf function's bindings in caller-saved registers, the other half of the plan. Counted, leaf functions pay 18% of the saves that are left, about 1% of all instructions.
Measured on that corpus: 903M → 858M instructions (slot reuse 903M → 888M, the loop 888M → 858M), binary 11.98 MB → 11.08 MB. test/parity/frame_reuse.mere reuses slots across siblings and arms and loops with eight register arguments from inside a match arm whose tuple pattern parks its pointer on the stack.
v0.1.617 — 2026-10-06
_The RISC-V backend rewrites a push/pop pair around straight-line code into two register moves, then propagates copies and drops dead moves. mere-ruby's corpus 165 on RV64 runs in another 9.5% fewer instructions, from an 8% smaller binary._
The emitter is a stack machine: an operand that must survive the evaluation of the next is pushed (addi sp, sp, -w; sd r, 0(sp)) and popped (ld r', 0(sp); addi sp, sp, w). Counted over mere-ruby those four instructions were 24% of everything run, and mv another 15%. Two passes now run over the item list before layout, so the binary, the listing and the debug map all describe the rewritten code (rv_listing_check holds them to that):
- push/pop → mv: when the code between a push and its pop has no label,
jump, branch, call or ecall, and names neither sp nor a spare temporary (t6, t5, t4), the pair becomes mv t, r / mv r', t. Nothing in between can have read the slot, and nothing can have run elsewhere and come back. Repeated until nothing changes, so nested pairs each get their own spare.
- copy propagation: inside a straight-line run, an instruction that reads
a register holding a copy made by mv reads the original instead (while neither has been written since), and a mv whose destination is written again before anything reads it is dropped. Only argument, temporary and saved registers take part, never sp, fp, gp, ra or zero, and every register's value is assumed live at the end of the run. The formats it reads are R, I, loads, S and U; any other opcode ends the run.
Measured on that corpus: 998M → 903M instructions, binary 12.99 MB → 11.98 MB. Pushes and pops that remain sit around calls, branches or nested evaluations that span them.
v0.1.616 — 2026-10-06
_On RISC-V a Map lookup is a runtime routine instead of prelude code: about 400 instructions became 91 for a str key and 42 for an int key. mere-ruby's corpus 165 on RV64 runs in 39% fewer instructions._
The RV prelude's Map is written in Mere, and every lookup went through rvmap_get → _mslot → _mprobe (one six-argument call per probe step, a tuple destructured and a bounds-checked vec_get per field), with the str hash and __str_eq called on the way. Counting every instruction of mere-ruby on RV64 (corpus 165, ten iterations) put the Map at more than half of 1.63 billion instructions, at about 400 instructions per lookup although a lookup takes 1.5 to 1.8 probe steps on average.
The lookup is now emitted as machine code:
__rv_mslot_wand__rv_mslot_sare leaf routines that hash the key, probe
linearly and return the slot (or -(slot+1) for the first tombstone or empty slot), the same contract as the prelude's _mslot_i / _mslot. The word routine mixes the key as _mhash_i does; the str routine compares length first and then a word at a time, with the last word masked to the length.
map_get,map_hasandmap_deletein user code call__rv_mget_*,
__rv_mhas_* and __rv_mdel_*; a missing key in map_get fails with the message the other backends give (map_get: key not found in Map (use map_has to check first)), and the prelude's own rvmap_get now says the same.
map_setfinds the slot with the routine and either updates in place
(__rv_mupd, which stores through __vec_set_rt so a value in an older container is copied into its arena as before) or inserts through the prelude. Insertion, growth, reindexing, compaction and iteration stay in the prelude; they run far less often.
__rv_str_hash hashes a word at a time (an FNV-style multiply per word, with the partial last word masked) instead of a byte at a time, and the str lookup shares its body, so the prelude's inserts and the runtime's lookups agree. Iteration order is unaffected (a Map iterates its entries in insertion order).
Measured on that corpus: 1,634M → 998M instructions; the Map routines and __str_eq together fall from 61% of the count to 36%.
v0.1.615 — 2026-10-05
_On the C backend a builtin's arguments run left to right under gcc as well: vec_set v (f 0) (f 1) ran f 1 first there. v0.1.614's CI went red on it._
C leaves the order of a call's arguments unspecified, and gcc evaluates them right to left. The backend had already decided this for user calls (every argument is bound to a __da temporary) and, in v0.1.450, for operators; a builtin or extern application was still written straight into a C call. So vec_set v (f 0) (f 1), map_set m (g "k") (g "v") and substring s (f 1) (f 4) ran their arguments backwards when the emitted C was built with gcc, and in order with clang -- which is what parity.sh builds with, so the parity suite could not see it. rv_exec_check builds its C reference with cc, gcc on the CI runner, and v0.1.614's new container_arenas (a random walk with two draws from one generator in one vec_set) printed another checksum there than on RISC-V.
Now an application whose head is a builtin or an extern, with two or more arguments that can have effects, binds those arguments to lets in order and applies the builtin to the variables (a closure literal counts as effect-free, so a builtin arm that inlines one still sees it). test/parity/builtin_arg_order.mere prints each argument as it runs; v0.1.614's C built with gcc printed f1 f0 gv gk f4 f1 ... for it, the interpreter's order is f0 f1 gk gv f1 f4 ....
v0.1.614 — 2026-10-05
_RISC-V containers keep their contents the way C's do: in arenas a store copies into, that a compaction frees and a recycle winds back. mere-ruby's corpus 165 and 285, which ran out of 256 MB on RV64, finish in 165 MB and 139 MB._
What was missing. v0.1.613 gave region blocks their rollback, and mere-ruby still did not fit: it stores into a long-lived table on almost every statement, each store raised the high-water mark, and its per-statement blocks gave back almost nothing. And vec_compact did nothing on this backend, map_compact only packed, map_recycle only cleared, so its collector freed no bytes. On C the same program is held at 90 MB by exactly those two things: a store copies its value into the container's own region, so a block gives back all of its garbage, and a compaction moves a container into a fresh region and frees the old one (measured: with the compactions turned off in the generated C, corpus 165 went from 90 to 129 MB; 285 was held by the blocks alone).
Arenas. An arena is a chain of power-of-two blocks, 1 KB up, carved from the bump heap a megabyte at a time (a carve inside a region block raises the high-water mark, so carves have to be rare) and given back to size-class free lists; a larger free block is split, nothing is merged. A Vec's cell has a fourth word, its arena. A container whose type says __heap is made in the shared default arena -- cell, buffer and Map tuple included -- as C makes it in the default region; vec_compact / map_compact copy the live elements into an arena of its own and free the old one (never the shared one); map_recycle empties a Map and resets its arena to the first block; vec_bytes / map_bytes answer an owned arena's capacity, 0 otherwise. A block- or parameter-region container stays on the bump heap as before.
Copy-on-store, typed at the call site. Values on this backend are untagged words, so a copy needs the type, and the prelude's Map code is compiled once for every value type (its values are erased to int; v0.1.613's note). So the copy is made where the type is known: vec_push / vec_set of a value with anything to copy call a per-type helper, and map_set is lowered to find, copy (the value; the key too when it is new -- C's rule), then update or insert. The copier is v0.1.613's __rcopy_<tag> run with gp pointed at an exact reservation that a per-type __ssize_<tag> measured first; a copy that ever ran past its reservation stops the program with that message rather than writing over the next thing in the arena. Containers and closures inside a stored value are kept as handles and, if they lie in an open block, keep it. Once any arena exists, a store into a container on the bump heap copies too, so nothing a container holds points into an arena that a compaction may free.
The contract is C's: a value read out of a container does not survive that container's compaction or recycle. RISC-V never freed anything before, so code that broke the contract ran there by accident; the first mere-ruby run on arenas crashed on two such reads, which turned out to be use-after-frees on C as well (AddressSanitizer), and are fixed in mere-ruby.
Also. The unchecked vec_set twin that loop versioning produces had no protect at all on RISC-V; it has the same as vec_set now. And test/parity/map_compact.mere leaves rv_exec_check's known differences at both widths: map_bytes has an arena to measure.
Measured on mere-ruby 770a113 (RV64, --ram 256): 165 and 285 complete (they ran out of memory on 612 and 613); startup leaves 78 MB of heap where 613 left 69. test/parity/container_arenas.mere (compaction, recycle, stores inside blocks, a random walk over twelve containers) prints the same lines on the interpreter, C, Wasm, RV32 and RV64 (LLVM has no map_recycle), and region_reclaim_check.sh --rv-only adds test/regionreclaim/arena_churn.mere: 60,000 stores into compacted and recycled containers inside blocks, 120 MB of values, in an 8 MB machine (v0.1.613 ran out of it).
v0.1.613 — 2026-10-05
_A region block on RISC-V gives its memory back again, behind a high-water mark; on LLVM a string stored into an older Map or Vec inside a block no longer dangles; and Wasm's high-water mark no longer misses a container a callee built and stored outward._
RISC-V. A region block reclaimed nothing on this backend. It once rolled gp back to the block's start, and that was unsound -- a map_set on an outer Map allocated the Map's new parts inside the block's range, and the rollback freed them -- so the rollback was taken out. It is back, with the Wasm backend's answer to the same problem: three runtime words (depth, the innermost block's mark, a high-water mark), a protect after every vec_push, vec_set and strbuf_push (a Map is Vecs since v0.1.611, so it goes through them), and at the closing brace gp returns to max(mark, high-water mark). The result is copied out by a per-type __rcopy_<tag> (str, bytes, a float's box, tuples, records, variants, recursive ones included) twice: above the garbage first, then down to where the block now ends, unless the two overlap. A block whose result is a closure or holds a container runs without rolling back, as every block did before. A failure out of a block skips the rollback, and try_or puts the depth and the block mark back. test/regionreclaim/pertree.mere: 100 trees of depth 14 run in a 16 MB machine, which v0.1.612 ran out of; region_reclaim_check.sh --rv-only is a CI step now.
Every store protects, an int's too. A version of this that skipped int stores freed a queue mere-ruby still held: the monomorphizer erases a leftover type variable to int, and one instance of the prelude's Map insert serves every Map whose values are words or pointers, so "an int" was a pointer there. The region words are per machine, so on bare metal another task's stores raise the mark into its own heap; the mark is only ever raised and a block never rolls back past its own gp, so that makes the block keep more, never less (riscv_bare_shell's background task counts into a Vec while the shell is in its per-command region). docs/bare-metal.md says what a scheduler has to do.
Wasm. Its protect raised the high-water mark only for a container older than the innermost block. A container a callee builds inside the block and stores into an older one is reachable from outside but lies above the block's mark, so when a later push in the same block grew it, the new buffer was above the high-water mark, the rollback took it, and the next block wrote over it: the new parity test read -14468027695667681 for a sum of 12350. A container below the high-water mark is protected now too -- once stored outward, that is where it is.
LLVM had no copy-on-store. Since v0.1.443 a value made inside a block lives in the block, and map_set / vec_push / vec_set stored it as the pointer it was, so a string put into an outer Map inside region R { } was read after R was released: five lines put a string into an outer Map and a Vec, ran one more block, and printed out of memory where the other backends printed v1 wx. The helpers now copy what they store into the container's region -- the C backend's rule since v0.1.30 -- but only while a block is open, so a program without regions runs the code it ran before. A Map copies a key only when it inserts it.
test/parity/region_stores_and_rollback.mere holds all three: results of every copyable kind, stores into older Maps, Vecs and StrBufs, a callee's container stored outward and then grown, nested blocks, a failure out of a block, 40,000 blocks of 1.6 KB garbage each (64 MB unreclaimed, run by rv_exec_check in 32 MB), and an int-keyed Map of queues made inside blocks (the case the skipped int protect broke). With this version it prints the same lines on the interpreter, C, LLVM, Wasm, RV32 and RV64; with v0.1.612, LLVM and Wasm did not.
v0.1.612 — 2026-10-04
_On RV64, str_of_float and float_of_str work on 64-bit words: a float is printed with about 5 KB of heap instead of 280 KB, and 1e+-5 is refused on RISC-V as it is everywhere else._
`str_of_float` (RV64). The limb library spelled a double by building its exact decimal expansion as a digit array and then trying %.{p}g for p = 12..17, parsing each attempt back: about 280 KB of heap and 7 ms on the emulator per call, never given back -- a script printing floats ran into the 256 MB. Now: Dragon4 (Burger & Dybvig's free-format algorithm) on bignums of 30-bit limbs gives the shortest digit string that reads back, the one nearest the value; when it has 12 digits or more that is exactly %.{p}g's answer, and when it has fewer, a fixed-precision pass gives the value's correctly rounded 12 digits (the same thing for a normal double, and not for a subnormal of few bits: the smallest prints 4.94065645841e-324, not 5e-324 -- rv_float_conv.mere caught the first version saying the latter). 1,000 calls: 8 MB and 1.4 s (they did not fit in 256 MB).
`float_of_str` (RV64). The lexical rules are the limb parser's, and anything that is not a plain decimal -- hex, inf, nan, not a number -- is handed to it, so its answers and messages are unchanged. A decimal D 10^k is converted exactly: for k >= 0 the top 57 bits of the product and a sticky bit, for k < 0 57-58 quotient bits of D 2^j / 5^-k by long division with the remainder as sticky, rounded once by the word arithmetic's pack (v0.1.608). Past 800 significant digits the rest becomes one sticky digit, which cannot move a rounding. 1,000 calls: 64 MB -> 32 MB, 1.9 s -> 1.3 s.
`1e+-5` was a float on RISC-V (both widths): the limb parser took any number of signs before the exponent's digits. strtod stops at the e, so it is not a number -- on RV now too, and 0x1p+-3 likewise.
Held to printf / strtod: test/float/rv_dec64.mere (random doubles by class, subnormals of a few bits included, each printed, read back, and read back from a longer spelling) and the existing rv_float_conv.mere now also run at 64 bits in scripts/rv_float_check.sh (with MEMU); 40,000 such lines over four seeds were identical to the C backend when this was made.
v0.1.611 — 2026-10-04
_On RISC-V a Map is a hash table: mere-ruby's startup there takes half the instructions, and a corpus file that ran past two minutes finishes in eleven seconds._
The RV Map was an assoc list. map_set prepended -- even over a key already there, so the list grew with every write -- map_get walked it, and map_iter walked a list of the keys it had seen once per element, the square of the length. Profiled on the RV64 emulator (through the debug map v0.1.610 made correct), that was 64% of mere-ruby's startup and 85-98% of the corpus files that timed out (_mseen, _mfind, __str_eq). C, LLVM and Wasm already had an O(1) Map.
Now the entries are kept in insertion order -- keys, values and a live flag, three Vecs -- under an open-addressing index (a power-of-two Vec: 0 empty, -1 a tombstone, n entry n-1; linear probing; grown at 3/4 full, tombstones counted). The meaning is unchanged and is every backend's: map_iter visits each live key once, in the order it was first set, with its latest value; setting an existing key keeps its place; a key deleted and set again goes to the end. Growing the index never moves an entry; map_compact packs the live entries and repoints the index without hashing again; map_clear / map_recycle empty it. A str key is hashed by a new runtime routine, __rv_str_hash (FNV-1a over the bytes, 31 bits), an int or bool key by mixing its bits; map_set now goes to the _i family by the key type too, as map_get already did.
Measured, mere-ruby on RV64: startup 939 -> 469 million instructions; corpus/67 from past 120 s (24.7 billion instructions) to 11.5 s, matching ruby. test/parity/map_random_ops.mere is a deterministic random walk of set / delete / get / has / len / iter / compact / clear over a small key space (str and int keys) whose output every backend must reproduce; RV32 and RV64 do.
v0.1.610 — 2026-10-04
_mere -rvs / -rv64s (the listing) and -rvg / -rv64g (the debug map) describe the binary -rv emits, word for word -- they were made from a different program, and a profile read through the map blamed the wrong functions._
The listing and the map skipped settling the regions. Q-127 settles every undecided container region (bind_region_params, default_container_regions) before code is generated; emit_program did, and emit_listing / emit_debug_map did not, so they generated code for a slightly different program. On 15 of the parity programs the listing's words were not the binary's; on mere-ruby every address in the map was off -- a sampling profile of its startup on the RV64 emulator, read through it, said 26% was do_class_alias and 7% lgamma's helper, where the truth is the Map (the next version). settle_regions is now one function the three paths share.
A wide listing encoded its jumps as J-type anyway. Past 512 KB of code a jump is emitted as auipc + jalr, and the listing printed it so -- after encoding it as a J-type first, which refuses anything over 1 MB: mere-ruby's listing stopped with "a jump of 2441396 bytes does not fit". It is encoded only when it is one now.
scripts/rv_listing_check.sh (in CI): for every parity program at both widths, each encoded word the listing shows is the binary's word at that address, and each function symbol of the debug map is a listing label at that address; plus a generated 1.25 MB program whose 12,030 jumps are all wide (v0.1.609 refuses its listing). 273 listings, 1.3 million words.
v0.1.609 — 2026-10-04
_RISC-V has a libm: an extern fn cbrt: float -> float (and twenty more of libm's names) is answered by the prelude instead of refused, and those, sin, cos, tan, atan2, exp, log and f_pow are correctly rounded there -- within ~0.52 ulp of the true value, sin 1e22 included._
The libm (lib/rv_libm.ml, appended to the RISC-V prelude). mere-ruby declares twenty-one libm functions as extern fn for its Math module, and on RV32IM / RV64IM every one was a named refusal: Math.cbrt, Math.hypot, Math.gamma, Float#% (fmod) and Math.ldexp stopped the script. An extern whose name is one of atan asin acos sinh cosh tanh asinh acosh atanh cbrt log2 log10 log1p expm1 erf erfc tgamma lgamma (float -> float), hypot fmod (float -> float -> float) or ldexp (float -> int -> float) AND whose declared type is libm's is now bound to __libm_<name> in the prelude (codegen_riscv's libm_bound); on a host the same declaration still links libm. Another type under one of these names is refused as before.
Correctly rounded, not "libm's bits". Each function is computed in double-double -- pairs of doubles carrying ~106 bits, built from the prelude's two_sum / two_prod with new __dd_add / __dd_mul / __dd_div / __dd_sqrt / __dd_exp -- and rounded once. Every function was prototyped operation for operation in Python (whose floats are the same IEEE doubles) and measured against a 50-digit reference over thousands of points: all are within ~0.52 ulp of the true value, fmod and ldexp exact. The Mere transcription was then held to the prototype bit for bit on 7,500 random calls on RV64. macOS's libm, measured the same way, is NOT correctly rounded for most of these -- cbrt, the inverse and hyperbolic functions, hypot, tan, erf, erfc, tgamma, lgamma are a ulp away on 3% to 49% of points (lgamma near its zeros by hundreds) -- so the gate, test/float/rv_libm_ext.mere, holds 502 results (mere-ruby's corpus/168 inputs, the IEEE specials and each function's edges, and random points) to the correctly rounded values in rv_libm_ext.expected, not to the host's. scripts/rv_float_check.sh runs it on RV64 and RV32 (256 MB there) with MEMU. The same 502, from the C backend linked against each host's libm: macOS disagrees on 39; glibc 2.39 (the CI image, arm64 and amd64) on 4, all lgamma, which glibc does not round correctly either -- so a RISC-V build of a program prints what it prints on Linux. erf is a positive-term series below 2 and erfc a 120-level continued fraction above it; tgamma and lgamma are Stirling's series past 12 with the shift in pairs, a zeta series within 0.1 of 1 and 2, and reflection below 0.
The builtins on the same pieces. sin, cos, tan and atan2 moved to lib/rv_libm.ml with double-double kernels; exp, log and f_pow keep their code with a more precise core under it (__fp_exp_core carries r^3/6 in a pair, __fp_log2p the series' first three terms and s's tail scaled by 2/(1 - s^2)). Against libm on the v0.1.606 sweeps: f_pow 47 -> 7 points apart, log 4 -> 0, exp 2 -> 3 -- and on every one of the remaining ten the prelude is the correctly rounded one. Three faults found on the way:
- the sin and cos series stopped at 1/17! and 1/14!, and r^16/16! at pi/4 is
1.6e-15: that was most of their documented "<= 10 ulp";
__fp_pio2_2, a piece of the three-piece pi/2, was written with fifteen
digits, not as fdlibm's double, so pi/2 was 6.5e-26 off and tan of the double nearest pi/2 was wrong in the ninth digit;
- past 2^20 that reduction is not exact, and nothing replaced it:
sin 1e22
answered 7.9e91. Arguments that large now go through Payne-Hanek (2/pi to 1,248 bits, as 24-bit chunks in an if-tree, so nothing is built at startup) and are as accurate as small ones.
v0.1.608 — 2026-10-04
_On RV64, a double is one 64-bit word while it is computed: float arithmetic allocates nothing but its result, 20,000 operations fit in 4 MB instead of 64, and run four times as fast. RV32 is unchanged._
Float arithmetic on RV64 without limbs. Both RISC-V widths computed a float operation in contrib/softfloat's 15-bit limbs -- the width that keeps a limb product inside RV32's signed 32-bit int -- building a record per operand per step: about 2 KB of heap per +, and a region gives nothing back on these targets. mere-ruby's ** cost ~600 KB there. On RV64 the int is 64 bits, so the IEEE pattern is one word and the 53-bit significand fits with room to spare: + - * /, sqrt, the comparisons, f_abs, f_neg, float_of_int and int_of_float are now the same algorithm (pack with guard/round/sticky, long division, digit-by-digit root, every NaN rule) on words, as top-level functions of ints that build no tuple or closure. Measured on the emulator: 20,000 float operations fit in 4 MB (64 MB before) in 0.56 s (2.22 s); 400 f_pow in 8 MB (128 MB) in 1.53 s (3.64 s). The decimal conversions (str_of_float, float_of_str) still go through the limbs.
The two are held to each other and to the hardware: rv_float_ops.mere's 3440 operations give the machine's bits on RV64 as on RV32; a new random-operand sweep, test/float/rv_float_fuzz.mere (zeros, infinities, quiet and signalling NaNs, subnormals, both exponent edges, cancelling neighbours), gives the machine's bits on all 16,000 results except where both sides are NaN; and 320,000 results of the same sweep under eight seeds were identical to the limb library's on RV64, NaN payloads and out-of-range int_of_float included. scripts/rv_float_check.sh runs both RV64 checks when MEMU is set.
`if __rv_xlen () == N` is decided at compile time (codegen_riscv's xlen_test): the branch for the other width is neither compiled nor counted as reachable. That is what lets the one prelude hold both arithmetics -- an RV32 image carries no 64-bit code (its size is unchanged), and an RV64 image no limb code (a float program's image went from 84,285 to 77,441 bytes).
v0.1.607 — 2026-10-04
_On RISC-V, print_bytes writes its bytes (a NUL included) instead of stopping, and a refusal on RV64 names RV64I rather than RV32I._
`print_bytes` on RV32IM / RV64IM was a stub ("needs a host") although a bytes value has been a str block there since v0.1.426 and the write goes by the length word: it is now that write, so a NUL in the middle goes out like any other byte, as on every other backend. docs/host-matrix.md: stub -> yes.
The machine a message names. Every message of this backend was written for RV32I and began with that name, so on RV64 mere-ruby reported "RV32I: fd_pipe is an extern fn". err and emit_abort now spell it RV64I on RV64; the 32-bit spelling is unchanged, because host_matrix.sh and rv_prelude_check.sh key on it. Two texts said --bare on the hosted path too: an unimplemented host service now says "needs a host service this target does not provide (a hosted run answers a fixed set of Linux calls; --bare answers none)", and an extern fn "has no C library to link against -- the program is the whole machine image" (the part mere-ruby matches on is as it was). test/rv/print_bytes_and_names.mere runs at both widths in scripts/rv_exec_check.sh (with MEMU).
v0.1.606 — 2026-10-04
_On RISC-V, f_pow, exp and log answer what libm answers: within 1 ulp everywhere, and the same bits on all but a few points in a thousand. 2.3 ** 3 is 12.166999999999998 there now, as everywhere else. An extern whose parameter is unit can be called with a unit-typed variable on C and LLVM._
`f_pow` / `exp` / `log` on RV32IM and RV64IM (found by mere-ruby on RV64, corpus 105 and 173). These are prelude Mere over softfloat there, and libm on every other backend. Measured against libm on 6000-call sweeps: f_pow disagreed on 3131 points by up to 20 ulp, exp on 303 by 1, log on 598 by up to 2; mere-ruby printed 2.3 ** 3 as 12.166999999999996 and 8 ** (1.0 / 3) as 1.9999999999999998. Now 47, 2 and 4 points, each 1 ulp (f_pow's worst is 0.71 ulp from the true value; on most of the 47 libm is the closer one). Three changes, all built from exact two_sum / two_prod (Dekker) pieces and rounded once at the end:
- the log's HEAD+TAIL pair carries the rounding of
s = (m-1)/(m+1), which
was up to 4e-17 -- a whole ulp of log near 1, and y times that in f_pow;
__fp_exp2takes its argument as a pair, sums 1 + r + r^2/2 exactly, and
scales the argument's tail by all of e^r: the first version scaled it by 1 + r and was six ulps off at |y ln x| = 190, where the tail is 1e-14;
- integer exponents |y| <= 32 walk the binary powering as pairs. The plain
walk stays as the first step: it is exact where the answer is exact, and outside [2^-960, 2^960] (where a pair cannot be split) its answer stands. The cost is time and memory: a f_pow is about 2.3x the work, and on these targets every float operation allocates (about 2 KB through softfloat's limbs, and a region gives nothing back), so a f_pow now takes ~600 KB of the heap instead of ~300 KB. test/float/rv_libm_points.mere holds 66 points of the three to libm's bits -- points that 605 got wrong -- and scripts/rv_float_check.sh runs it on RV32I when MEMU is set; RV64 gives the same bits.
An extern's `unit` parameter is dropped by its type (found while wrapping mere-ruby's externs). The C prototype drops a unit parameter by its declared type (extern int getpid(void);), but the call dropped only a literal (), so fn (u: unit) -> getpid u emitted getpid(mu_u) and clang refused the program; LLVM declared @getpid() and called it with an i64, which linked by luck. Both calls now drop the argument by the parameter's declared type, and a non-literal one is still evaluated, ahead of the call. test/ctests/extern_unit_param_variable.mere (clang) and a test_basic case (the LLVM call) hold it.
v0.1.605 — 2026-10-03
_A chain of literal ++ is one literal by the time it is compiled, and a list is copied along its spine by a loop: ten million elements no longer run the copy out of C stack, and a 6703-line text written one literal per line no longer costs the square of its length at startup. On Wasm, a region block's result no longer overwrites itself on the way out._
`"a" ++ "b"` is folded (found by mere-ruby). A text written as a chain of literals -- mere-ruby's Ruby prelude, one line per literal -- was concatenated when the program ran, every link copying everything before it: the copies came to 192 MB at every startup of mere-ruby, which is most of what -e 1 used, and on RISC-V, where a region gives nothing back, more than the 256 MB a program has there. fold_str_concat (lib/pipeline.ml) runs after parsing with the other lowerings, so mere fmt, which keeps the sugar, still prints the chain as it was written. Two literals become one; x ++ "a" ++ "b" folds its tail; "a" ++ x is left as it is. mere-ruby's -e 1: 244 MB -> 46 MB (with the joined prelude built in a region). Three unit tests used two literals to look at the code ++ makes; they put a variable on one side now.
A recursive variant's copier walks its spine in a loop. __mcopy_<tag> (copy-on-store) and __mdeep_<tag> (a region loop's carry) recursed into every node, and a list is one node per element: three million elements stored into a map ran out of the default 8 MB stack ("stack overflow (recursion too deep)"), and mere-ruby's Array.new(10_000_000) out of its 512 MB one -- in CRuby's test_method that was rest(*(1..10_000_000)), and the whole file with it. A constructor whose payload is a tuple ENDING in the type itself (a list's Cons) is now walked as a loop along that field, each new node linked into the slot the previous one left; the other fields and every other constructor are copied as before. scripts/deep_list_check.sh stores a 3M-element int list and str list into maps and carries one through a region loop under an 8 MB stack, and --poison puts the recursion back and requires a failure; in CI. The LLVM copier still recurses (a ten-million-element list carried out of a region overflows there); Wasm's is not measured.
On Wasm, a region block's second copy ran into its first (shown by the fold). A Wasm region block copies its result twice -- above the block's garbage, then down to where the block began -- and when the garbage is smaller than the value, the second copy writes over the first while it reads it: a record whose first field is a string had the string's bytes written over its second field (region R { rec_t { name = "abcdefgh", n = 4 } } read n as 1751606885, the bytes of "efgh"; a longer one ran out of memory). It took a string LITERAL to show it -- a string built in the block was garbage enough to keep the copies apart -- and folding turned region_result_boxed's "na" ++ "me" into one, and the parity gate went red on Wasm. When the copies would overlap, the first is now the result and the bump is left after it. test/parity/region_result_over_its_copy.mere; it fails on v0.1.604.
v0.1.604 — 2026-10-03
_A zero divisor fails on RISC-V as everywhere else: 7 / 0 printed -1 there, and 7 % 0 printed 7. And on the -rv path a second type error is shown in the program, not in the prelude._
RISC-V's div and rem do not trap: a zero divisor answers -1 and the dividend, which is defined behaviour for the instruction set and a silent wrong answer for the program -- the interpreter, C, LLVM and Wasm all fail with "division by zero" / "modulo by zero". Each division now branches on a zero divisor to one of two shared stubs (emitted only when the program divides) that fail with those messages, catchably: try_or and try_or_msg take them like any other failure. INT_MIN / -1 needed nothing: RISC-V answers the wraparound the interpreter does.
And on the same path: with two or more type errors, every one was shown as <rv-prelude>:2:9 with a line of the prelude under it. The report re-checks the file as it was given -- without the prelude glued in front -- and its positions were then moved again as if they counted from the prelude's first line; they are the file's own now (scripts/rv_prelude_check.sh gains the case).
Found while giving this backend try_or_msg (v0.1.599). test/parity/divide_by_zero_caught.mere is the same on all four backends and, through scripts/rv_exec_check.sh, on RV32 and RV64. Every RISC-V binary that divides changes; no other backend's output does.
v0.1.603 — 2026-10-03
_dom_canvas_put_pixels: a browser program puts a whole frame on a canvas in one call (Q-122)._
contrib/dom's canvas could set a fill colour and fill a rectangle, so a renderer drew a frame one 1x1 rectangle -- one host call -- per pixel, and the note beside the bindings put a blit off until "typed-buffer FFI". That never came, and was not needed: a bytes already crosses the Wasm boundary as its length and then its bytes, and Wasm has had ByteBuf since v0.1.585 to build one in.
dom_canvas_put_pixels canvas px w h draws a w x h image at the canvas's top-left corner from px, four bytes a pixel (RGBA), row by row; the glue copies them into an ImageData (one cannot be made over a view of a threaded module's shared memory) and calls putImageData. Fewer than w h 4 bytes is refused by name -- dom_canvas_put_pixels: a 2x3 image needs 24 bytes, got 16 -- rather than drawn short.
scripts/run_dom_headless.mjs gains a canvas whose 2d context records what was drawn, and scripts/dom_canvas_check.sh (in CI) builds test/dom/canvas_put_pixels.mere and checks the frame arrives pixel for pixel and the short one is refused. m3d's browser build and mbrowse's browser output were the two programs waiting on it.
v0.1.602 — 2026-10-03
_A top-level name defined a second time is warned about: everything above the second definition -- including a let rec ... and group -- still means the first._
Since v0.1.588 top-level functions may come in any order, and a second definition of a name shadows the first for what follows it. Nothing said so. Splitting mere-ruby's big let rec groups into separate definitions showed what that costs: nine names quietly came to mean an earlier definition of the same name -- three of them a function of the same type, which no type error can catch (Q-157). The warning is the precondition for doing that split safely:
warning: `f` is defined again at the top level (first at line 1). From here on it
means this one; everything above -- including any `let rec ... and` group that
uses it -- still means the first. Rename one if that is not the intent
It is given on every path (check, the interpreter, the compilers) when both definitions are the program's own: in the same directory, so two libraries with a private helper of the same name do not warn in every program that imports both, and not inside an installed package (.mere_modules/), which its user cannot edit. A definition that keeps a let fn promise is not a second one, and a program's own version of a prelude name is not warned about. In this repository and its downstream programs it names 6 files; in mere-ruby, 7 names (all_digits, hex4, hex_val, is_ws_ch, norm_ccc, rev_app_str, str_has_nul) -- which are what to settle before its groups are split.
v0.1.601 — 2026-10-03
_channel_sender ch h: a channel that is told who sends on it stops a receive from waiting once all of them have finished -- naming the sender that failed. On all four backends._
Since v0.1.586 a failed thread no longer takes the program down, and a channel did not know its senders: a receive waiting for a message a failed thread was going to send waited for good, with nothing to wake it. Joining the thread first fixes the one-shot workers (and par_map, v0.1.590); it cannot fix a producer that lives as long as the program -- an actor, a broadcaster, an accept loop.
channel_sender ch h registers thread h as a sender of ch. Once every registered sender has finished and the channel is empty, a receive stops waiting as it does on a closed channel:
channel_recvfails, catchably:channel_recv: sender thread 2 failed: fail:
boom -- the first registered sender that failed, with its message -- or channel_recv: every sender has finished and the channel is empty;
channel_recv_optandchannel_recv_timeoutanswerNone.
A channel with no sender registered behaves exactly as before, and a program that never calls channel_sender compiles to the same C as before. It is opt-in on purpose: the alternatives the investigation weighed -- closing every channel a failed thread held, counting senders, waking every receiver on any failure -- each broke a program that is correct today.
How each backend sees a sender end: the interpreter wakes the channels a thread was registered on when it ends; C and LLVM read the thread record (v0.1.586) and, with senders registered, look again every 50 ms while waiting; Wasm keeps each thread's status in the first 512 bytes of its private region of the shared memory, where any thread can read it, and a channel there takes at most 256 senders (past that, channel_sender fails by name: a worker cannot allocate the list a bigger one would need). test/parity/channel_senders.mere holds the answers -- values sent before a failure still received, the failure named, the all-finished case, and a timeout answered at once -- the same on all four.
v0.1.600 — 2026-10-03
_channel_close, channel_recv_opt and channel_recv_timeout on LLVM and Wasm, and a Wasm thread sees the program's top-level values: parity checks concurrency_channel and channel_unconstrained_elem on all four backends, which leaves 13 programs unchecked on LLVM and 6 on Wasm._
The three were interp and C only, and they are what a worker loop uses to end: mere-blog's app had no other gap on Wasm.
- LLVM. The channel gains a
closedword.channel_closesets it and wakes
every waiter; a send on a closed channel and a channel_recv on a closed, drained one fail catchably with the C runtime's messages. All receives go through one routine that waits (with a deadline, for the timeout) and reports whether it took a value, returned by value so a receive in a loop allocates no stack.
- Wasm. The host's channel gains
closedand a sequence word every send
and the close bump; a waiter waits on the sequence, so a close that lands between its last look and its wait still wakes it. One host call does every receive and writes the value into memory the module names.
- The option. Neither backend can build an option from inside a runtime
call, so before emission channel_recv_opt ch becomes let p = __chan_recv_pair ch in if fst p then Some (snd p) else None (Ast.rewrite_channel_ops), every new node typed from the original's, and a type nothing constrains erased to int as these backends erase it.
- A Wasm thread read every top-level value as 0. Top-level values are Wasm
globals, one set per instance, and a spawned thread is a new instance: a top-level function called from a thread read a channel at address 0 or a count of 0. A spawning program now exports them and the host copies their values into each new thread's instance. test/parity/thread_reads_toplevel.mere.
test/parity/channel_close_family.mere covers what each answers -- queued values surviving a close, the two failures, a timeout that expires and one that does not, and workers ended by a close -- the same on all four backends. Every program's LLVM IR changes (the channel runtime is in each); C, RISC-V and -t output do not.
v0.1.599 — 2026-10-03
_The RISC-V backend takes Maps keyed by int or bool, a local let rec ... and ..., a top-level function at another arity, and try_or_msg: mere-ruby builds for RV64 again (it stopped at the first of the four)._
mere-ruby's RV64 build -- the interpreter compiled for a CPU that is itself a Mere program -- had been refused since mere-ruby keyed its object table by object id. Fixing that showed the next refusal each time; there were four.
- Maps keyed by `int` or `bool`. The prelude's Map is an assoc list compared
with str_eq, which read an int key as a string's address, so such a key was refused. A Map whose key type is int or bool now goes to an rvmap_*_i family that compares with ==, chosen from the Map's own type so map_len and map_iter (no key argument) agree with map_set.
- A local mutually recursive group. Only a single local
let recwas
compiled; a group allocates every member's closure block and binds it before filling any captures, so members that call each other read each other's blocks.
- A top-level function at another arity:
f a bfor anfof three, a bare
f passed as a value (both refused, "no currying layer"), and f a b c d for an f of three. The first two are the closure the source means -- the given arguments evaluated once, in order, straight into a closure block whose code calls f with them and the rest -- built from values rather than frame slots, so the enclosing function's frame is unchanged. The third is the saturated call with the remaining arguments applied to its result. A function of one argument goes through its adapter as before.
- `try_or_msg`. It is
try_orwith a handler where the default was; the
unwind leaves the failure's message in a1 and the catch path calls the handler with it. In a program that calls try_or_msg, a fail the program wrote tags its message fail: as every other backend does, so the handler reads the same text -- and the prelude's own fails, which stand in for builtins' failures (random_int, map_get), do not, as theirs are untagged everywhere. A program that does not call try_or_msg keeps the bytes it had (its uncaught message is still printed untagged on this backend).
Four parity programs cover these (map_word_keys, local_mutual_rec, partial_top_fn, try_or_msg_catch); interp, C, LLVM, Wasm, RV32 and RV64 print the same. Of the 1,022 programs swept, every one that built for RISC-V before builds to the same binary, except two whose functions come out in a different order (the same 172 functions, each the same size: the table they are emitted from grew), and 63 that were refused now build.
v0.1.598 — 2026-10-03
_file_exists, sleep_ms and random_int on LLVM, and random_int's failure names the bound on C as it does everywhere else._
The three were the first wall for a dozen downstream programs on LLVM (file_exists alone for mengd, mreg, mraft, mk and mere-ruby). They give the C runtime's answers: whether the path names something (access(F_OK), since stat would need the platform's struct layout in IR), a sleep that does nothing for a count of zero or less, and a uniform int in [0, n) from rand, seeded once from time ^ pid. A libc function the program also declares as an extern is declared once.
random_int 0 failed with "random_int: bound must be positive" on C and with "... (got 0)" on the interpreter, Wasm and RISC-V; C and LLVM say "(got 0)" now. test/parity/file_exists_random_int.mere holds what of this is deterministic; docs/host-matrix.md shows the three as yes on LLVM.
v0.1.597 — 2026-10-03
_A tuple pattern nested in a let compiles on every backend: parity checks every program on C (206 of 206), and leaves 15 unchecked on LLVM and 8 on Wasm._
let ((a, b), (c, d)) = t in body ran in the interpreter and was refused by C, LLVM and Wasm ("nested pattern in let-tuple not supported ... use match"). The backends now see it as one flat let per level -- let (__lt0, __lt1) = t in let (a, b) = __lt0 in let (c, d) = __lt1 in body -- written by Ast.flatten_let_tuples after type checking, each new variable typed from the value's tuple type. A subtree with nothing to rewrite stays the same node, so the tables keyed on nodes still find it. It applies inside functions as well as at the top level, to any depth, with _ and functions as components (test/parity/nested_let_tuple.mere).
The C, Wasm, RISC-V and -t output of every other program in the repository and downstream is unchanged (4,086 comparisons).
v0.1.596 — 2026-10-03
_The LLVM backend erases a leftover type variable like the other backends, lifts a let rec written in the program's body, and stops mistaking externs for captures and t0 for a register: parity leaves 16 programs unchecked on LLVM (was 19), mpng builds with it and passes its 33 checks, and 42 programs that clang or the emitter refused now build._
- A type variable that survives to emission -- the result of a function
that never returns, a helper whose error type nothing constrains -- was refused ("unsupported LLVM codegen type element: 'a") where C and Wasm erase it to int. ty_tag erases it here too (llvm_ty_of already lowered it to i64), and the instantiation pass names the erased instance a residual call site asks for, as it does for C and Wasm (recover_promoted: without the concrete instances C and Wasm also get, which this backend finds through its own table). result_residual and free_result_thrown_away are checked on LLVM now. A tuple holding such a variable is declared under its erased name when nothing in it is a type constructor (let (f, g) = (fn x -> x, ...)).
- `let rec` in the program's body -- not inside a function -- was refused
("let rec inside an expression"); it is lifted under the host $main as the C backend has done. It was the first wall for mpng and medit2. region_outer_push_realloc is checked on LLVM, and test/parity/main_let_rec.mere covers a capture, mutual recursion and a let rec inside a lifted one.
- Externs are not captures. An inner function that called an
externtook
it for a free variable and passed a local that does not exist ("use of undefined value '%tcp_write'"); the C backend fixed the same thing in v0.1.61.
- A parameter named `t0` collided with the temporaries
fresh_regnames
%t0, %t1, ... ("multiple definition of local value named 't0'", mpng's chunk out t0 t1 t2 t3). Such names get a suffix with a dot, which no Mere identifier has.
Of 1,022 programs (test/, examples/, the downstream programs), 42 whose LLVM IR the emitter or clang refused now build and none that built stops; the C, Wasm, RISC-V and -t output of all of them is unchanged.
v0.1.595 — 2026-10-03
_floor, ceil and round on Wasm, and a negative zero printed as -0.0 there: contrib/raster's path library builds for Wasm, and parity leaves 9 programs unchecked on Wasm (was 10)._
The Wasm backend refused all three. floor and ceil are f64.floor and f64.ceil; they were held back because putting them in the eta-expansion list sent it into an infinite expansion, so a call is emitted directly and the bare name as a value is still refused. round is not f64.nearest, which rounds half to even where C and the interpreter round half away from zero: it is the truncation, moved one away from zero when the part cut off is at least a half -- computed from the truncation rather than as floor(|x| + 0.5), which rounds 0.49999999999999994 up because the sum is 1.0 in a double.
The Wasm host printed -0.0 as 0.0 ((-0).toFixed(0) is "0"); it prints -0.0 like the other three.
test/parity/float_round.mere covers halves, negatives, the 0.49999999999999994 case, negative zero and a value past 2^52. contrib_ctest unpins raster/path; docs/host-matrix.md shows the three as yes on Wasm.
v0.1.594 — 2026-10-03
_Five more places where building a program took time quadratic in its size are linear: mere-ruby's -c is 3.8 s (was 7.3) and its check 1.4 s (was 1.8), and every output is the same bytes._
Each was found by a program of one repeated shape, 2,000 and 16,000 times over; scripts/infer_scaling.sh now measures all five, and v0.1.593 fails each of them (42-85x for 8x the program).
- The top-level environment (
checkand every emitter). A program is a chain
of lets, and the environment at the thousandth one is a list a thousand long: every name bound early was found by walking past everything bound since. 16,000 top-level let g = map_new () took 17.9 s to check and 24 s to emit C. The type checker and the move checker now keep that chain as a persistent map, extended by each top-level let for its body (Env_index.extend) and moved along by the declaration loop (Env_index.track); a lookup walks only the bindings in front of it. check 0.38 s.
- The C emitter's variables in scope were the same list again; a map now.
-c 0.59 s.
- Instantiating polymorphic helpers. Asking where a helper is used visited
every instance of every multi-instantiated function; the instances' uses are filed by name, read back in the order the old walk produced them, which is the order that decides which instance the C lists first. 8,000 helpers used at two types: 43 s → 1.2 s.
- Lifting inner functions. Each lift turned every top-level name into a set
again (94 s for 16,000 functions with an inner let rec); the set is kept. 1.2 s.
- A long operator chain, as generated code writes one: its left spine was
collected by appending each link to the ones below it. 16,000 links: 1.6 → 0.08 s.
The C, LLVM, Wasm, RISC-V and -t output of 1,022 programs (test/, examples/, the downstream programs) and mere-ruby's C are byte-identical to v0.1.593's.
v0.1.593 — 2026-10-03
_Three walks in the C emitter that were quadratic in the program are linear: mere-ruby's emit is 6.6 s (was 7.8), and the C it emits is the same bytes._
resolve_vec_let_typeswalked the rest of the program once perletwhose
container type was still open. It is now one walk that records each binding's uses, honouring shadowing, and then unifies them in binding order -- the order the old per-binding walks ran in.
- The set of free variables a closure skeleton uses, and the bound names inside
free_vars, were lists searched with List.mem; they are a hash table and a string set.
The C for 1,022 programs (test/, examples/, the downstream programs) and mere-ruby's C are byte-identical to v0.1.592's.
v0.1.592 — 2026-10-03
_The C a standalone program compiles to has internal linkage, and a curried function's closure stages call its uncurried twin: mere-ruby's C is 40 MB (was 93), clang -O2 builds it in 54 s (was 185) in 1.4 GB (was 6.9)._
A top-level function f = fn a -> fn b -> fn c -> body is emitted twice: as f__direct(a, b, c), which every saturated call uses, and as the closure chain that a partial application or a first-class use goes through. The chain copied body into every stage and every stage's two-argument entry -- a number of copies that grows like the Fibonacci numbers with the arity (6 for three parameters, 56 for ten) -- and in a standalone program all of it had external linkage, so clang kept and optimised every copy whether anything reached it or not. mere-ruby's C was 92.9 MB, 59 MB of it closure stages; clang -O2 took 185 s and 6.9 GB and produced a 24 MB binary. Developers waited on clang for over 90% of a build.
- Internal linkage. The program's own definitions -- top-level functions,
their stages, their _as_value constants -- are static in a standalone program as they already were in a library: nothing outside the one translation unit links to them. Except under -g, where a debugger breaks on a function by its name and an unreferenced static one is dropped even at -O0 (scripts/debug_info.sh caught it).
- The chain calls the twin. The curried form's innermost body is now a
saturated call of the function itself, which the emitter sends to the twin, so the body exists once. Only where the twin takes no region arguments and no parameter is named like the function.
mere-ruby: C 92.9 → 40.2 MB, clang -O2 185 → 54 s, 6.9 → 1.4 GB, binary 24 → 5.4 MB, its emit 12 → 7.8 s; its corpus matches Ruby on 276 of 276. parity, examples_parity, debug_info, lib_check and downstream_cc_check (23 programs) pass.
v0.1.591 — 2026-10-03
_A function whose result is a free type, used at two types, is instantiated at both: the C build of such a program no longer fails in clang._
raise_exc = fn (msg: str) -> fail msg has type str -> 'a. Used for its value at v by one caller, it was resolved at v; called from a function whose own result is thrown away (so its type variable is erased to int), that call named the same v instance, and clang refused the C: "returning 'v' from a function with result type 'long long'". Seven lines reproduce it, and mere-ruby meets it as soon as its giant let rec groups are split. The recovery pass that already handles residual uses of multi-instantiated functions now also handles a function resolved at one type whose residual use agrees on every parameter and differs only in the result: it becomes multi-instantiated, and each call names the instance for its own type. A residual variable in a PARAMETER position is left alone -- erased to int it would name an instance whose closure parameter does not match (contrib raster's comparators showed it).
The same pass now reads an index of the emitted bodies -- each name's uses and its live references, filed per body in the order the walk visited them -- instead of walking all of them once per function per round (the new check would have made mere-ruby's -c 75 s, and a 16,000-member group 12 s). Every .mere in the repository and downstream emits byte-identical output on all five emitters (5,110 comparisons); test/parity/fail/result_free_two_types.mere holds the new case (it fails on purpose: 3, then fail: 500).
scripts/procsig_check.sh (v0.1.589) failed on CI only: a CI runner starts its steps with SIGPIPE ignored, and a shell cannot undo a disposition it was started with, so the scenes that need nothing inherited were the inherited one. They put the default back with perl first.
v0.1.590 — 2026-10-03
_What v0.1.586 left behind: par_map joins before it receives (a failing element hung every backend), a thread's failure is printed when it happens (a daemon never reached the exit that printed it), and C's closed channels fail catchably._
v0.1.586 stopped a spawned thread's failure from ending the program. Three consequences of that surfaced afterwards:
- `par_map` hung. It is lowered to a thread and a channel per element and
received without joining, so a failure in its function left the receive waiting for a message nobody would send -- on all four backends, and in a real program (mk). Every element's thread is now joined before its result is taken: the failure is raised in the caller, where try_or takes it, and par_map threads are no longer reported "never joined".
- A daemon's failures were silent. The failure line of a thread nobody
joined or detached went out at exit, and a server whose handlers' handles are dropped never exits. The line now goes out when the thread fails, whoever holds the handle -- Rust's panic message, the same four-backend text: mere: thread 1 failed: fail: boom. join still raises the failure again.
- C aborted on a closed channel.
channel_recvon a closed, empty channel
and channel_send on a closed one were abort() (exit 134, past every try_or) where the interpreter fails catchably. They fail by name now.
The fan-in examples join their workers before reading what they sent (parallel_compute, parallel_channel, concurrent_loops, pubsub), as the docs now advise. scripts/thread_fail_check.sh expects the new line and gains par_map_fail (four backends) with a poison that drops the join.
v0.1.589 — 2026-10-02
_A recycled map keeps its block's real capacity; no out-of-memory fail while the region lock is held; the pin that keeps an arena for a suspended coroutine follows the pointers it finds and remembers who it is for; a signal's disposition, a refused write to stdout and socketpair(2), which no extern can reach._
`map_recycle` ran off the end of the block it kept (found by mere-ruby). A recycle keeps the OLDEST block of a map's private arena and calls it the 4 KB seed. But a value bigger than a quarter of the seed gets a dedicated block (v0.1.307), chained in behind the bump block -- so while the seed is still the bump block, the oldest block IS the dedicated one, and r->cap = 4096 claimed 4 KB on its 2 KB: the next allocations ran into the heap beside it. In mere-ruby's frame pool that was a 2.8 KB exception message, a corrupted length, and a deep copy that asked malloc for the impossible. The recycle now keeps the block's own capacity (b->pad). test/uaf/recycle_dedicated.mere, built with ASan by scripts/region_uaf_check.sh (a case marked a), and its poison -- the old r->cap = 4096 -- goes red. (A map was the only container with a recycle; the region cache's release already kept best->pad.)
...and that malloc failure left the program stopped, silently. An allocation in the default region takes its lock, and a malloc that failed inside __lang_region_alloc -- a dedicated block, a new bump block, an in-place grow -- called __lang_fail_impl with the lock still held. A try_or above caught the fail; the next allocation waited on the lock its own thread held, at 0% CPU, forever. Each of the three now unlocks before it fails (add_block has a _try form for the one under the lock). Not under a gate: it needs a malloc that fails, which a test cannot ask for portably.
Two holes in the pin (v0.1.547), found by running a collection on a coroutine of its own. mere-ruby now collects in the middle of a method by switching to a coroutine and compacting there, so the method's stack is a suspended one -- the case the pin exists for -- and the first corpus run under a collection at every safepoint was a use-after-free in both of these ways.
- The pin read only the stack's own words. A stack holds a pointer to a node
-- a list cell, a tuple, a closure's env -- and the node holds the pointer into the arena being compacted. That arena was freed under the node. The pin now gathers what the stack REACHES, by the walk coro_scan_ints already makes: as far as the pointers go inside the coroutine's own regions, three hops into anyone else's (__lang_pin_reach). Once per stop, as before.
- A retired arena was freed under the coroutine that pinned it. Every later
retirement tries the retired arenas again, and asked only the SUSPENDED stacks -- so once the coroutine that pinned an arena was running again, and still holding the pointers that had pinned it, the next compaction anywhere freed the arena. (For a fiber: suspended across a compaction, resumed, and a map compacted while it runs.) A retired arena keeps the handles of the stacks that pinned it and is not tried while one of them is running (__lang_coro_pinners).
- ...and so a compaction under a pin does nothing. With the pin following
pointers, a suspended stack points into nearly every store, and compacting one retired the old arena and made the copy beside it: the bytes twice, for as long as the stack kept the word. mere-ruby's two thousand suspended fibers went from 6.9 s / 1.1 GB to 90 s / 6.5 GB. Compaction has no observable behaviour, so map_compact / vec_compact now leave a container where it is while a suspended stack points into its arena (a recycle still moves, since it must empty the map). The same walk also stopped being made twice: a suspended stack's coro_scan_ints gathers its pin as it goes. The benchmark is back to 6.65 s, at 1.66 GB -- what the sound pin keeps that the old one let go.
test/uaf/coro_pin_reach.mere and coro_pin_resumed.mere, ASan builds in scripts/region_uaf_check.sh, each with its poison (the reach skipped; a pinner's run not counted) going red. ⚠ The first version of the reach test stayed green under its poison: the string's address was still on the stack, in a dead frame's slot inside the live range, and a conservative pin rightly took it. The cell is made at the bottom of a 300-deep recursion now, so the copy is below where the stack pointer comes back to.
`proc_sig_noop`, `proc_sig_default`, `proc_sig_raise`, `proc_out_errno`. signal(2) and sigaction(2) move a disposition through a function pointer and a struct, so a program could only reach sigignore(2) -- SIG_IGN -- and an ignored signal STAYS ignored across exec(2). mere-ruby ignored SIGPIPE at its first pipe, every child it started ignored it too, and since the runtime dropped a write it could not make, a child loop { puts :ok } whose reader had gone printed into nothing for 94 seconds of CRuby's test_io, where ruby's child dies of SIGPIPE at once. ruby catches SIGPIPE with a handler that does nothing (a handler is reset to the default at exec) and keeps a disposition it inherited if that is not the default; proc_sig_noop is that, answering 0 installed, 1 kept, -1 refused. proc_sig_default and proc_sig_raise end a process of a signal, as an unrescued SignalException ends ruby. And __lang_write_all -- the write behind print_no_nl and print_bytes -- keeps the errno of a refused write for proc_out_errno (asking clears it), so an EPIPE is something a program can see. The proc_ family's errno slot (`proc_last_errno`) as for the others; emitted only when a program declares one of the names. `scripts/procsig_check.sh` (22 rows, in CI) runs each in the scene it is for -- nothing inherited, SIG_IGN inherited, a pipe whose reader has exited, a run that must end of the signal (exit 141) -- and four poisons (SIG_IGN for the handler, the inherited disposition not put back, the errno not kept, on both writers) each turn a row red. The third poison first went unnoticed: an un-restored handler also survives a raise, so the row that tells the two apart is a child's -- it inherits an ignore and not a handler.
`sock_pair`. socketpair(2) fills an array, so it joins sock_bind and the rest of the sockaddr family on the runtime's side: an AF_UNIX pair, both descriptors in one int (first * 2^20 + second), type 1 stream / 2 datagram / 5 seqpacket, -1 with fd_last_errno. mere-ruby's UNIXSocket.pair raised NotImplementedError, which seven of CRuby's test_io copy_stream cases met first. scripts/sockaddr_check.sh has nine more rows (82): the two ends talk, a closed end is the end of file, a datagram pair, an unknown type's EINVAL.
v0.1.588 — 2026-10-02
_Top-level functions may be defined in any order (Q-156): a function can call one defined below it, and two can call each other from separate declarations._
A top-level function could only call one written above it. The way round was to put everything in one let rec ... and ... -- mere-ruby's interpreter is a group of a thousand -- which also makes the members monomorphic in each other (Q-157), or a let fn promise per name, which almost nobody wrote (one, in four repositories).
Ast.order_toplevel runs once the top-level names are unique. Evaluation order does not move: a value keeps its place relative to every other value. A function definition, which does nothing when it is reached, is placed after the declarations it refers to (or earlier, when a value needs it), and functions that refer to each other across separate declarations become one group -- only those; a group the program wrote is never split. A value that needs a function that needs a later value cannot be scheduled and is left as written, with the unbound variable it always had. mere fmt prints the source order.
A program with no forward reference comes back exactly as it was, so nothing that compiled before changes: every .mere in the repository and in 18 downstream projects emits byte-identical C, LLVM, Wasm and RV32 output and the same -t answer as v0.1.587 (5,110 comparisons), and so does mere-ruby. The Q-156 probe retires. test/parity/toplevel_any_order.mere (203 programs) and four unit tests.
v0.1.587 — 2026-10-02
_mere fmt keeps the comment at the end of a constructor's line (lost comment lines 43 -> 26), and bytes_of_str is a view instead of a copy while no region block is open (Q-117)._
fmt (Q-154). A type written one constructor per line was printed on one line, and the comment at the end of a constructor's line had nowhere to go: a Top_type carries no positions. The parser now records the line of each constructor, and a type with a comment on one of them is printed one constructor per line with the comments back; every other type prints exactly as before. The gate's count also stopped counting // inside a string literal ("http://"): a line is a comment line when // is left on it after its string literals are removed. scripts/fmt_comments_check.sh: 7,527 comment lines in the examples, 26 lost (ceiling 43 -> 26; 29 by the old count).
bytes_of_str (Q-117). A str and a bytes have the same layout on every compiled backend -- the length, then the data -- and neither can change, so the str's own memory is the bytes. A line tool that handed each line to a SIMD load (which takes bytes) copied every line to do it: 2,000,000 calls on a 200-byte line peaked at 417 MB on C, and take 1.4 MB as a view. Only while no region block is open: inside one the str may be the block's and the bytes may be returned out of it, so it is copied into the current region as before. test/parity/bytes_of_str_view.mere holds both (202 parity programs).
v0.1.586 — 2026-10-02
_A spawned thread's failure has one meaning on all four backends (Q-090): join raises it again, a detached one is a line on stderr, an unclaimed one is a line at exit, and the exit status is the main thread's. C reports leaked threads too (Q-055)._
A plugin host reached for spawn to survive somebody else's code, and the four backends gave three answers. C and LLVM ended the process from the failing thread, racing the main thread's last write: the exit status was 1 in some runs and 0 in others, and stdout was sometimes cut. Wasm printed the failure and exited 0. The interpreter said nothing. A daemon whose detached handler failed once was gone -- mhttpd and mengd both.
Now every backend gives Rust's answer:
- the thread's failure is recorded, and `join` raises it again in the
joiner -- uncaught, it is the program's failure (fail: boom, exit 1); try_or takes it;
- a detached thread's failure is one line, and the program carries on:
mere: thread 1 failed (detached): fail: boom;
- one nobody claimed is one line at exit:
mere: thread 1 failed and was never joined: fail: boom;
- the exit status is the main thread's.
C and LLVM keep a record per thread { pthread, number, state, claim, message } on a list while nobody has claimed it; the thread body runs under the failure jump, so its failure lands in the record instead of in exit(1), and the record is freed when both the thread and its handle are done with it. Wasm's workers set a flag in their own instance, so an uncaught fail there hands its message to the host instead of printing it; the host keeps the record in the shared buffer the join already waited on. detach was refused on LLVM and Wasm; it is lowered on both now. LLVM's failure-message buffer is per thread.
The record is also what the leak report reads: MERE_THREAD_REPORT=1 now works on C, LLVM and Wasm (Q-055) -- still running, finished, never joined, died: <message>, never joined (only the interpreter can say what a thread was blocked on).
⚠ One consequence: a thread waiting on a channel for a message a failed thread was going to send now waits for good -- a channel does not know its senders -- where before the failure ended the process. scripts/spawn_stack_check.sh had exactly that shape (its "a fail on a spawned thread is that thread's" case), and hung; it now has main's try_or run while the worker fails and then joins it.
And a second: a worker POOL that let a handler's failure end the worker would lose a worker per failing request and then stop answering, silently. The two pools in contrib/http (serve_rd, serve_mt) now catch a handler's failure themselves -- one line on stderr, the connection closed, the worker back for the next request -- and with_rescue is still what turns it into a 500 (scripts/http_concurrency_check.sh checks the difference).
scripts/thread_fail_check.sh is rewritten from "pin the split" to "pin the contract": four programs on four backends (spawned_fail ten times each), the report, and eight poisons -- the thread not catching its own failure, join not raising it, a detached failure not told, the exit hook not telling, on C, LLVM and Wasm. scripts/thread_leak_check.sh gains the C column (the count; the wording is printed, since it races the real clock).
v0.1.585 — 2026-10-02
_ByteBuf, read_bytes and write_bytes on LLVM and Wasm (Q-016): mgunzip now decompresses byte-for-byte the same on C, LLVM and Wasm._
ByteBuf[R] -- random-access bytes, the buffer a binary tool builds its output in -- and the two whole-file bytes builtins were C and interpreter only, so a program like mgz's mgunzip ran on two backends of four.
- LLVM:
%mere_bytebuf = { data, len, cap, region, owner }, the C struct
field for field, with i64 lengths (an i32 one would make bytebuf_get bb 4294967297 a question about index 1). Allocated in the region the result type names; growth through the region's in-place grow; a bad index fails through the index-failure path, so try_or catches it with C's message. The owner field is set at creation, the two writes check it and the reads mark it shared, as on C (v0.1.575/582). read_bytes / write_bytes are the C runtime's.
- Wasm: a 16-byte header
{ data, len, cap }like StrBuf's, from the bump
allocator (a region block reclaims what it made), zeroed by hand (memory a block gave back is not zero), growing in place when it can and copying -- and protecting the copy -- when a block is open. One thread, so no owner check. read_bytes / write_bytes reuse the existing bytes host imports; a missing file fails catchably with the interpreter's sentence.
- Each runtime is emitted only when used: the Wasm output of 474 of 521 programs
(examples and parity) is byte-identical to v0.1.584's, and wasm_size_check is unchanged.
- Found on the way, on both:
write_file_bytesdid not check the 0..255
range (300 was written as 44 on LLVM, masked on Wasm). It fails with C's message now; test/parity/bytebuf_edges.mere had been asking since v0.1.279, on the two backends that refused it.
Parity: 201 programs (bytes_file_roundtrip is new: 512 bytes, every other one a zero, through a ByteBuf and back). Unchecked on LLVM 23 -> 19, on Wasm 13 -> 10: bytebuf_edges, region_bytebuf_reclaimed, region_rec_capture and (LLVM) raster now MATCH. mgunzip on a 2 MB member: identical output on all three, 4 MB peak on LLVM.
v0.1.584 — 2026-10-02
_of_json into a container is a type error at the call (Q-118), not a failure when the checkpoint is read back._
to_json writes a Vec as an array, a Map as an object and a StrBuf as a string; of_json can decode none of them -- a container has a region and an identity, and a decoder makes values. The program that met this was the one that checkpointed its state (usually a Map), wrote it without complaint, and failed reading it back: of_json: expected a variant value for Vec, at run time.
of_json, of_json_opt, of_json_like and of_json_opt_like whose target is a container -- or holds one, through a tuple, an option or list, a record field or a variant payload -- are now refused where the call is typed, naming the container and the way out (decode the contents as a list and rebuild it). A top-level binding is checked as soon as it is typed, because the interpreter runs it right after; the whole program is checked again once inference has settled targets that only a later use decides. Four unit tests.
v0.1.583 — 2026-10-02
_A curried inner function called with all its arguments builds no closures on LLVM and Wasm either (Q-142): mandelbrot takes 6 MB on LLVM (was 235) and runs on Wasm (was out of memory)._
An inner function is lifted to the top level with what it captures prepended to its parameters. The C backend has emitted an uncurried twin for it since v0.1.52 -- captures and every curried parameter in one call -- and sent a saturated call there. LLVM and Wasm had the twin for top-level functions only (Q-139), so the inner go i zx zy of a mandelbrot went through the curried chain: one closure environment per application per iteration. The same loop cost 16 bytes an iteration at top level and 104 nested in a function on Wasm; LLVM allocated 6,984,016 bytes for 96,000 iterations, and examples/mandelbrot.mere peaked at 235 MB there against C's 6 MB and ran out of memory on Wasm.
Both backends now emit <lifted>__direct beside the curried definition (which a partial application and a first-class use still reach) and send an exactly saturated call to it, a self tail call as musttail / return_call. The nested loop is 16 bytes on LLVM -- the whole run -- and 16 B/iter on Wasm, equal to the top-level loop (Wasm boxes floats). mandelbrot: LLVM 6.2 MB, Wasm runs, output identical to C.
scripts/inner_direct_check.sh measures it on all three compiled backends: the nested loop against the top-level one, and a partial application that must still give the right answer. MERE_NO_INNER_DIRECT=1 turns the twin off on every backend, and --poison uses it: C 10 MB, LLVM 7 MB, Wasm 10 MB.
v0.1.582 — 2026-10-02
_The rest of Q-179: a read by another thread freezes a container, an OwnedVec is moved rather than shared, sync type cannot vouch for a Map, a stateful --lib refuses overlapping calls, and the interpreter checks a program that spawns before it runs any of it._
v0.1.575 failed a container WRITTEN by a thread other than the one that made it, and left reads free so a table built once could be read by every thread. That left the race it could not see: the owner writing while another thread reads. A Map churned by its owner under a reader segfaulted 3 runs in 3 on C; a Vec compacted under one was a use-after-free.
- Freeze on first foreign read (interpreter, C, LLVM). The first time a
thread other than the owner reads a container, it becomes shared -- a bit in its owner field, so a write needs no second test -- and a shared container is read-only for good. The owner's next write fails with its own sentence (a Map another thread has read was written). The four races, 20 runs each on every backend that builds them: 200 of 200 fail by name, no signal. A table built once and only read still works. The rule refuses "threads read it, are joined, then the owner writes" -- the same answer the static check gives a captured container. A range-check-versioned loop's unchecked reads leave no mark and need none: its guard reads the Vec's length through the checked vec_len.
- OwnedVec has an owner.
spawngives away the OwnedVecs its closure
captures (the move the type check allows) and the new thread takes them; any other thread's use fails by name. Two threads pushing to one through a launcher's parameter used to lose elements or abort.
- `sync type` cannot hold a builtin container. The marker vouches that a
type is safe to share; a Map, Vec, StrBuf, ByteBuf, ListBuf, OwnedVec, File, ThreadHandle or Coro inside it (through tuples and unmarked user types) has no lock for it to stand for, and is refused at the type's line. No program in the repository or downstream used sync type.
- `--lib` with module state refuses overlapping calls. A library whose entry
file has a top-level container returns MERE_FAIL for a call that overlaps another thread's, naming why; four host threads filling a top-level Map had kept 191,819 of 200,001 entries. A library without module state still takes concurrent calls.
- The interpreter checks, then runs. It typed and ran each top-level
let
in turn, and the borrow and spawn-capture checks came after the last one -- so race_global, which every compiled backend refuses to build, ran its race on the interpreter first. A program that spawns is now checked whole before any of it runs (its warnings are not repeated).
scripts/owner_check.sh has eight new programs and nine new poisons (a read leaves no mark, the write does not say why, the guard's vec_len leaves no mark, OwnedVec use unchecked, spawn does not give the OwnedVec away -- on C and LLVM); scripts/lib_check.sh a stateful library under four host threads.
v0.1.581 — 2026-10-02
_A thousand-member let rec group stops costing a walk of the whole program per member: mere-ruby checks in 1.8 s (was 7.7) and emits C in 14 s (was 113)._
mere-ruby's interpreter is one let rec ... and ... group of about a thousand functions, and four passes were quadratic in it:
- The typer found a variable by walking the environment, and inside the
group the environment is every binding so far plus the group -- 4,226 names, 1,184 steps per lookup on average. While a group of 32 or more is inferred its environment is indexed once (Env_index); a lookup walks only what was consed in front of it (the member's own parameters and lets) and then asks the table.
- The move checker walked the same kind of list, and uses the same index.
- The unused-binding warning (
mere check) resolved names in a list too; it
is a map now (2.2 s at 25,600 top-level lets, 0.25 s at 16,000).
- The C, LLVM and Wasm instantiation search asked, for every function on
every pass, for its uses -- and each asking walked the whole program and every resolved body. One walk now indexes each body's uses by name, in the order and with the shadowing the walk had, and the question reads that.
Ast.walk also compresses the chains it follows (a link is only ever set on an unbound variable and never undone, so this changes no answer).
Nothing a program means changes, and that is checked rather than argued: every .mere in the repository and in 18 downstream projects emits byte-identical C, LLVM, Wasm and RV32 output and the same -t answer as v0.1.580 (5,045 comparisons), and so does mere-ruby (92 MB of C). mere-ruby: check 7.73 → 1.77 s, -t 10.5 → 2.8 s, -c 113 → 14 s, -ll 131 → 6.1 s.
scripts/infer_scaling.sh measures the group now, with -t and with -c, and the independent bindings under check. v0.1.580 fails it (64x and 66x for 8x the members); this release is 7-8x.
v0.1.580 — 2026-10-02
_parity.sh stops counting a miscompile as "unsupported", and the two it was hiding are fixed (Q-198, Q-199)._
A parity row whose backend refuses at emit time is UNSUP and passes, because a backend that does not implement a builtin yet is a documented limit, not a bug. The test for "refuses" was the word unsupported in the error -- and the LLVM and Wasm backends put that word in front of every codegen error they raise. So unbound variable: p on LLVM read as a documented limit. emit_kind now names the sentences that mean "not implemented here" (no lowering yet, not in the codegen subset, not available on Wasm, a region loop, ...) and anything else is a failure. Two rows went red at once.
An inner function's name met a variable of another function's (Q-198): an inner let rec cnt in one function and a plain let cnt = vec_new () in another lifted as one symbol, and LLVM -- and Wasm in a variant -- read the variable inside the other function's closures as the lifted function. mgz's inflate and deflate both have the shape (p, cnt); a captured variable named len failed the same way. The pass that renames colliding inner functions knew only other function names; it knows every variable's now. Three parity cases (200 in all).
A closure that calls an inner-lifted function, on Wasm: the lambda did not carry the callee's captures, so a closure written inside a region block that calls such a helper failed "inner-lifted capture v not in scope" -- the C backend's fix from v0.1.453, never ported. It had been in parity since then, counted as unsupported. It carries them now, and msha's sha1 builds on Wasm -- and so do four contrib libraries contrib_ctest.sh had pinned as Wasm gaps for the same sentence (font/font, html/entities, html/tokenizer, proto/gen), which are unpinned: a pin that starts passing is a failure there.
v0.1.579 — 2026-10-02
_v0.1.575's owner check made a versioned loop 2.8x slower once a thread existed; it is asked once per loop now._
v0.1.575 put the write-owner check into every write, the unchecked ones range-check versioning makes inside a loop included. Its A/B measured programs that never spawn -- where the check is one load of a global -- and stopped there. Once any thread has been started the check reads the thread id, which on Darwin is a call, and it is a branch out of the loop that kept clang from vectorizing it: axpy with a join (spawn ...) in front took 0.24 s where v0.1.573 took 0.09.
The unchecked writes carry no check now. The loop's guard -- evaluated once, before the fast copy -- asks __vec_owned of every container the fast copy WRITES, and a thread that does not own one takes the checked copy, whose vec_set fails by name as before. axpy with a thread started is 0.08-0.09 s on C again; LLVM 0.20 s (0.18 at v0.1.573). vec_reverse, vec_sort and lb_to_list write too and were not checked at all; they are, on the interpreter, C and LLVM. owner_check gains a reverse case and a versioned-loop case, and a poison that makes the guard stop asking.
And a guard that was refusing loops it should have taken: with a stride above one that lands on the bound exactly, the last index visited is N - stride, not N - 1, but the guard asked N - 1 + w <= len -- so a stride-2 loop over an exactly-sized Vec never took its fast copy. axpy_simd: 0.11-0.13 s to 0.08-0.09, same output. range_version_check holds both sides of that edge: a stride-2 loop over an exact Vec is dispatched, and one over a Vec a slot short fails at the same index as before on every backend.
v0.1.578 — 2026-10-02
_The self-hosted lexer reads every reserved word as a keyword (Q-153), and the new gates are in CI._
contrib/parser/lexer.mere read twelve of the words the compiler reserves -- view, region, import, trait, impl, derive, signature, drop, using, open, for, dyn -- as identifiers, and its parser matched TIdent "import", so a program the compiler refuses (a reserved word used as a name) went through it. They are one token now, TKw with the word, and the parser's two places that read import and region by name read the keyword. scripts/selfhost_lexer_keywords_check.sh (CI) DERIVES the list from lib/lexer.ml's keyword table, as the operators' gate derives operators, so a word reserved later is asked about without anyone remembering; its poison plants a word the self-hosted lexer does not know.
CI also runs owner_check (v0.1.575), lsp_coalesce_check and lsp_folding_check (v0.1.576). The README's test count is 2882.
v0.1.577 — 2026-10-02
_mere fmt keeps 50 more trailing comments: 93 lost lines over the corpus become 43 (Q-154)._
Two places, neither needing an end position. The first branch of an if chain did not take its trailing comment while every later branch did -- one call fmt_else_chain makes and fmt_if_multiline did not (26 lines). And a comment written after the ; that ends a declaration is now attached there, the last line the declaration's tree reaches standing in for its end (24 lines); where a closing delimiter sits alone on a later line the stand-in is short of the real end and the comment is left where it was, which is what happened to all of them before. fmt_comments_check's ceiling is 43; fmt_roundtrip still finds every formatted file re-reads and formats to itself. What is left is classified in the gate: variant lines in a type, let ... in, list literals reflowed across ,, nested if/else, and // inside strings.
v0.1.576 — 2026-10-02
_The language server stops re-checking keystrokes that are already out of date, and answers folding (Q-023, Q-147)._
Every didChange re-checks the whole program, and on mere-ruby's main.mere (44k lines, 93k with imports) a check takes 10.5-10.7 s. The server handled messages one at a time, so N keystrokes queued N checks -- a word typed bought a minute of diagnostics for text that no longer existed. A didChange carries the whole document, so a later one for the same file makes it worthless: before a didChange is handled, the messages that have already arrived are read, and an unbroken run of didChanges for one document is handled as its last. Nothing else moves -- a request between two changes is answered against the text it was asked about. Knowing what has "already arrived" needs the bytes in a buffer this code owns (an in_channel can hold a whole message that select does not see), so the server reads its own input now. scripts/lsp_coalesce_check.sh (CI): ten changes sent at once are two publishes, not eleven; a hover between two changes sees both checked; the same ten changes paced 0.3 s apart are eleven, which is the control that says the count counts.
textDocument/foldingRange is answered: a top-level declaration that spans lines folds, and so does a run of two or more comment lines. A declaration has no end position, but the parser knows where one ends -- the tokens one call of parse_decls consumed are exactly one declaration -- so it keeps each one's first and last line for the file being parsed (Parser.decl_spans), at no cost beyond walking the tokens once. scripts/lsp_folding_check.sh (CI) holds a three-line declaration and a comment run folding and a one-line declaration not.
v0.1.575 — 2026-10-02
_A container written by a thread other than the one that made it stops the program by name (Q-179 stage 2)._
v0.1.574 follows what a spawned thread reaches through functions whose definitions the program shows. A closure that arrives as a PARAMETER is not one of those, and that is the shape mere-blog's race had: its handler was handed to http_serve_mt_ctx, which spawns it, and the handler pushed to a Vec the program's thread made.
So Map, Vec, StrBuf, ByteBuf and ListBuf record the thread that made them, and a write from any other thread fails with one sentence, the same on the interpreter, C and LLVM, instead of hanging a probe loop (a Map filled from four threads, three runs in three) or losing entries (159,123 of 160,000). Reads are not checked: a table built once and only read is shared safely, and a read is the hot path.
What it costs. The program's thread is thread 1 and nothing asks a thread its number until something has been spawned, so a program that never spawns pays one load of a global per write -- on Darwin a thread-local read is a call, and vec_set sits in inner loops. On benchmarks/ (C, median of 21 runs on a loaded machine) crc32 is +1.4%, wordfreq +4.4%, the rest inside the noise; every output identical. The interpreter checks the same writes with the domain id. A --lib build turns the check off at mere_lib_init: a host calls in from whichever thread it likes, one call at a time, and module state is written from each. OwnedVec is single-owner by construction and is not checked.
The struct sizes the LLVM runtime had written as numbers (24 for a Vec and a StrBuf, 32 for a ListBuf) are computed from the types now that the types grew. scripts/owner_check.sh (CI) runs a Vec, a Map and a StrBuf written through a parameter on all three backends, plus a program that shares a read-only table and must still run; its poisons take the check out of the emitted C and LLVM.
v0.1.574 — 2026-10-02
_A spawned thread is checked for what it REACHES through functions, not only what it mentions (Q-179)._
A function's type says nothing about what the function touches (TyArrow _ -> true in both is_send and is_sync), so spawn (fn () -> effect "x") with effect pushing to a top-level Vec was accepted. mere-blog shipped exactly that: eight HTTP workers pushing to one Vec, and the same push from eight spawned threads kept 159,123 of 160,000 entries (that standalone shape is refused now). Q-080 recorded the hole in 2026-08, and it was left open while mere-ruby's fibers depended on it; they have run on coroutines since 2026-09-28.
Now every function the spawned expression mentions whose definition the program shows -- a let or let rec binding, a partial application of one, a data literal holding one, a module member -- is followed to what it mentions, and a value that is not Sync at the end of that walk refuses the spawn. The message names the path (through w > fill > put), because the offending binding is usually several calls from the spawn.
A container nothing writes is shareable: if every occurrence in the program is the first argument of a read builtin, threads may read it together. mengd's inflate tables -- built with vec_of [...] at the top level, read by a thread per connection -- are that case, and nothing in them changes. (Range-check versioning turns a proved vec_get into __vec_get_unchecked before this pass runs, which made the tables look written until that name was on the list.)
What it does not see is a function value whose definition is not visible -- a parameter, a channel's message, a call's result -- so a library that spawns the handler it was given is not checked at its spawn. mere-blog's race is that shape: its handler reaches the Vec, and the handler is handed to http_serve_mt_ctx, which spawns it. Its pre-fix source is still accepted here; it was fixed in that repository (2a1bc87), and a run-time check is the next stage for exactly this reason. Nothing in the 712 programs of this repository or the 19 downstream repositories is refused. The test that pinned "a Map inside a captured closure is not seen" as accepted pins it refused; its neighbour, the parameter case, stays accepted and says why. A program with no spawn skips the pass (mere-ruby's check: 7.97 s before, 7.91 after).
v0.1.573 — 2026-10-02
_spawn checks what any argument mentions, and a ByteBuf is not shared across threads (Q-179 stage 0, Q-197)._
spawn's capture check looked at one shape: a lambda written in place. Every other argument went through as an ordinary application -- spawn (w slot), a partial application whose argument is a Map, was not looked at at all, and four threads filling one Map that way hung the C build in three runs of three (a concurrent resize breaks the open-addressing probe loop). Now any argument is checked for what it mentions, the way a lambda's captures are: spawn (w slot) is refused with the sentence a captured Map already got. A function value alone (spawn body) still passes -- an arrow type says nothing about what the function reaches, which is the next stage's question (Q-179).
ByteBuf was missing from the Send and Sync classifiers' container case while the value restriction and the region list had it, so a top-level ByteBuf written by a spawned lambda was accepted as shareable. All three now read the one list of region-parameterised containers. Nothing in the 311 examples or the 19 downstream repositories was relying on either.
v0.1.572 — 2026-10-02
_An annotation's region name is a variable unless a block of that name is open around it (Q-196)._
fn (b: ByteBuf[R]) -> .. is how the documentation writes a function over any buffer -- patterns.md says to "write R as a type variable", the tutorial shows fn (v: Vec[R, int]) -> vec_len v -- and it was not one. The name went into the type as the region NAMED R, and that is the region of every region R { } in the program. Pass a buffer to such a function, open an unrelated block called R later in the same function, and the compiler said the buffer escaped R.
mengd and mtar have done exactly that since they were written, and stopped compiling at v0.1.564: that version made a let-bound ByteBuf stop generalizing (Q-193, a soundness fix), which had been hiding the collision. A Vec, never generalized, failed the same way in every version. Neither repository was in test/downstream/REPOS, so nothing here said so for eight versions.
Now an uppercase region name in an annotation -- a parameter's or (e : t) -- names a block only when a block of that name is open around the annotation; otherwise it is a variable, one per name within the annotation, as 'r always was. Inside region R { } an annotation's R is still that block, and a container annotated with it still cannot leave it. mengd and mtar as they were written compile again (both have since been changed to 'r, which means the same thing on every version). The two tests that pinned the old printing -- (Vec[R, int] -> int) for an annotation outside any block -- print the variable now, and the tutorial and patterns.md say which is which.
test/downstream/REPOS gains mengd, mtar, mvm, mhttpd and mreg (nineteen rows), and test/downstream/CC compiles all five.
v0.1.571 — 2026-10-02
_The downstreams' emitted code is compiled now, and the first run found two LLVM bugs._
downstream_check asks whether mere -c succeeds in each of fourteen repositories; nothing asked whether a C compiler then accepts what came out. The runtime is text pasted into every program, so a change to it reaches every downstream at once and every gate that stops at emission is blind to it -- v0.1.566-567 shipped such a change and a mere-ruby build went looking.
scripts/downstream_cc_check.sh (CI) compiles each repository's emitted C with that repository's own flags (test/downstream/CC), and the LLVM IR of the four the LLVM backend emits whole; nothing is linked, so only headers are needed. A failure says which of three things it was -- a compile error, a timeout, or a kill (signal 9 is usually the out-of-memory killer; mere-ruby and mbrowse peak at 6-7 GB) -- and --poison holds each sentence to its cause. A row whose code, compiler and flags hash to an earlier pass is reused rather than rebuilt (the hash is of the code, not of which files a commit touched), which CI keeps in a cache. Each row prints its seconds.
Its first run failed one row: mwasm's LLVM IR did not build (use of undefined value '@mere_vec_int_get'). The positioned-file runtime calls the Vec[int] helpers, and file_openrw, file_size, file_fsync and file_close pulled that runtime in without registering the instance, so a program whose only file builtins were those had no helpers to call. The instance is now decided once from the runtime blocks being emitted. Built, mwasm then failed "out of memory" on its first ++: args () handed out argv's own pointers, from when a str on this backend was a plain NUL-terminated pointer, and strs have carried a length header since -- str_len of the argument hello was 32194994947519841. Arguments are copied into the default region now (never freed, so a region block around the caller cannot take them away). mwasm built on LLVM prints what it prints on C. scripts/llvm_host_args_check.sh (CI, poisoned) holds both.
And it found what killed clang in v0.1.566-567. The first full run reported mere-ruby's row as a compile error, and it was not one: clang's driver survives when its frontend is killed and exits 1 itself, printing Killed: 9 -- the gate reads that message now. What killed it is in a log: mere-ruby's mspec/rss_guard.sh, which a test sweep runs beside it, kills -9 any process over 6 GB whose command line contains mere-ruby, and a clang compiling .../mere-ruby.c peaks at 6-7 GB. The guard logged that clang's whole command line. v0.1.569 said nothing on record explained the kill of 2026-10-01; this is the explanation that fits it (the same file, the same signal, a sweep running), though that night's log is not there to say so. The gate names its temporary files after the row, not the repository.
The budget the plan had for peak RSS is not here: the same compile's peak moved by a fifth between runs on a loaded machine, and a budget that flaps is a gate nobody reads.
v0.1.570 — 2026-10-02
_mere -c --region-sites: which source lines filled the default region._
The default region is never given back, and a container goes there when nothing decided its region -- most often a buffer a function makes and uses but does not return, so it is in no type (Q-134). MERE_REGION_STATS could say how much went there, not from where. mgit's default region was 11.5 GB; finding that 8.3 GB of it was one line, mgz's 64 KiB inflate window made once per git object, took 23 rebuilds with one allocation site at a time moved to the current region.
Compiled with --region-sites, each container the program sends to the default region gets a region of its own named after its source line, which forwards every allocation to the default region (the forwarder v0.1.563 introduced) and counts what passes through it -- the struct, the storage, every growth and every copy stored into it. MERE_REGION_STATS=1 then lists the lines, most bytes first: mgit's list starts mgz/inflate.mere:341: alloc_total=8347724352, in one run. Where memory lives does not change, so the output and the default: total are the same with the flag and without it; without the flag the emitted calls are what they were (the runtime carries the forwarding check either way, a branch on a field __lang_region_live already reads). C backend only.
scripts/region_sites_check.sh (CI) holds all three on a small program of mgit's shape and goes red when a line is not charged or keeps its own memory. It was run on the CI image before it was pushed.
v0.1.569 — 2026-10-02
_Four gates that answered for their environment, and a correction to v0.1.568._
CI was red from v0.1.563 to v0.1.568, every time on one line: region_uaf_check --poison said the LLVM poison for kept_then_written stayed green. The poison frees a region struct the fix keeps, and whether the program then notices is the allocator's business: on macOS the freed bytes are reused and the output goes wrong; on glibc they are still there and the writes land in live memory. The poison now fills the struct with 0xAA.. before it frees it, and goes red on both (a segfault on Linux, a hang the gate's alarm ends on macOS). Poisons can be written as perl: expressions, which can add a line where sed cannot. MALLOC_PERTURB_ was tried and is not enough. The gate is run on the CI image before it is pushed this time; v0.1.563 added it without that.
tty_raw_check's Ctrl-C leg failed whenever the gates were started from a background job: a non-interactive shell starts & jobs with SIGINT ignored, an ignored signal survives exec, and the program under test never died of the Ctrl-C it was sent. Its pty driver restores SIGINT in the child, and scripts/bounded.sh restores SIGINT and SIGQUIT for every command it runs. Load had nothing to do with it: 24 of 24 runs failed started with &, and 24 of 24 passed with SIGINT restored, under the same 30 busy loops.
fmt_comments_check read fmt_roundtrip_check's probes: the latter writes examples/__fmtroundtrip_$$__.mere next to each example and deletes it, the former globs examples/*.mere, and _ sorts first -- so POISON 6's "first example that formats" was sometimes a file that was about to vanish, and the plain legs counted 289 files or died of a syntax error on an empty count. The probes are dot files now (.fmtroundtrip_$$.mere), which no *.mere glob picks up -- the shape other gates already use. Four other gates glob examples/ and are covered too.
downstream_check gives each repository 900 s instead of 420, and prints how long each took. v0.1.568's CI failed a second way: mere-ruby's main.mere outlived 420 s. The same mere-ruby commit had taken 222 s in v0.1.567's run fifty minutes earlier, and v0.1.562 and v0.1.569 emit it in the same time here (120 and 123 s of CPU) -- the runner, not the compiler. mere-ruby is about three times the size it was when 420 was chosen, so the margin had already gone.
The reason v0.1.568 gave is withdrawn. It said the inlined pool check made clang -O2 get killed on mere-ruby's C "every time". That was seen twice, while the full gate sweep and a simulation ran on the same machine; clang peaks at 6-7 GB on that file. On a machine doing nothing else the C that v0.1.567 emitted compiles and links (187-206 s, 5.2-6.0 GB peak, three ways), and nothing on record says what killed it. The noinline stays -- it costs nothing measurable and keeps the pool out of every creation site -- but it fixed nothing that was shown to be broken.
v0.1.568 — 2026-10-01
_mere-ruby builds again: the stack pool's size check made clang -O2 give up on it (v0.1.566-567)._
v0.1.566 made a pooled stack carry its size and the pool check it on the way out, a call to __lang_fail_impl (which does not return) inside the take loop. Inlined into every coroutine creation, that sent clang -O2 over a cliff on mere-ruby's 93 MB of emitted C: the compiler was killed after 136 s, every time, where the same file without the check compiles in 168 s. Nothing in this repository's gates compiles a downstream's C -- downstream_check asks whether mere -c succeeds -- so it shipped twice; a build of mere-ruby against v0.1.567 is what found it. The pool functions and the failure path are out of line now (noinline, C and LLVM): the pool is not the hot path, and churning 64 coroutines at a time costs what it did (0.06 s for 2000 rounds). mere-ruby builds in 164 s and its corpus matches.
v0.1.567 — 2026-10-01
_A recycled or compacted container that escaped its block lost its struct (Q-195)._
A Map or a Vec can move its storage to an arena of its own -- map_recycle, map_compact, vec_compact -- and after that its region names that arena, while the struct the handle points at still lives where it was made. Runtime retention (v0.1.557) decided from region alone: for a map made in region R { }, recycled, and captured by a closure that outlived R, it saw an arena that is not a block, did nothing, and R was freed under the handle (the closure's next write segfaulted on C). A Map and a Vec now carry home, where their struct lives, and their copier keeps both. LLVM has no map_recycle or compaction and is not affected.
Found by simulating Q-134's proposed fix -- every container whose region nobody decided placed in the runtime current region instead of the default one -- on mere-ruby, whose frame pool recycles every map it hands out: corpus program 100 segfaulted, a bisection over the 1,867 allocation sites and AddressSanitizer (with the region cache off, so a freed block is really freed) named the route.
region_uaf_check gains recycled_escape and a poison that drops the home retention; ROUTES gains its row.
v0.1.566 — 2026-10-01
_coro_new_sized, and a stack pool that follows the program's peak instead of stopping at 16._
The pool. A finished coroutine's stack was kept for the next one, up to 16 per thread; past that every coroutine paid an mmap, a guard mprotect, a first-touch fault and a munmap, about 10 us. A server with fifty connections in flight is past it whenever more than 16 finish together -- the likely cause of mpoll_coro making a coroutine per connection serving fewer requests than its worker pool. A list now keeps up to the most coroutines the thread has had alive at once (at least 16). That is memory the program already had at its peak, so the peak does not rise; the pool only stops giving it back below it. madvise on the way in was measured and left out: 0.6 s more for no RSS saved on stacks that touch a page or two.
| 2000 rounds of K coroutines alive, then finished | before | now |
|---|---|---|
| K = 16, C | 0.02 s | 0.01 s |
| K = 64, C | 1.0 s | 0.05 s |
| K = 64, LLVM | 0.16-0.23 s | 0.06 s |
`coro_new_sized : int -> (Coro['m] -> 'm -> CoroExit) -> Coro['m]` makes a coroutine on a stack of at least that many bytes, rounded up to a power of two of at least 64 KiB, at most 1 GiB; the pool keeps one list per size. Only the pages a stack touches are resident, so a small stack saves address space rather than memory -- for a program that keeps very many coroutines alive, or a big one for one that recurses deeply. Overflowing a small stack is named like any other. Rewritten before type-checking like coro_new (partial application and the bare name included); interpreter, C and LLVM; Wasm and RV refuse it by name.
LLVM maps its stacks with mmap, as C does (the flags are Darwin's or Linux's, told apart at run time by the weak symbol the overflow handler already uses). posix_memalign made a page resident for every stack before anything ran on it, so a suspended coroutine cost two pages on LLVM and one on C: 10,000 of them were 335 MiB against 171. They are the same now (164 MiB).
coro_check gains sized (both sizes, a partial application, a size refused), sized_overflow and sized_reuse -- a small stack finished first must not be handed to a coroutine of the program's own size -- with a poison on each backend that makes the pool ignore the size. A pooled stack carries its size beside the pool's link word and taking one checks it: a stack of another size would put the stack pointer outside its mapping, and what that does depends on what is mapped there -- the first version of this poison stayed green in one full run out of two, which is how that was found.
v0.1.565 — 2026-10-01
_A finished coroutine keeps nothing on LLVM either: generational handles and an env of its own, the C runtime's v0.1.558._
On C, v0.1.558 made a finished coroutine free everything (Q-183). LLVM still handed the program the record's address, so the 24-byte record could never be freed, and a body's env went to the default region: a million coroutines made and finished held 33.7 MiB, and 49.7 MiB when the body captured something.
Now a handle on LLVM is what it is on C -- a slot and a generation, 24 bits each, below 2^48, resolved through a per-thread slot table -- and reaping frees the record and bumps the slot's generation, so a handle from before answers "has finished". Coro and CoroExit are i64 in the IR (the message word needs no conversion for a Coro any more). The lambda written as coro_new's argument has its env malloc'd as the coroutine's own and freed when the body ends; any other closure's env is where it was.
| a million coroutines made and finished | C | LLVM before | LLVM now |
|---|---|---|---|
| no captures | 1.48 MiB | 33.7 MiB | 1.50 MiB |
| capturing an int and a bool | 1.48 MiB | 49.7 MiB | 1.52 MiB |
coro_check's LLVM cap drops from 128 MiB to C's 16 MiB, gains a poison that keeps the record (47 MiB, red), and its finished-check poison now removes the check that a stale handle resolves to nothing.
v0.1.564 — 2026-10-01
_The last inner-function shape reaches its caller's region, and a let holding a ByteBuf is not generalised (Q-193): mgit's never-freed allocation drops from 34.7 GB to 11.5 GB._
v0.1.562 gave inner functions region parameters of their own and left one shape out: an inner function of one member of a let rec ... and group calling another member. That is the shape of mgit's object store -- store_read reaches loose_read through an inner in_packs, and the three are one group -- so mgit's reads still went to the default region.
The group's marks. When an inner binding is generalised, a region variable it does not quantify belongs to something that outlives it, and its "allocation" mark is removed so a call site inside a block cannot decide it. A group's own variables are in that position too -- they appear free in an inner function of a member -- and losing the mark meant the group could not make them region parameters. Now a variable keeps its mark when it is the group's: level 1 or deeper, and in the type of a let rec group still being inferred (both places that infer groups -- the typer's Let_rec and the pipeline's top-level infer_top_rec -- say which those are). Every other variable is unmarked as before. v0.1.562 tried keeping every mark, and a call inside region BLK then decided a region that mere-ruby's benv shares with a global pool; those are level 0, and this rule leaves them alone.
ByteBuf in the value restriction (Q-193). A let is generalised unless its type mentions a mutable container, and the list had Map, Vec, OwnedVec, StrBuf, Channel, ListBuf and Coro, and not ByteBuf. So in let (out, _e, ok) = zlib_inflate raw hp in, out's region became a fresh variable at every use, and the region the inflater allocated the buffer in was decided by nobody: the default region. The same program written with a Vec passed the caller's region.
Neither alone moves mgit; together, counted by the runtime on the Mere repository (17,505 objects, output identical to git):
| default region (never freed) | blocks (given back per object) | |
|---|---|---|
| v0.1.563 | 34.7 GB | 1.8 GB |
| v0.1.564 | 11.5 GB | 24.8 GB |
What is left in the default region is what no function's type names -- mgz's Huffman tables and bit readers, made and dropped inside the inflater. That is Q-134 proper.
region_params_check asserts all four inner-function shapes as ^param down to the emitted calls (four on C, four on LLVM) and a ByteBuf let (test/regionparams/bytebuf_let.mere); both go red on v0.1.563.
Swept, with imports resolved, against v0.1.563: 982 files compile under both, none differently; 111 compile under neither (unvendored imports, deliberate refusals) and are not counted. mere-ruby's main.mere compiles under both. (The sweep v0.1.559 reported -- 1,072 files, no difference -- measured nothing: it timed each run with timeout, which this machine does not have, so every file exited 127 under both compilers.)
`mere install` names the other cause of an integrity error (Q-194). Before v0.1.400 a package's content hash covered the installed tree's .git, which is different in every clone, so a lock entry written then fails on every fresh clone with "content changed since mere.lock". Four repositories had such entries (their hash lines are re-computed now; the revs did not move). For a whole-repository package -- a subdir never had a .git to hash -- the error now says so, and that deleting the lock and installing again keeps the revs, which mere.toml pins.
v0.1.563 — 2026-10-01
_Four use-after-frees the types could not see: a retained container written after its block, a thread's env, an OwnedVec element, an LLVM channel message._
All four were accepted by every backend and safe by every verdict in test/escape/ROUTES, and all four read or wrote freed memory once the block's arena had been reused. They were found by enumerating, route by route, what can reference a value after its block ends -- the survey that R2 (Q-134) needs before it can put more allocations in blocks.
A container retained by v0.1.557, written after its block (Q-190). v0.1.557 fixed the closure route (Q-188): a container captured by a closure stored somewhere older marks its block "kept", and the block's memory is handed over instead of freed. Then the block's struct was re-initialised and put back in the cache -- but the escaped container still names that struct as its region, so a later vec_push allocated in whichever block took the struct next, and was freed with it. The witness for v0.1.557 only READ the value afterwards, which is why it passed. Now a kept block's struct is never reused: it stays as a forwarder to the region that adopted its memory, and allocation and growth through it follow the forward (C: fwd; LLVM: field 5, and the struct is no longer freed). One struct per kept release is the cost.
A thread spawned in a block (Q-191). spawn handed the thread its closure's env as it was; a str captured inside region R { } was read by the thread after R was reused. The thread now gets a copy, in the default region, through the env's copier -- as coro_new does since v0.1.558. (LLVM cannot express this today: detach and a block returning a ThreadHandle are unsupported there. It marks the open blocks kept, the conservative form it uses for closure stores.)
An OwnedVec element (Q-192). An OwnedVec is malloc'd and outlives every block, and owned_vec_push stored the element as it was. It is copied into the default region now, on C and LLVM, as every other container copies what it stores into its own region.
An LLVM channel message. The C backend copies a message into the message's own region; LLVM sent a str as the pointer it was, and the receiver read the churn's bytes. It is copied into the default region now.
scripts/region_uaf_check.sh (CI) runs each route after forcing reuse and compares with the interpreter on C and LLVM; --poison undoes each fix in the emitted code and must go red (six poisons). ROUTES gains the four routes as SAFE rows and bytebuf_out as GATED (the region-escape rule refuses a ByteBuf out of its block; the codegen-side container check does not list ByteBuf, and this row is what says the type rule covers it).
v0.1.562 — 2026-10-01
_An inner function takes region parameters of its own._
A top-level function's allocations follow its caller's region (v0.1.464): the region travels in as a leading argument of its __direct twin. An inner function's did not -- a let rec in a body, a let-bound lambda, a lift inside a lift -- so whatever it allocated, and whatever the functions it called allocated, went to the default region even when the function around it had been handed a block:
let pr = fn (n: int) -> let v = vec_new () in let _ = vec_push v n in v;
let sr = fn (n: int) ->
let rec ip = fn (k: int) -> if k == 0 then pr n else ip (k - 1) in
ip 2;
region R { let a = sr 5 in ... } // a's Vec: R now, the default region before
The typer already quantified an inner function's allocation regions and named them (it must: v0.1.559 stopped it, which took region polymorphism away, and v0.1.560 withdrew that). What was missing was a parameter to carry them and a call that passes them. Now they lead the lifted definition (and its __direct twin on C), a direct call passes what it bound -- read off the call the way a top-level call's are -- and a use as a value, through the closure adapter, passes the default region, as a top-level function's closure form does. A self-tail call stays a jump only while it hands its regions back unchanged. C and LLVM.
test/regionparams/inner.mere: an inner let rec, a let-bound lambda and a lift inside a lift now read ^param in --dump-region-params (which learned that an inner function's parameters are parameters), and region_params_check asserts the emitted calls on both backends. region_inner_poly still holds.
Not reached: an inner function of one member of a `let rec ... and` group calling another member. The group's region variables are unmarked when the inner binding generalises, because a marked variable it does not quantify might equally belong to an enclosing value, and then a call inside a block would claim it -- keeping the marks made mere-ruby's benv a region BLK value again. Telling the two owners apart is a change to generalisation. It is also the shape mgit has (store_read reaches loose_read through an inner function, in one group), so mgit's reads are unchanged: 33.7 GB of its allocation still goes to the default region, counted by the runtime.
v0.1.561 — 2026-10-01
_A coroutine transfer carries a value, typed by the one that receives it (Q-184)._
coro_new : (Coro['m] -> 'm -> CoroExit) -> Coro['m]
coro_transfer : Coro['a] -> 'a -> Coro['b] -> 'b
coro_exit : Coro['a] -> 'a -> CoroExit
coro_root : unit -> Coro[unit]
coro_switch : Coro[unit] -> unit
A switch carried nothing, so both programs built on coroutines kept a mailbox beside it: mpoll_coro a table from fd to the event bits that woke it (and a second one to hand a pooled worker its fd), mere-ruby a global slot and a kind. Now coro_transfer c v me suspends me, resumes c with v, and returns what me is handed next; a body is handed its own handle and its first message, and ends with coro_exit next v.
It is typed, and that is why `coro_self` is gone. Coro['m] is the type of what that coroutine receives. The 'b a transfer returns is the 'm of me, and the runtime checks that me is the one running -- so what arrives was checked, at the sender, against the receiver's own type: handles come only from coro_new (typed by the body) and coro_root (the thread's own stack, which receives unit). "The coroutine running now" has no type to check a sender against, so coro_self could not be kept; coro_self and the old body shape (fn () -> .. returning the coroutine to run next) are errors that say what to write instead. A let bound to coro_new .. is not generalised, like a channel, or one coroutine could be sent an int here and a bool there. CoroExit is the name of what a body returns rather than Exit, which is a name programs would want for themselves.
A message is a scalar for now -- int, bool, float, unit or a Coro. It is one word in the runtime, read by the receiver with no copy. A boxed message would have to be copied into the receiver's region, and mere-ruby's Fibers run in the default region, so every resume would allocate there. move_check refuses anything else by name, and C and LLVM check again at the type they emit.
How it is built: coro_new is rewritten before type-checking into two internal builtins -- a raw coro_new whose body is handed its handle, and a read of what was sent -- so the runtimes stay untyped and each backend reads the word back at its type with ordinary code; the literal fn me -> fn m -> .. form makes no closure beyond the user's (a wrapper would copy an env per coroutine, the default-region growth v0.1.558 removed). Interpreter, C and LLVM; Wasm and RV refuse each builtin by the name the program wrote.
The two users, moved: mpoll_coro loses both tables -- a connection is a Coro[int] handed the bits that woke it, and a pooled worker is handed its next fd the same way (verify.sh: every mode, all checks). mere-ruby keeps its mailbox, because what a Fiber passes is a Ruby value and a Ruby value is not a scalar; its coro_self became coro_root (the only stack without an entry of its own is the main one), and a finished fiber's entry is dropped last, just before coro_exit. Its corpus (267/267) and core/fiber, thread, enumerator, lazy and threadgroup are unchanged. A million switches cost what they did (0.11 s, C -O2).
coro_check has the values of each kind on three backends, the third argument that is not the one running, and the refusals: a str message, coro_self, the old shape, and a message type fixed once.
v0.1.560 — 2026-10-01
_v0.1.559 is withdrawn: an inner function is region-polymorphic again._
v0.1.559 made a region parameter reach a call made from an inner function by not letting the inner function quantify the allocation's region. That gave every inner function ONE region for all of its calls -- the enclosing function's -- and a helper called once for something that outlives a region block and once inside it is then a region-escape error:
let f = fn (u: unit) ->
let helper = fn (k: int) -> let v = vec_new () in let _ = vec_push v k in v in
let keep = helper 40 in
let t = region B { let tmp = helper 2 in vec_get tmp 0 } in
vec_get keep 0 + t; // v0.1.559: `keep` now holds a value from region `B`
mere-ruby's main.mere has that shape and stopped compiling. The typer is back to quantifying, and the lifted body no longer captures the enclosing function's region parameter either: a closure can outlive the call that made it, and its env would then hold a region pointer whose block has been released and whose struct the runtime reuses for the next block. The emitted C for mere-ruby is byte for byte v0.1.558's again. test/parity/region_inner_poly.mere is the shape above; region_params_check pins the four inner-function calls at ? and goes red when they are taught.
What v0.1.559 was after needs the other half: an inner function taking region parameters of its own, passed at each call the way a top-level function's are.
Two things v0.1.559 said that were wrong. Its sweep of 1,072 programs found nothing newly refused because it ran without their imports resolved, so mere-ruby failed on an import under both compilers and compared equal; and the gate that would have caught it, downstream_check, reports "did not run" (exit 3) when MERE_DOWNSTREAM is unset, which a local run of every gate counted as a pass. And "mgit's peak memory does not move" was one pair of peak-RSS readings on a busy machine, where the same binary on the same input read 5.6 and 15.4 GB. Counted by the runtime (MERE_REGION_STATS=1), v0.1.559 moved mgit's never-freed allocation from 33.7 GB to 10.6 GB -- which is what v0.1.560 gives back up, until inner functions take region parameters.
v0.1.559 — 2026-10-01
_A region parameter reaches a call made from an inner function._
A function that takes a region parameter (v0.1.464) passes it to what it calls -- unless the call is made from something lifted out of its body. Then the callee got the default region, whatever region the outer function had been given:
let pr = fn (n: int) -> let v = vec_new () in let _ = vec_push v n in v;
let sr = fn (n: int) ->
let rec ip = fn (k: int) -> if k == 0 then pr n else ip (k - 1) in
ip 2;
region R { let a = sr 5 in ... } // a's Vec: the default region, not R
It had two causes. The inner let rec ip generalised the allocation's region itself, so its region was a variable nothing would ever bind; and a lifted body was handed no region parameter to pass on. Now an inner binding never quantifies an allocation's region -- only a top-level function can take one, since the backends look it up by source name -- so the variable is the enclosing function's, which makes it a parameter, and a lifted body or a closure receives it as a capture, the way a block's region has been captured since Q-131. Unmarking a region the binding does not quantify is a top-level decision too, so a let rec ... and group keeps its marks.
C and LLVM both pass it: an inner let rec, a closure, a lift inside a lift, and the inner binding of a group (test/regionparams/inner.mere, four ? before, four ^param now, asserted in region_params_check down to the emitted calls; test/parity/region_inner_param.mere holds the four backends to one answer).
Found by mgit: its store_read reached every allocation through an inner function, so none of its reads took a region. They do now (0 call sites named / 1 forwarded before, 7 / 15 now), and the output is unchanged. Its peak memory is not: what it spends is zlib_inflate's buffers, made, used and dropped inside the inflater, and a region that appears in no function's type is not something an argument can reach. That is the rest of Q-134.
No program in the repository or the downstream packages (1,072 files) is refused, or compiled differently in outcome, on C or LLVM.
v0.1.558 — 2026-10-01
_A finished coroutine keeps nothing (Q-183): its env is its own, and its handle is a slot and a generation._
A server that made a coroutine per connection grew by about 120 bytes per connection for as long as it ran -- 13.9 → 23.1 MiB over 80k connections, and mpoll_coro had to write a pool of workers to stay flat. What a finished coroutine kept, measured exactly on C: its 24-byte handle record (32 with malloc's rounding) and its body's env, both for good.
The env is the coroutine's own. coro_new copied it into the default region, which is never freed. An env copier can now be given a marker region that makes the env STRUCT malloc'd -- freed when the body ends -- while what its fields point to goes to the default region as before (a captured string may be kept by something else; the struct's address never leaves it). And the lambda written as coro_new's argument is made that way to begin with, so nothing of it is ever in the default region; its captures are deep-copied into the default region as the copy would have made them.
The handle is a slot and a generation. The program held the record's address, so the record could not be freed without a stale handle switching to some other coroutine. A handle is now 24 bits of slot and 24 of generation, below 2^48 (out of the range coro_scan_ints callers number their own handles in). Reaping frees the record, the stack, the saved state and the slot, and bumps the generation; a handle whose generation does not match answers "has finished". A program keeps its peak of live coroutines.
Measured: a million coroutines made and finished hold 1 MiB (34 MiB before); the Q-183 probe retains 0 bytes per coroutine (64 before); mpoll_coro with a coroutine per connection, its body written as a lambda at the coro_new, is flat over 100k connections (3.4 / 2.9 / 3.3 MiB at 20k / 50k / 100k) with no pool. mere-ruby's corpus 267/267. coro_check.sh caps C at 16 MiB for the million and has a poison that keeps the records; poisons 5, 6 and 8 follow the new runtime. LLVM is unchanged and still keeps the 24-byte record (its cap stays 128 MiB).
Also: migration_check.sh and live_soundness_check.sh both started PostgreSQL on port 54329, so a gate runner running them at once failed one of them; migration_check uses 54330.
v0.1.557 — 2026-10-01
_A closure that leaves its region block no longer reads freed memory: runtime retention (Q-188)._
A use-after-free that every check accepted. A closure made inside region R { } and stored somewhere that outlives R kept pointing into R:
let o = vec_new ();
let _ = region R { let v = vec_new () in let _ = vec_push v 700 in
vec_push o (fn (u: int) -> vec_get v u) };
... blocks that reuse the arena ...
(vec_get o 0) 0 // interpreter: 700. C: -1. LLVM: segfault.
An arrow type names no region, so the escape check cannot see it. On C the closure's env was copied out (v0.1.290) but a captured Vec is a handle into R; on LLVM, which allocates an env in the current region and has no env copier, even a captured int read garbage. Found while planning Q-134, by reading the env copier, and confirmed by running it.
The fix is at run time, not in the types. Every store of a value goes through a copier into the destination's region (__mcopy_). A container's copier -- still an identity, containers are handles -- now asks whether the container's region is a block younger than the destination; if so, the block is marked, and its release hands its memory to the destination instead of freeing it (or, when the destination is not a block open on this stack, to a region that is never freed). LLVM has no copier on that path, so a store of a value whose type holds a closure marks every open block above the destination -- conservative, and never unsafe. Wasm already kept a callee's allocations (Q-132) and reads the right values. A wrong guess costs memory, never safety: nothing depends on the types being precise.
Measured: mere-ruby's corpus (267/267) and csv benchmark mark nothing, so its memory is unchanged by construction. MERE_KEEP_TRACE=1 prints each block that is retained and where it went. test/parity/region_closure_escape.mere holds four backends to the interpreter's answer (v0.1.556's C prints -2); test/escape/ROUTES has the route, SAFE, Q-188.
v0.1.556 — 2026-09-30
_Three small things the dogfoods walked into: Vec[__heap, T], let rec drop, and a region block before ||._
`Vec[__heap, T]` is the default region's Vec, written out (Q-185). The tutorial names __heap as that marker and int Vec expands to it, but the bracket form took only an UPPERCASE first name as a region -- so Vec[__heap, int] made a Vec whose region slot held a type called __heap. mere check and the interpreter accepted it; the C backend stopped at the vec_new that filled it ("missing Vec result type"), which is where mgit's pack table met it. __heap in that position is the region now.
A reserved word after `let rec` is named as one (Q-186). let rec drop = said "expected 'ident = expr' after 'let rec'". v0.1.538 taught the pattern position to say "drop is a reserved word"; the name after let rec is not a pattern, so it never did, and two dogfoods a month apart walked into drop.
A region block is an operand (Q-187). It is closed by its }, like a parenthesised expression, but it was parsed at the level of let and if, so region R { e } || x ended at the } and || x was a syntax error. It is parsed as an atom now. test/parity/heap_region_annotation.mere has both.
v0.1.555 — 2026-09-30
_A listener on a chosen address, bind(2) on a socket, the address a socket has, and the errno fd_ keeps._
tcp_listen binds INADDR_ANY and nothing could ask which address or port a socket had: bind(2), getsockname(2) and getpeername(2) take a struct sockaddr *, which no extern can spell. A ruby on top of this answered TCPServer.new("localhost", 0).addr with the wildcard where ruby answers ::1, could not raise EADDRINUSE or EADDRNOTAVAIL or EACCES from a bind, and read the bound port back from lsof(8). And fd_read / fd_write answer -1 without errno, so a connection the peer reset read as a bad descriptor and every refused socket write had to be called EPIPE.
The runtime keeps the sockaddr, the way it keeps struct rlimit (v0.1.550): an address goes in as the text getaddrinfo(3) reads and comes out as the text getnameinfo(3) writes, and the family comes out as a name, because AF_INET6 is 30 on macOS and 10 on Linux.
tcp_listen_at (host) (port) (backlog) -> int the fd; -1 errno kept; -2 did not resolve. "" is the wildcard sock_bind (fd) (host) (port) -> int numeric host, the socket's own family; 0 or -1 sock_local_addr (fd) -> str "inet 5000 127.0.0.1", sock_peer_addr (fd) -> str "inet6 5000 ::1", "unix 0 <path>"; "" on failure fd_last_errno () -> int errno of the last of these or of any fd_ call, 0 on success
tcp_listen_at does what ruby's TCPServer.new does (ext/socket/ipsocket.c, init_inetsock_internal): each answer of an AI_PASSIVE lookup in order gets socket, SO_REUSEADDR and bind, the first that binds is the listener, and when none does the errno is the last answer's. fd_* return what they returned before; every call now stores the errno of its own syscall in the per-thread slot fd_last_errno reads, and a refusal the runtime decides itself (a negative fd, an unknown mode, a read of size 0) is EBADF or EINVAL.
scripts/sockaddr_check.sh, 73 rows, in CI. Each row is a port the kernel gave the probe, compared with what the other end of a connection reports, or a refusal it provoked: a port in use, an address the host does not have (TEST-NET-1), a bind of a socket already bound, and a reset provoked by closing a socket with an unread byte in it after poll(2) said the byte had arrived. The errno numbers differ by platform exactly here (EADDRINUSE 48 and 98, ECONNRESET 54 and 104, ENOTSOCK 38 and 88), so the expected ones are read from the host's own <errno.h> by a C program the gate compiles. Run on macOS arm64 and on Linux arm64 and x86-64, as root and as an unprivileged user. Poisoned six ways -- errno not per thread, errno dropped by fd_*, the last answer bound instead of the first, the socket's family ignored by sock_bind, no SO_REUSEADDR, the family names swapped -- and each moved at least one row.
⚠ Three of the poisons moved nothing on the probe's first version. The wildcard row accepted :: or 0.0.0.0, which is right -- a host without IPv6 answers only the second -- and so it also accepted a listener bound to the last answer. What does not depend on the host is that a listener on a name is reached by a connect to the same name; that row catches it. The SO_REUSEADDR row needed a port in TIME_WAIT, which an RST does not leave, so it closes a connection from the server side first. And sock_bind binding "" on an IPv4 socket is only a question where :: is answered first, so that row binds exactly that.
Interpreter, LLVM and Wasm are unchanged, as for fd_*: the interpreter says it has no mock for the extern, LLVM and Wasm leave it an import.
v0.1.554 — 2026-09-30
_A gate that did not run says so in its exit status: 2 could not answer, 3 optional._
Ninety-four places in scripts/ printed "skipping" or "SKIP" and exited 0, so "passed" and "did not run" were one status. A runner had to read the words to tell them apart -- mgate did, and read three passes that count their skips ("(13 skipped)") as skips -- and the 65 gates that could not fail for a missing tool were this shape. Every one of them now says what it is:
| status | class | |
|---|---|---|
| 0 | passed | |
| 1 | failed | |
| 2 | could not answer | a tool it needs is missing, too old, or failed to set up (73 places) |
| 3 | optional, not run | the author wrote "this check is optional", the tool is one the gates runner is not meant to have (psql, sdl2-config, wasmtime, wasm-tools), no PostgreSQL at all, an architecture with no pattern, MERE_DOWNSTREAM unset (21 places, 15 gates) |
| 201 | timed out | scripts/bounded.sh |
In CI a 2 is red: tool_preflight --required says the runner has every tool the gates probe, so a gate that cannot find one is a broken runner. A 3 is not: the fifteen gates that can exit 3 run through scripts/gate.sh, which turns 3 into 0 and prints that the gate is optional and did not run. The 26 header comments that said "Skips (exit 0)" say which status they mean.
scripts/skip_exit_check.sh (CI) keeps the old shape from coming back: an echo that says SKIP, skipping or skip-ok with exit 0 on the same line or the next. Its poison plants both shapes.
These are the classes mere test (v0.1.552) and mgate already report.
Two more things the gate runs found in the gates themselves. normalize_conformance.sh wrote its scratch program into examples/ under a fixed name, and a stopped run left it there -- where it was committed once, with this machine's temp path in it. It carries the process id now, is removed on INT and TERM too, and .gitignore names it. And the first poison of skip_exit_check.sh planted a probe for a tool called frob in the usual command-lookup form, which tool_preflight --required read as a dependency of the gates -- in the comment explaining it, too.
v0.1.553 — 2026-09-30
_A table a helper builds at module init is module state under --lib, not a call's._
Under mere -c --lib every top-level function body is typed inside the call region (v0.1.455), so a helper that builds and returns a container has Vec[__call, _] in its type. Called at module init it runs outside any call and the container lives in the default region -- but the value's type still said __call, and the library-boundary escape check refused it:
let vec_of = fn xs -> ...; // builds and returns a Vec
let tbl = vec_of (Cons (3, Cons (4, Nil)));
pub let get = fn (i: int) -> vec_get tbl i;
region escape across the library boundary: tbl now holds a value built during a call ... build it at module init -- which is what the program did. Every table mgz builds that way was refused (with carets on the functions after it, not on the table), so mpng-ruby, which links mgz into a library, had not built since v0.1.455. Found by running every downstream repository's checks at once.
A non-function top-level value's type now says __heap where it said __call: at init the call region is the default one. lib_check.sh has the case, and its store-into-module-state case is still refused.
v0.1.552 — 2026-09-30
_mere test: the checks a package declares, found and run._
Q-169 had half an answer since v0.1.520: mere doc came in, mere test did not, because the question was the discovery convention (*_test.mere? a tests/ directory?) rather than anything technical. The convention is now read off the packages that exist: of the downstream repositories with checks, eighteen have a verify.sh at their root, three have *_test.mere files, and none imports contrib/test. So:
[test] run = ["sh verify.sh", ...]in mere.toml is the list, run in order
from the manifest's directory -- a list and not a glob, as test/contrib_ctests.txt is, because a check that should have been found and was not says nothing
- without a list, the directory's
verify.sh; with neither, exit 2 and a line
saying so
- each exit status is a class, the one the gate runner mgate uses: 0 PASS, 2 CANNOT, 3 SKIP, 201 TIMEOUT, else FAIL.
mere testexits 1
on a FAIL or TIMEOUT, 2 on a CANNOT, 0 otherwise
- each command sees
MEREandMERE_BIN(this compiler) and, for a
checkout's build, MERE_ROOT. ⚠ The first version set MERE to the checkout, on the belief that most verify.sh files take one: run over all 21 downstream repositories it turned 14 red with "is a directory". Counted, 12 run $MERE and 9 build $MERE/_build/...; the 9 accept either now
Before this, mere test compiled a file called test -- the [_; path] arm took it. scripts/test_cmd_check.sh (CI) runs a fixture with one check per class and a verify.sh beside the list; its poison removes the list, and the check that the list ran in order is what goes red.
v0.1.551 — 2026-09-30
_Four holes the dogfoods walked into: a record named like the prelude's result, a builtin passed as a value, a region captured from the wrong side of a lift, and a gate that could not run beside its own poison._
Three new programs written against the language in one day -- a gate runner (mgate), a git reader (mgit) and a coroutine-per-connection server (mpoll_coro) -- and each stopped on something no earlier program had written.
A record and a variant sharing a name are refused. type result = { name: str } passed mere check and ran on the interpreter, and mere -c died inside the compiler: Invalid_argument("List.combine") in subst_variants, the record's zero parameters zipped against the prelude's ('a, 'e) result. v0.1.474 and v0.1.525 refused a variant redeclared with other constructors and a record redeclared with other fields; the cross case was the one left, and the name a program reaches for first is one the prelude already owns. Now: type result is declared both as a variant (Err of _ | Ok of _) and as a record ({ name }), from either order. The capability records Logger and Metrics stay replaceable by a program's own declaration in any shape, as before. Measured before landing: 0 of 1,193 .mere files in this tree and downstream are newly refused.
A builtin is a value on every backend that can apply it. list_iter hs join did not compile on C while list_iter hs (fn h -> join h) did: each backend eta-expanded builtins in value position only for the names a phase had listed. The rest are now expanded too, on C, LLVM and Wasm -- of fourteen single-argument builtins tried, twelve were refused on C and all twelve compile (is_digit, utf8_len, hex_of_bytes, print_int, file_exists, ...); on LLVM and Wasm the ten tried were refused as values while applying them worked. A builtin with no direct-call lowering comes back to the value path from inside its own expansion and is refused there, recognised by the synthesized callee's identity. ⚠ The first version guarded with a flag set around the synthesis; the expansion becomes a closure adapter whose body is emitted when the adapters are drained, the flag was down by then, and test_basic ran for an hour. test/parity/builtin_as_value.mere holds four backends to one output.
A region opened inside a lifted function is not a capture from outside it. fn -> let rec loop -> region R { let b = bytebuf_new 0 in let rec go (pushing to b) } emitted C that did not compile: use of undeclared identifier '__region_R'. The names bound inside a lifted body held the block as R, and a region travels as a capture named __region_R (Q-131), so the inner helper's capture was threaded into loop from its call site, outside the block. Both spellings are bound now. test/parity/region_rec_capture.mere.
`fmt_roundtrip_check.sh` can run beside its poison. It writes each formatted example next to the original under a fixed name, and a gate runner running the plain and --poison invocations together (mgate) found 56 and 48 "refused" files on a main that is green run by run: the two overwrote each other. CI runs one step at a time, which is why nothing had seen it. The probe names carry the process id, and cleanup removes only its own.
Recorded, not fixed: a record field annotated Vec[__heap, int] loses the Vec's element type at the vec_new that fills it (C), and a Vec of region-polymorphic records at top level is the Q-042 shape; mgit keeps its packs in parallel Vecs instead. And Q-134 has its second witness: mgit reads every object inside region R { } and peaks at 1.9 GB on this repository, because the allocations are made by the functions the block calls.
dune test 2870/0; every gate in the workflow plus the two unwired ones run through mgate, against the same run on v0.1.549: fmt_roundtrip green again, first_run green after the README's counts (193 parity programs, 2870 tests).
v0.1.550 — 2026-09-30
_Resource limits, scheduling priority and flock(2): proc_* and file_flock._
Five syscalls a program could not declare with extern fn, each for its own reason: getrlimit / setrlimit move the limits through a struct rlimit *, getpriority / setpriority take an id_t -- <sys/resource.h> is in every emitted program already, so the prototype Mere writes meets the real one and is a "conflicting types" error -- and flock is in <sys/file.h>, which is not included. Like file_* (v0.1.509) and fd_* (v0.1.522) the C is written in the runtime with the real headers, emitted only when a program declares one of the names: two headers, no library, no change to the build line.
The platform's numbers differ exactly here -- RLIMIT_NOFILE is 8 on macOS and 7 on Linux, RLIM_INFINITY 2^63-1 and 2^64-1, EWOULDBLOCK 35 and 11 -- so they stay on the C side and the runtime defines the contract:
- a resource is asked for by name (
proc_getrlimit "NOFILE"), the names
ruby's Process.getrlimit accepts, present only where the platform defines them; proc_rlimit_names lists them and proc_rlimit_resource hands out the platform number for a caller that has to show it (ruby's Process::RLIMIT_*);
- a limit of
-1isRLIM_INFINITYgoing in and coming out;
proc_rlim_const "INFINITY" gives the platform's value in decimal, because on Linux it does not fit an int;
proc_getrlimittakes one snapshot andproc_rlimit_fieldreads soft or hard
out of it -- file_stat's rule, so the two limits come from one moment;
- the priority selector is 0 process / 1 group / 2 user, and the lock bits are
1 shared / 2 exclusive / 4 non-blocking / 8 unlock;
file_flockanswers three ways:0,1when the lock is held elsewhere
and the caller asked not to wait (ruby's false), -1 otherwise.
errno is kept, which the earlier families do not do. getpriority answers -1 for a process at priority -1, so its failure is visible only in errno, and errno is gone by the time a Mere program can ask. Every call stores its own errno (0 on success) in a per-thread slot, and proc_last_errno reads it.
scripts/proclimit_check.sh, 55 rows, in CI: limits set and read back, a nice value raised and read, a lock taken through one open and refused through another, and the refusals (unknown name, soft above hard, bad selector, no such process, closed fd). Run on macOS arm64 and on Linux arm64 and x86-64, as root and unprivileged. Poisoned six ways -- a refusal answered as -1, the non-blocking bit mistranslated, infinity not mapped coming out, infinity not mapped going in, getpriority's errno dropped, a failed snapshot left readable -- and each moved at least one row. ⚠ The probe's first run broke itself: a "soft above hard" row asked for the current hard limit plus one, NOFILE's hard limit is unlimited on macOS and crosses as -1, and the soft limit became 0 descriptors. Every open after that row failed. The row names both of its numbers now.
Interpreter, LLVM and Wasm are unchanged, as for file_*: the interpreter says it has no mock for the extern, LLVM and Wasm leave it an import.
v0.1.549 — 2026-09-29
_What a suspended coroutine still holds: coro_scan_ints._
A program that keeps its own handles -- integers into tables of its own, as an interpreter written in Mere does -- and collects them itself could not see what a coroutine that stopped in the middle of an expression still held in C variables of its stack: something it had just taken out of a table. mere-ruby lost the object in [q.pop, Fiber.yield] to a collection and crashed when the fiber resumed (2 of 4 shapes). coro_scan_ints c lo hi f hands f every integer in [lo, hi) that coroutine c's stack can reach:
- a suspended coroutine from where it stopped, the running one from the call up
(registers spilled first);
- a word that points into a live block is followed and the 16 words from
there read the same way -- without limit inside the coroutine's own regions, at most three hops into any other block. Only the allocated part of a block counts (a recycled region's block still holds what was there before), and a word inside no live block is never read through. That takes a list of every live block, which the region runtime now keeps (one per thread; the shared default region's under a lock of its own);
- conservative, never unsafe: it hands over numbers, not values;
- a suspended coroutine's findings are kept until it runs again.
Why a range and not a bound: mere-ruby's handles were small integers, and so were its loop counters and its Ruby Integers. Every one read as some handle, and what those kept kept more -- a dropped Enumerator's fiber came out "reachable" and the reclaim gave back none of 9000. Handles numbered from 2^48 are above every user-space address and every counter; the program asks for that range. (Reading a table's arrays from the middle still reads the neighbouring handles; mere-ruby therefore decides which fibers can run again without the scan, and uses the scan only for what to keep.)
The interpreter and LLVM hand over every integer in the range -- the superset the contract allows. Wasm and RV refuse it by name, with the other coroutine builtins. scripts/coro_check.sh --poison: fixtures scan and scan_high, and poisons for not following pointers, not reading suspended stacks, not reading the running stack, and not following the coroutine's own regions; run on macOS arm64 and Linux arm64. mere-ruby's csv benchmark against v0.1.548's mere-ruby: parse_line 5.03 -> 5.19 s, generate_line 1.26 -> 1.40 s, the rest unchanged.
v0.1.548 — 2026-09-29
_Asking whether a suspended coroutine pins an arena is cheap enough to ask on every recycle._
v0.1.547's question -- does any suspended stack point into this arena? -- rescanned every suspended coroutine's stack each time it was asked, and mere-ruby asks it constantly: it recycles a map at every method return (the frame pool), and CSV.parse_line leaves up to a thousand suspended fibers behind. Its csv benchmark did not finish in an hour. Three changes:
- A coroutine can only point into a block that existed while it ran. Blocks
carry the epoch they were allocated in (bumped by every switch), each coroutine the epoch it last stopped in, and the live list is kept most-recently-stopped first -- so a walk ends at the first coroutine that stopped before every block in question.
- A recycle keeps its region's first block and frees only the growth after
it, so only those blocks are asked about. They are recent, which is what makes the walk above short.
- Each coroutine's candidate words are gathered once per stop -- only the ones
inside the address range this thread's blocks have spanned -- sorted, and binary-searched per block.
The csv benchmark (2000 rows, best of three) against the build before fibers could suspend: parse 1.47 -> 1.44 s, quoted 26.6 -> 26.3 s, headers 1.62 -> 1.52 s, parse_line 4.88 -> 5.09 s, generate_line 1.20 -> 1.40 s. scripts/coro_check.sh --poison still passes, the compaction poison included.
v0.1.547 — 2026-09-29
_Compacting a map or a Vec while a coroutine is suspended is safe (C)._
A coroutine that stops in the middle of an expression can hold a pointer into a store's arena in a C variable of its own stack -- a string it read out of a map, say. map_compact / vec_compact move the store into a fresh arena and free the old one whole, and map_recycle frees its growth in place, so the coroutine resumed into freed memory: held: lost at -O2, a heap-use-after-free under ASan.
Now the runtime asks before it frees. A program that makes coroutines installs a hook (on the first coro_new) that scans every suspended coroutine's stack -- from the stack pointer its switch saved, which covers the registers the switch pushed, to the stack's top -- for a word that falls inside one of the arena's blocks (low three bits masked, for tagged values). If one does, the arena is retired instead of freed, and tried again at the next retirement; a recycle moves the map to a fresh arena and retires the old. The scan is conservative: a word that only looks like a pointer keeps an arena a little longer, never frees one early. A program without coroutines frees at once, as before.
test/coro/compact.mere holds a string across a compaction and a churn that reuses the freed memory; scripts/coro_check.sh runs it on the interpreter and on C, and its new poison takes the retirement out and goes red. (The LLVM backend has no map_compact / vec_compact at all -- it refuses them by name -- so there is nothing there to retire.) Run on macOS arm64, Linux arm64 and x86-64 Linux.
This is what mere-ruby needs to compact while fibers are suspended -- until now it skipped compaction then, and got back entries but not bytes.
v0.1.546 — 2026-09-29
_A free-result function used at two types, and with its result thrown away, no longer calls an instance nobody emitted (C and Wasm)._
let boom = fn (msg: str) -> fail ("boom: " ++ msg);
let as_str = fn (b: bool) -> if b then "fine" else boom "str";
let as_bool = fn (b: bool) -> if b then true else boom "bool";
let thrown = fn (u: unit) -> let _ = boom "thrown" in 0;
boom's result type is free (it ends in fail), so it is instantiated once per type it is used at: str -> str and str -> bool. The call in thrown throws the result away, its type stays a variable, and the backend names that call's instance with the variable erased to int -- mu_boom__str__int on C, $boom__str__int on Wasm. Nothing emitted it. The compiler said nothing; clang said "call to undeclared function" and wat2wasm "undefined function variable". The monomorphizer already recovered the same hole for a function that never concretized at all (it emits the erased instance for a call site that names one); a function that did concretize, at more than one type, fell outside that. It is covered now, in the same fixpoint: a recovered body may call a function the other recovery has to supply.
The C and Wasm backends name a residual call that way and now get the instance; LLVM refuses such a program by name, as it did. Two further holes showed on Wasm, both found by the self-hosted codegen's bootstrap test: a recovered instance's body calls other polymorphic functions at types the main pass never saw (the fixpoint now collects those too), and a backend lists a name's specializations from the instance table, which the recovered ones were missing from.
Found by mere-ruby's scheduler, whose delivery of a pending exception used re_raise for its value where a dozen other sites discard it. test/parity/free_result_thrown_away was red before and is green after.
v0.1.545 — 2026-09-28
_A finished coroutine keeps 24 bytes, not its saved state._
A Coro is a value the program may go on holding after the coroutine has finished, and switching to it must then fail by name, so the record a handle points at is never freed. It held everything: the saved registers' jump buffer, the parked runtime state, the stack's bounds. A million coroutines made and dropped held 339 MiB on C and 522 MiB on LLVM after they had all finished.
The record is two now. The small one -- state, owner, and a pointer to the rest -- is what the handle points at and what the checks read. The large one is freed when the coroutine is reaped, along with its stack going back to the pool. The same million hold 34 MiB on both backends.
scripts/coro_check.sh bounds that program's resident memory at 128 MiB on C and LLVM, and a poison per backend takes the free back out; both go red. The measurement is /usr/bin/time (BSD -l or GNU -v); where there is neither, the check fails and says the question was not asked, rather than passing. Run on macOS arm64, Linux arm64 and x86-64 Linux (emulated, with C and IR emitted ahead).
mere-ruby's Fibers moved onto the coroutines the same day, and that is where a million of them is an ordinary program: one Enumerator stepped per CSV line.
v0.1.544 — 2026-09-28
_LLVM lowers the coroutines too, without assembly._
The C runtime switches stacks with a few lines of assembly per target. The LLVM backend's output is IR that is not specific to one machine -- its runtime decides the platform at run time, off one icmp -- so it cannot carry a register swap for arm64 and another for x86-64. It does not need to. Between two stacks that already exist, a switch is _setjmp on the way out and _longjmp on the way in; the one step that has to set the stack pointer, entering a coroutine for the first time, is llvm.stackrestore onto the top of its stack followed by a call. Measured the same as C: two million switches in 0.11 s at -O2. Checked at -O0 through -O3 on macOS arm64, Linux arm64 and x86-64 Linux, and with _FORTIFY_SOURCE.
What a switch carries is C's four pieces plus this backend's ListBuf depth word. The guard page is posix_memalign + mprotect rather than mmap, because the MAP_* flags differ between Darwin and glibc and the protections do not. The thread's own record is on the heap, made on first use: a large thread-local lands in each thread's stack on glibc (v0.1.541).
One refusal C does not have: a coroutine made while a region block is current fails by name. This backend's closures carry no env copier, so the body's env would be released with the block; C copies it out.
Found on the way: a region block whose result is a Coro asked the variant tables about it, as if every named type were a user variant, and the backend stopped with unknown variant type.
Both runtimes now keep finished stacks for reuse, up to 16 per thread, linked through a word at the bottom of each: making and finishing 100,000 coroutines on C went from 1.10 s (0.64 s of it in mmap / munmap) to 0.02 s of user time. And the guard below each stack is 64 KiB instead of one page, because clang does not probe a large frame on Linux and a frame bigger than the guard steps over it. That also answered v0.1.543's open note: under x86-64 emulation on an arm64 Mac the overflow at -O2 died unnamed with a one-page guard, and is named with this one. The emulator runs on 16 KiB host pages and a 4 KiB guard is a quarter of one. Real x86-64 (CI) named it either way. There is no poison for the guard's size: a frame between one page and 64 KiB is not something a Mere program can be made to have on purpose.
scripts/coro_check.sh now runs every fixture on LLVM at -O0 and -O2 as well, with seven poisons on the emitted IR beside the seven on the emitted C; each turns its fixture red. test/coro/pingpong became 1000 x 1000 switches: as a single recursion a million deep it measured the stack at -O0 rather than the switch. Run on macOS arm64, Linux arm64 and, with C and IR emitted ahead, x86-64 Linux under emulation: every fixture and every poison as expected.
v0.1.543 — 2026-09-28
_Same-thread coroutines: coro_new, coro_switch, coro_self (interpreter and C)._
Until now the only way to leave a stack standing and come back to it later was spawn: one OS thread per suspended computation, a switch that is a channel round trip between two threads, and a thread's creation cost per suspension. A coroutine is a second stack on the same thread. coro_new : (unit -> Coro) -> Coro makes one without running it; coro_switch : Coro -> unit hands the thread to it; coro_self : unit -> Coro names the one running. The body's result is the coroutine to hand over to when it ends, because transfer is symmetric and there is no parent to return to.
Swapping registers is the easy part (a few lines of assembly per target, arm64 and x86-64). What a switch must also carry is every piece of runtime state that belongs to a stack rather than a thread. On C there are four: the current region, the stack of open blocks a fail unwinds, the innermost try_or's jump buffer, and the stack bounds the SIGSEGV handler compares a fault against. Leave any one on the thread and the program reads freed memory, releases another stack's blocks, jumps onto a stack that is not running, or reports its own overflow as a segfault. The interpreter carries its three (call depth, the call stack a failure prints, the current region's id) and builds symmetric transfer from OCaml 5's asymmetric effects: only a thread's root runs a driver, and a coroutine that switches returns its continuation to it. A fail no try_or on the coroutine's own stack catches is uncaught on both, rather than landing in a handler the switching stack had open.
Coro is neither Send nor Sync. A body whose captures include a container made inside an open region R { } is a type error: the body's env is copied into the default region when the coroutine is made, which is right for a value and wrong for a handle. Switching to a finished coroutine fails with a message, as does a body that ends by naming itself. On C a coroutine's stack is the size stack in mere.toml asks for, or 8 MiB, reserved below a guard page, and its overflow is named. LLVM does not lower these yet; Wasm and RV refuse them by name, with the reason.
Found on the way: fail and exit in a position that wants a Vec, Map, Channel or Coro emitted (Vec___rp12_int){0} after the noreturn call -- a struct nobody declared -- and clang refused the program. The three sites that make that unreachable value (with the match fallthrough, which had the pointer case since v0.1.51) now share one function. test/parity/fail_pointer_result was red before and is green after.
scripts/coro_check.sh (CI, with a poison): eleven fixtures in test/coro/ on the interpreter and on C at -O0 and -O2, the refusals, and one allowed capture. The poison edits the emitted C to drop each of the four carried pieces, the env copy, and the two runtime checks, and the fixture that exists for each must go red; all seven do. Run on macOS arm64 and Linux arm64 (ubuntu 24.04, clang 18). On x86-64 Linux under emulation the fixtures and poisons are green except the overflow at -O2, which dies unnamed; the same emulator also loses the name for a plain main-thread overflow at -O1 about half the time while a hand-written C program with the same handler never does. Real x86-64 is CI's to answer.
v0.1.542 — 2026-09-28
_LLVM: two threads can allocate at once (Q-180)._
The LLVM backend's allocator took no lock on the default region, which every thread shares (the C backend's does), and its current region and stack of open blocks were process globals (C's are thread-local). Two spawned threads building strings at the same time -- sharing nothing -- died with out of memory in 3 runs of 3; doing it inside region R { } blocks, they hung or crashed. One thread alone was fine, which is why nothing saw it: no gate ran two allocating threads concurrently on LLVM.
The default region now takes a one-word spin lock (cmpxchg: a zeroed pthread_mutex_t is not an initialized one on macOS), taken only for the default region; a block region belongs to the thread that opened it. The current region, the open-block stack and the ListBuf depth word are thread-local -- small words only, since a large thread-local lands inside each thread's stack on glibc (v0.1.541). The allocation meter adds atomically.
scripts/threads_alloc_check.sh (CI, with a poison): two threads allocating at once, in the default region and inside block regions, three runs each on C and LLVM under a time bound. The poison takes the fix back out of the emitted IR -- the lock calls, then thread_local on the current region -- and each must turn a fixture red. Run on Linux (ubuntu 24.04, clang 18) and macOS.
v0.1.541 — 2026-09-28
_LLVM: a spawned thread's alternate signal stack is its own block, not a thread-local one._
v0.1.540 made LLVM's 128 KiB alternate signal stack thread_local so that every thread could name its own overflow. On glibc, static TLS lives INSIDE each thread's stack, so a thread created with a small requested stack had no room for it and pthread_create refused: stack = "128K" made every LLVM program on Linux fail with "this host refused the stack this program asked for". macOS allocates TLS apart from the stack, which is why it passed there, and spawn_stack_check.sh's poison caught it on CI's Linux. The buffer is a plain global again, used by the main thread; a spawned thread mallocs a block of its own, installs it, and disables and frees it when it ends (the C backend has done it that way since v0.1.540). Checked on Linux (ubuntu 24.04, clang 18.1.3) and macOS: spawn_stack_check and stack_request_check with their poisons, thread_fail_check, thread_leak_check.
v0.1.540 — 2026-09-28
_A spawned thread gets the stack the program asked for, and its overflow and its fail stay its own._
`spawn` honours `stack` (Q-178). Q-168 let a program say stack = "512MB" in mere.toml, and the C and LLVM backends run main's work on a thread of that size. spawn called pthread_create with no attributes, so every thread a program started had the host's default -- 512 KiB on macOS, ulimit -s on Linux -- whatever it had asked for: the recursion main finishes overflowed one spawn away. With a request, both backends now create every spawned thread with the same size; without one the emitted code is unchanged. The request is still address space and not memory: 64 threads under a 512 MB request hold about 2.6 MiB resident.
A spawned thread's overflow is named. The bounds the SIGSEGV handler compares against were process globals holding the main thread's stack, and sigaltstack -- which is per thread -- was installed on the main thread only. A spawned thread that ran out of stack therefore died with the handler unable to run: no output, exit 132 or 139, the shape v0.1.271 removed everywhere else. The bounds are thread-local now, and a spawned thread sets its own bounds and alternate stack on entry (C: a 64 KiB block allocated for the thread and released when it ends; LLVM: the existing buffer, now thread-local).
LLVM: a fail on a spawned thread no longer lands in another thread's `try_or`. The C backend has had one jmpbuf per thread since v0.1.310; LLVM still had one global, so a fail on a spawned thread with no try_or of its own jumped into the one main was sitting in -- onto another thread's stack -- and main carried on with the handler's value as though it had failed itself. It is thread-local now, and the fail takes the thread's uncaught path (the message, exit 1), as on C. Found while writing the gate below.
A thread that cannot be started is a fail. pthread_create's return code was ignored on both backends, so a refused thread left its handle unset, the closure never ran, and the program went on to join it. It is a fail now ("spawn: the host refused to start a thread"), catchable with try_or like any other; so is a stack size the host refuses.
scripts/spawn_stack_check.sh (CI, with a poison): deep recursion inside spawn finishes with the request and is named as a stack overflow without it, on both native backends; a spawned thread's fail takes its own uncaught path; 64 threads under a 512 MB request stay small. The poison asks for 128K and requires the spawned recursion to stop -- a request that was not reaching spawn would pass everything else on a host whose default is large.
v0.1.539 — 2026-09-28
_The second lexer catches up with the first, and the file's last line keeps its comment._
The self-hosted lexer read ten operators wrong (Q-153). contrib/parser/lexer.mere split |> -- the language's own pipe -- into Pipe and Gt, split <|, <<, >>, <> and <- the same way, and failed outright on @@, ?!, ? and &. selfhost_check.sh stayed green throughout, because the bootstrap corpus uses none of them. All ten are tokens now. The self-hosted parser does not read most of them yet, which is why adding them breaks nothing; what it needed first was the text split where the compiler splits it.
scripts/selfhost_lexer_ops_check.sh (CI, with a poison) does not list the operators: it derives them from lib/lexer.ml's arms -- every 'c' when ... s.[i + 1] = 'd' is a two-character operator, every 'c' -> advance 1 a one-character one, 40 in all -- and requires each to be one token in the self-hosted lexer. Against the old lexer it names the ten. // is left out on purpose (it begins a comment in both), and the backslash arm, the one written as two characters in the OCaml source, is asked for by name after the first version of the gate missed it. Reserved words are a separate gap and are not asked: the self-hosted parser reads import as an identifier, so the lexer cannot make it a keyword alone.
`mere fmt` keeps the comment on the line a file ends with (Q-154, slice 6). sum xs // 15, the shape an example file ends in, lost its comment whenever the expression was the tail of a run of let ... in: the formatter only puts a trailing comment back where it emits the whole line, and never asked at the end of a let run -- correctly, in general, because a caller may append the ; or ) that closes the run, and a comment placed before it swallows it. The one tail nothing is ever written after is the expression the whole file ends in, so that one is recognised by identity and asked. Comments lost on the examples corpus: 99 -> 93. What is left, by the line it was on: an if / then inside an expression 30, a line ending in the caller's ; 26, a variant or match arm 12, let ... in 7, , 6, else 5, other 4 -- plus 7 that are // inside a string literal and not comments.
v0.1.538 — 2026-09-27
_Four known failures, each of which had a written record and none of which had been fixed._
`f_min` / `f_max` with a NaN or a signed zero (Q-177). The interpreter answers with OCaml's Float.min / Float.max and Wasm with f64.min / f64.max; C and LLVM compiled a < b ? a : b, which is not symmetric in a NaN -- false with the NaN first, so f_min nan 1.0 was 1.0 while f_min 1.0 nan was NaN -- and returned equal zeros in argument order, where -0 is below +0. The RV32/RV64 prelude had copied the C spelling on purpose. All three now transcribe OCaml's stdlib step for step:
c = y > x || (!signbit y && signbit x)
min = c ? (isnan y ? y : x) : (isnan x ? x : y)
max = c ? (isnan x ? x : y) : (isnan y ? y : x)
The first version used a + b for the NaN case, and rv_float_check.sh caught it: a signalling NaN comes back from OCaml as it went in, and an add quiets it -- same NaN, different bits, and that gate compares bits. The transcription returns the operand itself. test/parity/float_minmax.mere covers every pair of eight values on the parity backends, and test/float/rv_float_ops.mere now asks f_min / f_max of every pair on RV32I, bit for bit against the machine (3440 lines).
`run` reported a shell killed by a signal in OCaml's numbering (Q-088). v0.1.360 wrapped the command in a subshell so the shell run waits on usually survives to report 128 + signal itself. When the shell itself is the one killed -- kill -9 $$, since $$ is that shell even inside the subshell -- the interpreter read the status from waitpid, which carries OCaml's encoding (Sys.sigkill = -7), and answered 121 where the C backend answered 137; 117 for SIGTERM, 116 for SIGUSR1. OCaml's named signals are now translated to the system's numbers: 1-15 other than SIGBUS are the same on Linux and macOS, and the rest are chosen by the system the compiler was built for, which dune writes into Platform_config from %{ocaml-config:system}. Three rows in run_status_check.sh kill the waited-on shell; with the translation removed they fail as 121 / 117 / 116.
`mere -t` accepted programs the build refuses (Q-083). The type query runs no borrow, move, Send or exhaustiveness check, so a non-exhaustive match typed there and was refused by mere -c -- and a measurement once made with -t concluded that a check did not exist. -t and -te still exit 0 with the type, which is what editors and scripts asking a type question need, and now print the build's own diagnostic on stderr under "note: mere -t answers a type; it does not accept the program". The findings come from Pipeline.type_notes, which runs the build's checks and puts back the warnings they collect, so -t prints nothing new for a program the build accepts.
`view` was reserved and nothing said so (Q-084). let view = 5 failed with "expected pattern", and view was in neither docs/reserved-names.md nor language-reference.md's Keywords block -- which was also missing trait, impl, dyn and derive. The lexer's reserved words are a table now (Lexer.keywords); the parser names one when it is used as a name ("view is a reserved word, so it cannot be a name here"), and scripts/keywords_doc_check.sh (CI, with a poison) requires every word in the table to be listed in both documents.
And a warning that was false on every backend. Writing that section turned up the other half of reserved-names.md: "a top-level name that collides with a C keyword or libc/libm symbol ... will be a compile error at codegen". It had stopped being true -- every backend prefixes a top-level binding (mu_div on C) -- and a program binding div, malloc, printf, memcpy, strlen, case and main printed the same answer on the interpreter, C, LLVM, Wasm and RV32. The warning is gone. TYPE names are not prefixed on the C backend and do still collide (type wait against union wait), so that warning stays, and reserved-names.md and patterns.md say which is which.
dune test 2868/0, parity 188.
v0.1.537 — 2026-09-27
_utf8_chars split a stray continuation byte onto the character before it, on all three compiled backends._ test/parity/prop_utf8.mere had pinned the disagreement since contrib/prop found it: for "A" ++ chr 128 the interpreter gives two characters and C, LLVM and Wasm gave one, so codepoint_at s 0 was 65 on one and a refusal on the others. The file's own note said the difference had to be under the prelude, because the span logic is one Mere source compiled by all four.
It was the builtin. Each compiled backend implemented utf8_chars as a BACKWARD walk -- start at the end, step back over continuation bytes, take what is left as one character -- which gives every continuation byte to whatever precedes it, a lead byte or not. The interpreter splits FORWARD by the lead byte's span, and so did the compiled backends' own utf8_len, which is why the two builtins disagreed with each other inside one backend. No backward walk can reproduce the forward rule: a lead byte does not check what follows it (0xE0 'A' 'B' is one three-byte span), so the boundaries are only known from the front. All three now record the starts going forward (C: malloc; LLVM: malloc, because a string can be megabytes; Wasm: the bump region) and build the list from the back, as before.
The three .expected pins are deleted; prop_utf8 is MATCH on all four.
v0.1.536 — 2026-09-27
_A cached test prints nothing._ With dune on the PATH (v0.1.535), first_run_check got one step further on CI and failed on the next thing: dune test printed no "N passed" line . Its own comment explained why, as a feature -- "dune caches it, so on an unchanged tree this costs nothing" -- and a cached test action is not re-run, so it prints nothing, so there is no count to read. CI runs dune runtest a few steps earlier, so on CI the tree was always unchanged and the check could never pass there; locally it passed whenever something had just been rebuilt, which during a change is always. Reproduced by running it twice in a row, which fails on the second. It asks dune test --force test now: the suite the count comes from, re-run, about four minutes.
v0.1.535 — 2026-09-27
_CI had been red since v0.1.531, on two steps that were right about something and wrong about where._ The Gates job failed on tool_preflight and first_run_check for every push from v0.1.531 to v0.1.534, while every gate passed locally -- including those two, which is how four versions went out on a red build. ⚠ The local sweep in v0.1.533 ("run every gate") was run on a laptop, and the laptop is not the runner.
`first_run_check`: the one real finding. v0.1.533 derives the README's test count by running dune test, and the CI step ran the script without opam exec --, where dune is not on the PATH: FAIL first_run[counts]: no dune. The step runs under opam exec -- now. tool_preflight had said so -- dune — scripts/first_run_check.sh would skip -- in the same log, under four false alarms.
`tool_preflight`: four false alarms, three causes.
initdbandpg_ctl: the PostgreSQL install step sat between two gates,
well after the preflight, so the tools were not on the PATH yet when the preflight looked. The gates that use them ran and passed. The install is with the other tools now, before the preflight.
riscv64-elf-objdump: the preflight resolvescommand -v "$OBJDUMP"from
the script's default, the Homebrew name, while CI runs the gate as OBJDUMP=riscv64-linux-gnu-objdump sh scripts/rvd_oracle_check.sh. On CI (GITHUB_ACTIONS set) the name CI passes the gate is the one asked about; off CI the default still is, because that is how a developer runs it. A sixth poison replaces CI's name with one that does not exist and requires the preflight to report it.
brew:tls_server_check.shasks Homebrew where OpenSSL's headers are on
macOS; on Linux they are on the default path and the gate runs 19 checks without it. It is in ALLOWED_ABSENT, with that reason.
The preflight step runs under opam exec -- too, so that dune -- a tool a gate CI runs needs -- is asked about in the environment that gate gets.
`fma_check`, new in v0.1.534, was red on its first CI run. Its LLVM leg compiled the .ll with $CC, which on the Ubuntu runner is gcc, and gcc took the IR for a linker script. The LLVM leg uses clang now (LLVM_CC); the C leg keeps $CC, so the runner checks gcc's answer there as a second compiler. ⚠ A new gate is supposed to run once on a CI-like Linux image before it is pushed; this one was not, and its first Linux run was the CI run. It has been since: x86-64 Ubuntu (gcc 13 at -O0 and -O2, clang 18 for the IR, glibc's fma as the oracle) and arm64 Ubuntu under dash, all six families matching. A leg that prints no checksums at all is now reported with what it did print -- under node 18, which has no tail calls, the report had been "differs in:" and an empty list.
v0.1.534 — 2026-09-27
_fma: one rounding, asked for by name — and two LLVM bugs its gate found on the way._
`fma a b c` and `f64x2_fma` compute a * b + c rounded once (Q-176). a * b + c written out is unchanged: it rounds twice, on every backend at every optimization level, which is what the C backend's #pragma STDC FP_CONTRACT OFF has guaranteed since v0.1.315. Fusing is now something a program asks for, and what it gets is a different answer that is the same everywhere: fma is one of the operations IEEE-754 requires to be correctly rounded, like + - * / and sqrt.
Why it is worth a builtin, measured on mat4xvec4 through the C backend: the f64x2 kernel written with f64x2_fma runs in 100 ms against 140 ms for f64x2_mul + f64x2_add, and clang turns each pair of lane fmas into one fmla.2d. It is not the benchmark row, because that row's claim is that the scalar and lane programs print the same bytes.
Lowering: the interpreter uses Float.fma, C calls fma(3) (per lane for f64x2_fma), LLVM calls llvm.fma.f64 / llvm.fma.v2f64. Wasm has no fma instruction and the host's `Math` has none either, so $__lang_fma computes it in the module, in integers: the three significands normalized to 54 bits with bit 0 clear, the 108-bit product built from 32-bit halves (Wasm has no 64x64->128 multiply), c aligned against it with a sticky bit for whatever is shifted out, one add or subtract, and ONE rounding -- the conversion of the top 63 bits to f64, pre-rounded at the subnormal's last place when the result will be subnormal so the final scaling does not round again. The structure is musl's. About 12 ns a call on node 24, four times the unfused arithmetic. The alternative was to refuse fma on Wasm, and a program would then write a * b + c for that one backend and get a different answer there -- the thing the pragma exists to prevent. The RISC-V backends refuse the name: they have no double-precision unit, and the 90-bit working width of the software floats there has no room for a 106-bit product.
`scripts/fma_check.sh` (CI) holds all four backends to the C library's fma(3), bit for bit, on six families of inputs: all 10,648 triples of 22 special values, fully random bit patterns, products at the edges of the exponent range, exact and near cancellation, and ties decided by a single sticky bit. The last family exists because random inputs essentially never land on a tie: while the software fma was being written, dropping the sticky bit of a shifted-out c passed 20 million random cases and failed 10,272 of 20 million once ties were in. The gate's own poison is fma redefined as a * b + c, which must not match (it differs in 6 of 6 families). Before it was transliterated, the algorithm ran 250 million cases against the hardware fma natively and the WAT ran 50 million under node, all bit-identical; two of its sticky paths (the product shifted right, or out entirely, against a much larger c) could not be made to change an answer and are argued unobservable rather than tested -- a 53-bit by 53-bit product cannot have 64 zero bits between two set ones.
LLVM: `x != x` said a NaN was not a NaN. Float != was the ordered fcmp one, which is false when either side is a NaN; the interpreter, C and Wasm give IEEE-754's answer, true. It is une now. Found because the fma gate folds NaN results with exactly that test, and the LLVM leg alone disagreed -- on the one family that produces NaNs. test/parity/float_edges.mere asked ==, < and > of a NaN and never !=; it asks all six now, and fails with the old predicate.
LLVM: a local binding of a builtin's name leaked into `main`. main was emitted with the inner-lift view of whichever fn had been emitted last, so a let atan2 = ... inside a fn made every atan2 in main call that local: atan2 1.0 1.0 printed 2.0. main's own local fns are not lifted through that table at all, so its view is the empty one now. Found by the fma parity case, which binds a local fma and uses the builtin elsewhere in the file. ⚠ Only the LAST host's lifts leaked, so the first version of test/parity/local_shadows_builtin.mere passed with the bug in place -- a later fn with its own lifted local had replaced the table. The fn that shadows is the last one in the file now, and the file fails without the fix.
Also found and left for its own version: f_min / f_max with a NaN argument return NaN on the interpreter and Wasm and the other operand on C and LLVM, depending on argument order.
dune test 2862/0, parity 187 (fma, local_shadows_builtin).
v0.1.533 — 2026-09-24
_Running every gate, rather than the ones that looked affected._ A full local sweep of the 124 gates CI names found three red, and two of them had been red on main.
`section_coverage` — mine, from v0.1.528. wasm_stdin_host_used is a gated runtime section, and a gated section with no row in test/parity/SECTIONS is text that is emitted by nothing and therefore validated by nothing. I added the flag and not the row. test/parity/read_stdin_host.mere is that row. ⚠ Its calls are emitted and not executed, on purpose: the harness gives a program no stdin and does not close it, so a real read_line () blocks until the timeout — the first version of the fixture did exactly that. The guard is args (), empty here and unknowable at compile time. section_coverage checks that a section is EMITTED; wasm_stub_check.sh is the gate that feeds it real bytes.
`trailing_trim_check` — red since v0.1.510. fail_reason_check.sh uses sed '$d', which the gate forbids outside an allowlist. The use is legitimate and the same shape as parity.sh's: on the plain Wasm host the diagnostic goes to stdout when stderr is empty, so the last line IS the message and both halves are then compared. It is on the allowlist now, with that reason.
`first_run_check` — the README's own sentence came true. It said the parity count is re-derived "because a number in a README rots silently — and the test count is not ... and nothing checks it". The parity count had drifted 180 → 185 and the test count 2796 → 2856. Both are derived now: the test count by running dune test and reading its own "N passed" line, which is what the number is about. A sentence that names its own failure mode is not a check.
⚠ Nothing found today was found by the build. Two of these were red on main while four versions were pushed, because each push ran the gates that looked related. The sweep is the thing that had not happened.
dune test 2856/0, parity 185+30.
v0.1.532 — 2026-09-24
_A gate in no workflow has never had to pass._ v0.1.531 asked whether a gate CI runs could skip. This is the other half: three gates were in no workflow at all, and one of them was red.
rv_exec_check.sh is the differential between the RISC-V backend and the C backend on the hosted side — the half qemu_virt.sh does not cover, and one of the three boundary checks a conference talk cites. Nothing ran it. Run on 2026-09-24 it exited 1: int_width_boundary, added to the parity suite for a different question (Q-039 pins where the interpreter's 63-bit int and the compiled backends' 64-bit one part company), differs by construction on a 32-bit target — bit_shl 1 60 is 1152921504606846976 on the C backend and 0 there. It is in KNOWN_DIFF now with that reason, which is what the list is for: a name in it that stops differing also fails.
64 bits: 105 passed, 0 failed, 2 known-different. 32 bits: 98 passed, 0 failed, 4 known-different. ⚠ 84 programs are refused by `-rv` at compile time and do not run, so "98 agree" has a denominator of 186, not 98 — host_matrix.md holds the reasons.
tool_preflight_check.sh grew the second question: every `scripts/*_check.sh` must be named in a workflow, or listed with the reason it is not. 90 gates, 0 unwired. The two exceptions are named: rv_float_check.sh needs an emulator checkout and minutes of emulator time, and bigstr_check.sh allocates gigabytes by design and no runner can host it. Two more poisons (a gate falling out of every workflow; the gate glob collapsing), five in total.
⚠ Found by re-measuring numbers for a talk, not by the build. The build had nothing to say, because the check was not in it.
dune test 2856/0.
v0.1.531 — 2026-09-24
_A gate that did not run is not a gate that passed._ 65 of this repo's gates exit 0 when a tool they need is absent, and 46 of them are named in `ci.yml`. Both halves of that are reasonable on their own — nobody should need psql to work on the parser, and a build should not go red because a laptop lacks wat2wasm — and together they mean this: the day a package is renamed, the best-effort install fails, the gate skips, and CI is green with the check gone.
qemu_virt.sh is the case that matters. It is the differential against an emulator nobody here wrote, which makes it the strongest oracle in the repo, and it was installed with continue-on-error: true and skipped cleanly when absent.
scripts/tool_preflight_check.sh fails in CI, before any gate runs, naming the gate that would have disappeared. The tool list is derived, not typed: every command -v in scripts/*.sh, mapped back to the gate that probes for it, and a tool counts as required when a gate that probes for it is named in ci.yml. 33 tools across 93 gates; 31 required.
⚠ The first version missed the gate it was written for. qemu_virt.sh probes command -v "$QEMU", and a scan that reads only literal names sees a variable and moves on — so qemu-system-riscv32, the whole argument, was not in the list. $CC (39 sites) and the loop variables went the same way. There are four shapes, and the derivation resolves all four: a literal, a variable with a VAR=${VAR:-tool} default, a for t in clang wat2wasm node loop, and a have() { command -v "$1"; } helper.
⚠ And it printed "0 of them are needed by a gate CI runs" and called itself ok. The derivation had silently produced nothing — which is the failure this file exists to describe, written into the file about it. There is a floor now, and its poison is one of three: the tool this was built for going missing, the gate-to-CI link breaking, and the floor being a real number. Each asserts the message it should print rather than a non-zero exit.
⚠ A shape-based scan gets false positives for free. The prose "have to" in a comment parsed as a tool named to. Comments are stripped first, and the one genuine false positive left — command -v mere, which three gates use to find the compiler and which they fail loudly without — is named in ALLOWED_ABSENT with its reason rather than quietly dropped.
dune test 2856/0.
v0.1.530 — 2026-09-24
_Rename handed back a program that does not compile._ Top_let's binder is a pattern and has carried a position since the beginning. Top_let_rec's was a bare string, so everything that had to point AT the name pointed at the VALUE instead — a different line as soon as the fn is written under the =, which is how most of mere-ruby's ~2,500 top-level functions are written.
Three things were wrong, and the third is the one that touches files:
let | let rec (before) | |
|---|---|---|
| go to definition | the name | the fn on the next line |
| document symbol | the name | the value |
| rename | declaration + uses | uses only |
Top_let_rec was not in the match that says "a top-level declaration's own name is an occurrence too", so a rename rewrote every use and left the definition alone. The local let rec ... in had the same hole.
Top_let_rec and Let_rec now carry (name, its own position, value). That is 76 sites across 15 files, and the point of doing it as a widening rather than a side table is that fst, snd and List.assoc stop type-checking the moment a pair becomes a triple: the compiler named every place that had to decide, including four in the backends and two in the REPL. A pass that SYNTHESISES a binding — while desugaring, the RV fast-path clone, monomorph's copies — has no name in the source and passes the value's loc, which is exactly what every binder got before.
⚠ The automated part of the edit widened three tuples that were not rec binders (a (name, ty) list, a (expr, params) list, and eval's placeholder refs). Each one type-checked as a different error a step later and had to be put back. Mechanical is not the same as safe; the type checker caught all three, which is the argument for the widening.
scripts/lsp_binding_position_check.sh drives the real server over JSON-RPC and asks the same two questions of five binding forms. ⚠ Every fixture puts the value on the line AFTER the binder, because a one-line fixture cannot tell the two positions apart — and the gate asserts that about itself before it trusts any of its own rows. On the previous compiler it reports six failures naming fn as the definition; on this one, none. Five unit tests cover Query directly and three of them go red without the fix.
A stale exemption went with it. decls_roundtrip.sh carried one program that could not round-trip, with the rule "if it starts passing, FAIL, because the exemption has outlived its reason". Q-125 fixed the underlying gap in v0.1.525 and the rule fired the first time anyone ran the script afterwards. 172 programs, no exemptions.
dune test 2856/0, parity 184+30.
v0.1.529 — 2026-09-24
_Only the editor was red._ mere -t resolved import "dep.mere" against the current working directory. Every other path that reads a file — -c, -ll, -w, check, fmt, fix, --decls, --header, --suggest-regions, --dump-region-params — resolves it against the file's directory, which is where the file's author wrote it. So a program the build accepts was rejected by the type query, with cannot resolve path (tried: /dep.mere) naming a path nobody wrote.
The tool most likely to be on this path is an editor, and an editor is the thing least likely to share a working directory with the file it is showing.
One line: -t passes ~base_dir:(Filename.dirname path) like its twelve siblings. The other four handlers that looked like they were missing it were measured and were not — they compute it a few lines further down.
scripts/type_query_imports_check.sh asks five paths the same question from a directory that is not the file's, and is poisoned twice: the fixture has to actually depend on its import (otherwise a compiler that never reads import passes every row), and a genuinely missing import has to be refused with a message that says where it looked. It reports -t as the only failure on the compiler from ten minutes ago.
This is NOT the other half of Q-083 — -t skipping the borrow and capture checks is deliberate, documented in mere --help, and unchanged.
dune test 2851/0, parity 184+30.
v0.1.528 — 2026-09-24
_The two builtins the gate could not see were the two that were broken._ read_stdin and read_line answered the empty string on plain Wasm, under a comment saying a browser host has no stdin. True of a browser. False of scripts/run_wasm.js, which has an fd 0 — and false of docs/host-matrix.md, which called both yes. This is the rest of the job v0.1.350 did for run, env_var and file_exists, and it was left undone for the same reason it was invisible.
An empty string is what EOF looks like, so the program did not fail. It read nothing and carried on.
Both now go through the host. scripts/mere_host.js grows the one stdin reader every Node host shares — it has to be one, because the buffer is: two independent read_lines would each hold half of a line. Checked against the interpreter and the C backend on four inputs, multi-byte included: three lines consumed in order, a short read at EOF, empty input, and ま/る split across a chunk boundary by construction. A host with no stdin answers 0, which becomes a real empty str rather than a null pointer, because a caller reading the length header of 0 reads whatever is at address -4.
The exclusion list is where the bug lived. scripts/wasm_stub_check.sh said so in a comment:
# read_line / read_stdin / file_openrw are left out: their answers # depend on stdin or on writing a file, neither of which is fixed # across a run.
Stdin is fixed by feeding it and a file is fixed by choosing its path. All three are probed now; the gate went red naming both stubs with their evidence (C=mere_stub_probe_line Wasm=) before the fix and reports 0 after.
A fourth probe was measuring nothing. read_file_bytes was probed with bytes_len, and read_file_bytes answers Vec[R, int], so the expression was a type error and the gate had recorded refused for it since the day it was written — a word that reads like an answer.
The gate now states what it is a fraction of. "10 builtins probed, 0 stubs" says nothing about the eleventh. The Wasm backend's host surface is the set of (import "env" ...) names it can emit, and it is 39. Fourteen probes reach 32 of them; the other 7 are named with the reason nothing probes them — memory is not a function, exit_proc's answer is an exit status that exit_status_check.sh owns, and the five concurrency imports have no answer that is fixed across a run. The list is checked in both directions, so a skip that stops being true fails instead of drifting into fiction. Poisoned four ways, each asserted by the message it should print rather than by a non-zero exit.
Found while adding the flag: three flags from v0.1.350 were never reset. wasm_run_host_used, wasm_env_host_used and wasm_fexists_host_used are per-emission state, and a second module emitted in the same process inherited the first one's host imports — one more name its host has to bind or fail to instantiate. Every JS host in the tree happens to bind all three, which is why nothing noticed. Four unit tests emit a program that reaches for the host, then one that does not, and read the second.
dune test 2851/0, parity 184+30, wasm_stub_check 14 probed / 0 stubs / 39-name surface, --poison 4 caught.
v0.1.527 — 2026-09-24
_Q-120 (b): mere --ffi-header writes the C side's header, so a shim stops copying the layout by hand._
extern fn carries float, record-by-value and bytes on the C and LLVM backends. What it did not carry was any way for the author of the other side to KNOW what those look like — the mu_ field prefix, the field order, and the fact that a boundary int is C's 32-bit int and not long long are all this generator's choices. Shims wrote them out by hand:
typedef struct { double mu_x, mu_y, mu_z; } v3; /* copied */
⚠ That is an ABI held together by two people agreeing. An extern declaration is a promise, not a check.
mere --ffi-header <file> prints the prototypes and the structs they carry, and only the types the boundary reaches — an internal record does not leak into a file the other side compiles against. Nothing new is computed: the prototype builder was already in the C backend and is now one function both callers share, because writing that mapping twice would be two answers the first time either moved.
⚠ `--header` was already taken, by the other direction — Mere compiled as a shared library, for a C caller. The compiler said so (this match case is unused) rather than the two silently becoming one. Two headers, two guards: MERE_FFI_H stays with the library one, this is MERE_EXTERN_H.
scripts/ffi_header_check.sh builds a shim that includes the header, links it against the emitted C, and checks the answer; regenerating gives the same bytes; reordering the record in the Mere source keeps the shim right. Its poison is the hand-copied struct, now stale, giving 607 where the header gives 418.
⚠ The first version of that poison passed for the wrong reason. The shim computed a dot product, which is the same number however the fields are permuted, so a stale copy still looked right. The weights make the order readable.
Still untested and therefore still unclaimed at the boundary: tuples, variants, and passing a Mere closure as a callback. The header does not name them, because naming them would be making the same unchecked promise this replaces.
2,847 unit tests, parity 214/214.
v0.1.526 — 2026-09-24
_Q-161: the syntax hint's window is the statement now, not the failing line._
v0.1.507 gave syntax errors a help: line that answers a spelling from another language. All 18 catalogue rows answered — but only because the evidence sat on the line the parser stopped at:
def f(n): def f(n):
return n 0
The first gets help: \def\ — a function is \let name = fn (x: int) -> body;\. The second was silent: ): reads as a type annotation, so the parser walks on and fails at line 2, and def is a line back.
The window walks back to the nearest ; now, capped at two lines. Both numbers are measured: at a cap of 1 the window is the old line-only one and the new row goes silent; at 2 it answers, all 18 old rows still answer, and every poison still goes red. Wider buys nothing that was asked for.
⚠ The fear this question recorded did not materialise. Q-161 said widening the window widens the false-positive window for var / case / val / mut — words that are real identifiers in these repositories. It does not: bound_names walks the WHOLE token list, so "does this file bind that name?" never depended on the window. That was checked at each width rather than assumed, because the question asked for it to be.
Two poisons were added, and the first one written was weaker than it looked:
- The
;is a wall — a spelling three statements back is not offered as the
explanation. ⚠ But that fixture passes at ANY width, because the ; stops it before the line cap is reached, so it says nothing about the cap.
- So the cap gets its own: run the catalogue with the window one line narrower
and the cross-line row must go red. Without it the number in the source would be decoration.
19 catalogue rows, 5 poisons. 2,847 unit tests, parity 214/214.
v0.1.525 — 2026-09-24
_Q-125: a record declared inside a module could not be named in an annotation — and not only from outside. From inside, either._
module M { type t = { a: int }; let get = fn (v: t) -> v.a; }
⚠ That did not compile, in the module that declared the type. The question was recorded as "cannot be named from OUTSIDE the module", and it was also narrower than recorded: variants were fine all along. M.v worked, v worked, inside and out.
The asymmetry is where the qualification lands. A variant's lands on the CONSTRUCTOR — M.A still builds a v — so the type stays canonical and any annotation matches. A record's landed on the TYPE NAME: the literal became M.t { … } and every annotation resolved to t, so the two could never meet.
A record's identity is its bare name now, the way a variant's already was. The canonicalisation is in the parser, at the two places it builds a qualified record literal or pattern — ⚠ one place rather than the 87 sites that read a record name across the typer, the exhaustiveness checker and both native backends. Putting it in the typer first worked for the typer and left the C backend emitting two structs, M__t and t, for one type.
⚠ Records had no redeclaration check at all, which variants have had since v0.1.474. type t = { a: int }; type t = { b: str }; was accepted in silence and the second won. That was survivable only while M.t and t were different types; making them one would have merged two different records without a word. So the guard lands first, in the same shape as the variant one and with the same kind of message — and the variant wording is not touched, because three unit tests and scripts/doc_claims_check.sh's catalogue pin it verbatim. Restating a record identically is still fine; twelve files here restate 'a list.
scripts/module_type_check.sh asks fifteen questions, and asks the twins side by side: every spelling that must work is asked of a record AND of a variant, because the bug was not "records are broken", it was "records and variants disagree" and nothing compared them.
2,847 unit tests, parity 214/214 with no skips.
v0.1.524 — 2026-09-24
_The same control the Wasm assertions got now covers C and LLVM: 352 substring assertions, every one of them able to fail._
v0.1.523 gave the 97 Wasm assertions a control and found 41 vacuous and eleven FALSE. The same shape was in codegen: (121) and llvm: (138) with no control at all. ⚠ This time there were no lies — measured, not assumed. It is prevention, and it is worth the same as the find: this exact hole kept eleven wrong claims green for a month.
Sixty of the C assertions could not be reached from outside. They are written as let out = codegen "…" in … assert_contains … out …, which no regex over the source can follow — which is why they had never been measured. Moving the check INTO the assertion makes the shape irrelevant: assert_c name out needle does not care how out was produced.
The rule is written once. Three backends print three different things, so what changes between them is how a line says "a function starts here", what ends a body, and how a comment is spelled — not the idea. Unifying the three found two bugs in the version that had been shipping:
- ⚠ "a line at column 0 ends a body" is wrong for LLVM: a basic-block label
sits at column 0 INSIDE a function, so that rule ended every body at its first label and left 5,103 of 5,319 runtime lines in.
- ⚠ A function header starts a body wherever it is written. The Wasm runtime
has one (func nested a level deeper than the other 57; requiring top-level indentation kept its whole body, and with it six assertions went vacuous.
Fifteen C and thirty-nine LLVM assertions take the runtime exemption, and the number is pinned. Two reasons, both written down: the claim is about the runtime rather than about a program ("the idiv helper still uses sdiv", "region_alloc bounds-checks before bumping"), or ⚠ the control uses the same generated name — anon_0_fn is what the first anonymous function is called and the trivial program already has one, so a subject's adapter cannot be told from the control's by name.
scripts/wasm_assert_strength.sh asks all three: a floor on each population, a ceiling on each exemption, and zero assertions left in the unchecked form.
⚠ Not covered, and counted so it is visible: 292 assert_contains remain under prefixes that do not name a backend (vec: 15, owned_vec: 7, of_json: 5, …). Some of them read C output. The prefix cannot be used to sort them, so they are left, and said so.
2,847 unit tests, parity 214/214 with no skips.
v0.1.523 — 2026-09-24
_The Wasm substring assertions now have to be capable of failing, and eleven of them turned out to be false._
test/test_basic.ml checks the Wasm backend by compiling a small program and asserting that some string appears in the emitted module. That is a check only if the string would be ABSENT from a program without the feature, and 41 of the 97 were not: the module carries the whole runtime — 58 $__lang_* functions for the program 0 — so a needle naming an opcode matched whatever was compiled. scripts/wasm_assert_strength.sh has reported that number since v0.1.201.
⚠ The record said none of them were false. Eleven were. Asked about the USER'S code instead of about the module, i32.mul is i64.mul, i32.eq is i64.eq, i32.and is nothing at all, and unreachable and (loop $lp are not in the program's own code either. They are survivors of the i32→i64 widening — the same class as the twenty fixed in v0.1.201, still here because that sweep only asked what was false MODULE-WIDE, and module-wide every one of them was true. A needle that matches the runtime is not merely weak evidence; it can hide a claim that is wrong.
The check moved into the test. assert_wasm compiles the control program 0 once and refuses a needle that also appears in the control's own code, so a vacuous assertion fails the moment it is written rather than being counted by a script somebody has to remember to run. What counts as "the program's own code" is decided BY THE CONTROL — every function the trivial program also defines is boilerplate, $main excepted — because a list of name prefixes has to be maintained, and the first thing such a list got wrong was $show_WCgCol8, generated for the user's own type and named like a runtime helper.
Twelve assertions keep a documented exemption: (module, the exported memory, the imported puts and the like are the skeleton every module has, and asking them to be discriminating would be asking for a lie. That number is pinned, because an exemption nobody counts is how 41 of them got here. Four were deleted as duplicates of a call-site assertion on the same program.
⚠ And the script that measured this reported "0 assertions, 0 vacuous" and exited green once the renaming was done — its denominator had gone to zero. It holds a floor now, and its poison adds an assertion in the old form to a copy of the source to prove the detection works.
93 wasm assertions, 0 vacuous, 0 false. 2,847 unit tests, parity 214/214.
v0.1.522 — 2026-09-24
_Q-052: a local let was writing a top-level binding that happened to share its name, on the LLVM backend, silently. Four lines reproduce it._
let n = 7;
let f = fn (u: int) -> n;
let _ = print_int (codepoint_of "a");
print_int (f 0)
The interpreter, C and Wasm print 7. LLVM printed 1. codepoint_of is the STANDARD LIBRARY's, and it binds its own n inside — nothing had to be imported, no closure trick, no volume of allocation. Any program with a top-level let n that reached it got a wrong answer, which is the failure direction with no symptom.
The backend decided "does this let initialise a file-scope global?" by asking whether the NAME was registered in top_globals_llvm. A name is not a binding. llvm_in_top_level_body now marks the one context where the question makes sense — the top-level let spine of the main body — and every other let is a plain local, whatever it is called.
⚠ The same bug was found and closed on the Wasm backend a month earlier (wasm_in_top_level_body: a local let entries overwrote a KV strbuf pointer and kv_save then wrote 0 bytes). That fix shipped with no gate, and the twin stayed broken. scripts/toplevel_shadow_check.sh asks both backends now, plus the whole corpus for the emit-time footprint: no @mu_* global may be stored from a function other than @main. That was 23 sites across 6 examples before this and is 0 now — and five of those six still printed the right answer, so the behavioural question alone would have called them fine.
Two pinned divergences come off. capture_after_call was Q-052's own case. llvm_loop_guard_global was recorded separately as "a loop whose bound is a top-level binding stops after one iteration on LLVM… whatever n is read from after the lifted thunk has been called is not the global that holds 4" — the same root cause, diagnosed as a loop-guard problem and never connected. Parity's stale-pin detector is what said so.
⚠ Q-052's recorded narrowing was wrong in three places, and all three were written as measurements: it is not the closure's captured value (no closure is needed), it does not need allocation volume (four lines), and the global is not intact (it is the only thing broken — the direct reads that looked fine were using an in-flight value instead of reloading).
2,850 unit tests, parity 214/214 with no skips.
v0.1.521 — 2026-09-23
_mere fmt stops deleting a file's imports and writing somebody else's code into it (Q-174), and the gate that could not see that now asks the question it was missing._
import SPLICES: by the time anything downstream sees a program, the import statement is gone and the imported file's declarations are sitting where it was. For a compiler that is the whole point. mere fmt printed that, so formatting a file DELETED its import lines and copied the imported declarations in — nine imports became zero and seven externs became nineteen on one example — and mere fmt -i saved it over the source.
⚠ Both existing fmt gates were green on it. The inlined output type-checks and is stable on a second pass, so "what fmt writes is still a program" and "formatting twice gives the same file" both pass. They are two spellings of is the output broken; nothing asked whether it is the SAME program. A third question does now: every example with imports keeps every one of them (92 of them), and a fixture pins that the imported file's declaration is not written into the output. Its own poison, because a ceiling poison cannot reach a bug whose symptom is a count going to zero.
The parser records each of the entry file's imports as (index in the declaration list, path as written, how many declarations it spliced, source line), so the formatter undoes the splice exactly rather than guessing from positions — a nested import is already inside its parent's count, and an import that brought nothing in because something else had already pulled that file in is still a line the source wrote, so it still comes back.
Two things followed from it.
extern declarations can be placed now, which recovers the 12 trailing comments v0.1.520 had to leave: placing them used to pick a line from whichever copy of a duplicated name came first, and the duplicate was the imported file's. The ambiguity guard stays, because two externs of one name in ONE file is still something a person can write.
And fmt_comments_check.sh was measuring against a contaminated output. LOST_CEILING went 53 → 111 → 99 in this version and nothing was newly lost: the output used to carry the imported files' declarations, and contrib is full of string literals containing // ("http://", Url._cred), so the subtraction was crediting the output with another file's text. Measured separately, strictly lost comments of the entry file went 151 → 108 over the same corpus and no file lost more of its own. The number is not comparable across this version, and the gate says so where the ceiling is set.
v0.1.520 — 2026-09-23
_A program can say how much stack it needs, mere doc exists, and the gate that was about to certify the next ceiling was not guarding its own denominator._
Q-168: `stack = "512MB"` in mere.toml. Deep recursion has been DIAGNOSED since v0.1.271 -- it says "stack overflow (recursion too deep)" instead of exiting 139 with an empty stderr -- but nothing in the language could ask for more. The answer lived outside the build in three spellings: -Wl,-stack_size on Darwin, ulimit -s on Linux, --stack-size on node. So whoever RAN a program had to know a fact about the PROGRAM.
Translating the request into a link flag could not reach: mere -c emits C and never invokes the linker. But main is a thing this compiler writes, so the C and LLVM backends now put the program's work on a thread sized to the request and join it. Measured both ways on both: two million frames overflow without it and complete with it. A trivial program built with the same request holds 1.5 MiB resident, so it is address space and not memory. Wasm and RV32IM refuse by name, because there the stack belongs to the JS engine or the linker script. A size this cannot read is refused rather than ignored -- a request nobody honours is worse than no request, because the program looks like it asked.
Q-169: `mere doc`. Both halves already existed with no exit between them: --decls knows every top-level name and its inferred type, and doc_above (v0.1.506, hover) knows the comment block a definition was written under. This is those two joined and nothing else. --decls --json gained doc and line, added and not instead of -- decls_json_check.sh still rebuilds the text output from those fields, byte for byte.
An undocumented name is printed with no block under it. Leaving it out would merge "this file does not export that" with "nobody wrote a comment", and the second is what a reader is trying to find. mere doc also lists only the file's OWN names: import splices, so the walk sees the imported file's top-level names too, and --decls prints them on purpose (they are in scope, and that output is for pasting) -- but "what does this file document" must not answer with somebody else's names.
⚠ The comment ceiling went 96 → 53, and finding the rest of it found two defects instead.
The tractable half landed: the comment on an if ... then line was keyed on the THEN-BRANCH's line, and the branch is on the NEXT line, so the key never matched and 42 of these were dropped. It is keyed on the condition's line now. When the branch starts on the then line the source wrote the comment after the BRANCH, and the branch keeps it.
The other half did not, and the reason is worth more than the twelve comments it was about. extern declarations carry a position the parser already keeps, so placing them looked like four lines -- and it destabilised two example files. mere fmt prints the SPLICED program: a file importing contrib/http/query.mere comes back with extern fn http_current_body TWICE, both now in one file, so the name is not a key and whichever line is chosen is right in one pass and wrong in the next.
Which is the small symptom of something larger, now Q-174: mere fmt -i DELETES a file's import lines and writes the imported file's declarations into it. Nine imports became zero and seven externs became nineteen on one example. The idempotence and round-trip gates could not see it -- the inlined output still type-checks and is stable on a second pass. Loc.file already says which declarations came from an import, which is what a fix would read.
And the gate under all of this was not counting its own corpus. lost is a difference taken over the files fmt accepted, so a file that starts being refused leaves BOTH sums and the gate gets greener -- a false pass, which is the one thing a ceiling cannot show. Four of the 291 examples do not format and nothing said so. fmt_comments_check.sh now classifies refusals by reason the way region_params_check.sh does for the same corpus, names the three that are deliberate, fails on one nobody listed, fails on a name that starts formatting again, and holds a floor of 280 measured files. Three poisons, because an impossible ceiling cannot reach a bug that makes the number smaller.
v0.1.519 — 2026-09-23
_The two ceilings v0.1.518 measured are gone, and one of them was not what the last entry said it was._
v0.1.518 wrote that the three non-idempotent files were "a comment moving one indent level" -- placement drifting rather than anything being lost. That was wrong, and it was wrong in the reassuring direction. Measured properly, a comment on line 380 of the first pass is on line 805 of the second: it moved 425 lines, onto code it does not describe. The diff looked like indentation because the - and + lines carried the SAME TEXT and only the leading whitespace differed -- the line numbers were never read.
The cause was that comment placement is decided by source line, and Loc.line counts within its OWN file. import splices another file's declarations in, so the entry file's line 380 and an imported file's line 380 were the same key. Placement now only considers comments whose location carries no file -- the entry file's own.
Three more things, from the same corpus:
- The formatter prints
dyn Trait eagain. The parser desugars it to
Trait__pack e unconditionally -- it is in the parser, where keep_sugar does not reach -- so the formatter reconstructs it, and only for names the file actually declares as traits.
- Trailing comments survive on match arms and on
else ifheads. Over the
corpus: 7,584 comment lines in, 7,488 out -- 188 lost → 96. The restriction is narrow on purpose: a comment is only attached where the formatter itself emits the newline, because attaching it to a final arm or a closing else puts the caller's ; inside the comment -- that broke 13 files before the rule was narrowed.
mere checkreports top-level bindings nothing reads, in files that mark their
exports with pub (Q-146). Files that mark nothing stay silent: without pub there is no way to tell a library's surface from dead code, which is exactly why this was deferred in the first place. The gate asks both directions.
fmt_roundtrip_check.sh now runs at CEILING=0 and DRIFT_CEILING=0 over all 293 example files: what mere fmt writes type-checks, and formatting twice gives the same bytes. fmt_comments_check.sh is at LOST_CEILING=96.
v0.1.518 — 2026-09-23
_Q-173 closed to a measured ceiling: keep_sugar guarded the desugaring and nothing else._
Five passes run after it and were not guarded at all -- range-check versioning (Q-108) splitting a loop into f__rvfast / f__rvslow, inner-function uniquifying renaming go to go_uq3, par_map lowering, top-level shadow uniquifying, main reservation. All of them are preparation for CODE GENERATION, and mere fmt printed them back: 19 example files came out carrying names nobody wrote, and formatting the output split the already-split loops again.
That is the sentence v0.1.504 wrote about echo, one layer down -- a formatter that edits the source it formats is a tool people stop running. Formatting skips them now.
Non-idempotent example files: 109 → 3. The three left are a comment moving one indent level on the second pass, which is placement drifting rather than anything being lost -- the inline comments landed by v0.1.511 are put back by source line, and a line moves when the code around it reflows.
scripts/fmt_roundtrip_check.sh owns both questions now, over the whole corpus and with measured ceilings: what the formatter writes still type-checks (1 left, a trait's internal __pack constructor) and formatting twice gives the same file (3 left). fmt_comments_check.sh says in its header that the idempotence question moved -- formatting ONE fixture twice is what let 109 files drift unseen.
v0.1.517 — 2026-09-23
_Q-172: what mere fmt wrote was not a program, and the cause had five layers._
The parser FLATTENS a module -- its members become top-level bindings called M.foo -- and the formatter printed that: let Bignum.base = 1000000000;, which is not syntax. mere fmt -i rewrites in place, so 52 of the 293 example files came back as something the compiler refused. The formatter's own tests were all small single-file samples; the corpus had never been asked.
Putting the block back is a grouping pass, not a tree rewrite: only the binding name loses its prefix, because a qualified self-reference is valid inside the module it names -- the parser registers the module before parsing its body for exactly that. Then four more layers, each found by fixing the one above:
- Constructors keep their prefix wherever they are USED, and the rule is
the OPPOSITE inside and out. M.C is unspellable inside module M (a parse error in a pattern, unbound in an expression) and REQUIRED outside it -- examples/module_scoping.mere has Red in two modules and tells them apart that way. Stripped inside the block only.
- A type declared in a module was printed outside it. The declaration
carries no trace of the block, but Top_ctor_alias ("Traffic.Red", "Red") does, so the type goes back where its constructors were declared.
- That ownership lookup misfired on a shared name. Keeping the last
owner of Red put type Light inside Mood. A type is placed by the first constructor that names exactly ONE module.
- `A [1, 2]` does not type-check. A list literal is fine as a FUNCTION
argument (sum [1, 2, 3] runs) and not as a constructor's -- both A [] and A [1, 2] come back as constructor A requires an argument. The formatter wrote the bare form in both.
⚠ And parenthesising the literal everywhere broke the self-host cross-validation: contrib/fmt/fmt.mere and this formatter are held to byte-identical output on a set of samples, and that one does not parenthesise the function case. The parens went to the constructor site instead. Two formatters agreeing is what stopped a fix that was too wide.
scripts/fmt_roundtrip_check.sh asks the corpus directly now: every example the compiler accepts is formatted and the OUTPUT handed back to the compiler. 52 → 1, and the one that remains is a trait's internal __pack constructor reaching the output -- a ceiling, measured, so the next cause shows up as the number moving.
v0.1.516 — 2026-09-23
_Q-171: a catch now releases the region it jumped over, on the backend that had never been able to._
A fail raised inside a region R { } and caught by an outer try_or LONGJMPS PAST THE BLOCK'S EXIT, so the release written there never runs. The C backend has carried an active-region stack for exactly this since v0.1.31, and made it release rather than leak in v0.1.301. The LLVM backend had neither, and kept the region struct in an alloca -- on a frame that is gone by the time anyone could free it. The loop segfaulted at a HUNDRED iterations where C ran twenty thousand; Wasm and the interpreter were fine.
Two changes. The struct comes from malloc, and every live block is on an active stack, so a catch releases what it jumped over by depth alone -- try_or and try_or_msg save @__lang_region_active_n on the way in and call @__lang_region_unwind on the failure path.
Simpler than the C version on purpose. This backend REFUSES region loop, so its blocks are strictly LIFO and a release pops the top; C carries a scan-and-shift because its region-loop swap releases the entry under the one just pushed. Porting that too would have been code with no caller.
scripts/region_unwind_check.sh runs the loop on every backend that can build here, at a scale where leaking the 1 MiB block could not fit in memory -- so passing is evidence the blocks come back, not just that the answer is right. Its poison is the shape of the bug: one catch passes on the broken backend too, which is why the gate runs twenty thousand.
⚠ What it does not claim: total memory. Mere reclaims nothing by default, so peak RSS still grows with the iteration count on every backend, and by different amounts (N=50000: C 2.6 MiB, LLVM 9.7 MiB). That is allocation shape, not this bug -- a leaked block would have been fifty gigabytes.
v0.1.515 — 2026-09-23
_Q-173: the formatter escaped five characters and the lexer writes seven._
A carriage return went out of mere fmt as a RAW CR. The lexer reads that as a line break -- newline in string literal -- so 69 of the 315 example files could not be formatted twice. That was the visible half. The other half is worse: formatting the output again DROPPED the byte, so mere fmt -i on a file with CRLF in a string, which is every Redis and HTTP string in contrib, silently changed what the program sends. A formatter that rewrites in place may not lose a byte. \0 had the same hole.
The fix is two lines in escape_string_for_fmt, and the rule behind it is that the set it writes has to be a SUBSET of what the lexer reads back: n t r 0, backslash, quote, and the interpolation braces. Anything else stays a raw byte and round-trips, because the lexer copies unknown bytes through -- only these two were read as something else.
scripts/fmt_comments_check.sh asks the new question directly: a string holding CR, NUL, tab, backslash and a quote keeps its LENGTH through a format, and formatting twice gives the same file. Its fourth poison is the old behaviour -- drop an escape and the length changes.
_And the number it was hiding._ Non-idempotent files went 109 → 52, while the count of files whose output does not PARSE went 21 → 33: a dozen files used to fail at the carriage return before they could reach the other defect. Two independent faults, and the first one was masking the second.
v0.1.514 — 2026-09-23
_Q-166: pub at the top of a file, and the boundary the splice had not erased._
import "path"; splices the imported file's declarations into this one, so after an import there is a single top-level namespace. The language reference drew the obvious conclusion and said file-level visibility had "nothing to enforce": a library's helper was as reachable as the function it was written for, and mere could not report an unread top-level name because it could not tell a file's surface from its insides (Q-146).
The boundary is not gone, though — it is in Loc.t. Every token carries the file it came from, set by the lexer for anything that arrived through an import, which means a REFERENCE can be checked against the file that made the binding. pub let at top level now says the file has decided its surface, and the check is the module one (v0.1.504) one level out.
Opt-in per file, as it is per module: a file that marks nothing exports everything, which is every file written before this. pub stays a contextual keyword — a program that binds it as a name is unaffected.
The rule that keeps it usable: a file that binds the name ITSELF is unaffected by what another file decided about its own copy. The single namespace means two files may bind the same top-level name; without that, one library marking pub would make a common name unusable in a file that never imported it. scripts/file_privacy_check.sh holds all four directions with two poisons — including that removing the marker makes the refused call succeed, so the gate is watching this failure and not some other one.
⚠ This is visibility, not separate compilation. The splice still happens: the program is still one translation unit and mere -c still reads the whole tree. What changed is what may be REFERRED to.
_And the claims gate caught the reference again._ File-level visibility has nothing to enforce is retired wording now. The first row written for the new sentence reported PHRASE GONE from a file that said exactly what it claimed — the phrase had wrapped across a line, and the check is a fixed-string grep per line. The header says so now.
v0.1.513 — 2026-09-23
_Q-167: four spellings the lexer had never heard of, and three of them were name errors rather than syntax errors._
0b1010 came back as `unbound variable: b1010`. The lexer read the 0, stopped, and handed b1010 to the typer as a name — so the diagnostic was true and useless, and nothing in it said "this language has no binary literals". Same for 0o17 and 1_000. int_of_string had understood all three all along, and the digit separators too: what was missing was the lexer, not the conversion.
Binary (0b / 0B), octal (0o / 0O) and _ between digits, in every base and in a float — 1_000_000, 0xFF_FF, 0b1010_1010, 1_000.5. Each _ has to be followed by another digit, so 1_ is still the integer 1 next to an identifier, which is the reading a program that binds _x already relies on.
`\uXXXX` is the one that is not a number. Exactly four hex digits, encoded as UTF-8, so "\u3042" is three bytes and utf8_len counts it as one character. A surrogate half is refused with a sentence rather than written out: it is not a character, str is bytes, and nothing downstream would put a pair back together. Beyond the BMP is written as the character itself — this file is UTF-8 and the lexer copies unknown bytes through.
`scripts/doc_claims_check.sh` caught its own docs. The gate landed two days into its own life with No Unicode escape (\uXXXX) and no octal or binary literal syntax and no digit separator in its catalogue; the moment the lexer learned them, it went red naming docs/language-reference.md:613 and :614 before either sentence had been touched. Both wordings are retired now, which is the other half of the same gate.
_Measured and not fixed._ The self-host lexer (contrib/parser/lexer.mere) splits every one of these into three tokens — including `0xFF`, which this compiler has had since v0.1.46. It was behind before this change and is behind by three more spellings now; that is Q-153's ground, and the number is in it.
v0.1.512 — 2026-09-23
_Q-039: the shift count had four answers, one of them undefined._
bit_shl x n with n at or beyond the width meant four different things. The interpreter defined it. The C backend's << on a signed long long was undefined behaviour — and not theoretically: UBSan on the emitted code says left shift of 4611686018427387904 by 1 places cannot be represented in type 'long long', and a count of 64 or more is undefined on its own. LLVM IR calls a shift by 64 or more poison. Wasm masks the count mod 64 by spec, so bit_shl x 70 quietly meant x << 6 there.
The contract was already written down, in the RISC-V backend: a count at or beyond the width gives zero for a left shift and the sign bit for a right one, chosen there "to match what the other backends give". Three backends had never implemented it. They do now — C shifts in unsigned long long and guards the count, LLVM masks the count so no poison is produced at all and selects, Wasm tests the count before shifting — and the interpreter's one disagreement went with it: a NEGATIVE count used to return x unchanged from bit_shr, which no compiled backend could reproduce because they compare the count as unsigned. One rule, four backends, test/parity/shift_counts.mere.
And the boundary between the two integer widths is pinned rather than avoided. The lexer refuses a literal above 2^62-1 on every backend and says why; nothing enforces that on a computed value, and 2^62 is exactly where the interpreter's 63-bit int and the compiled backends' 64-bit one part company: bit_shl 1 61 is one number everywhere, bit_shl 1 62 is two. test/parity/int_width_boundary.mere declares that difference to the harness, which reads it as DIVERGE — so moving it is a failure rather than a surprise, and the corpus stops quietly staying under 2^62 without saying why.
No `bit_ushr`, and the reason is that boundary rather than effort. A logical right shift is a statement about the top bit, and the top bit is what the two widths disagree about: bit_ushr (-8) 1 is 2^63-4 on the compiled backends, a value the interpreter cannot hold. Documented in the stdlib reference next to the mask convention the varint code here already uses.
v0.1.511 — 2026-09-23
_Q-154: the comments the formatter was still deleting, and two holes found under them._
v0.1.505 taught mere fmt to keep comments written in column 1. The other two kinds were documented as dropped: an indented comment belongs to an expression and the tree has no field for one, a trailing comment belongs after a node whose extent no Loc.t records. The lexer had both all along -- it records a position for every comment and the collector threw away everything that was not in column 1 -- so what was missing was placement, not collection.
Indented comments go back into the run of `let`s they were written in. That run is the one place this formatter emits its own indent, which makes it the one place a comment can be put back without guessing a column; 516 of the 654 in the examples corpus sit directly above a let. Placement uses a watermark rather than draining everything written before the line, so a comment written above a whole block is not dragged down to the second binding -- it falls through to the old fallback and is printed above the declaration. Moved, which this formatter has always preferred to lost.
A trailing comment goes back at the end of the line it was written at, when that line is one the layout emits in one piece -- a let whose value fits, or a one-line declaration. 375 of 636 are on a line that starts with let.
On examples/: 8,181 of 8,426 comment lines survive, against 7,168 before. What is still dropped is a trailing comment on a line the formatter does not emit whole (an else, a match arm, a multi-line binding), and scripts/fmt_comments_check.sh now holds a measured CEILING on the loss rather than a note -- the next slice shows up as the number going down, a regression as it going up, and a third poison proves the ceiling can refuse.
Two properties this arc measured and did not fix, both older than it. On the 315 example files: mere fmt output does not parse for 21 of them, and formatting is not idempotent for 109. The gate's idempotence check formats ONE fixture twice, which is why neither number was visible; it says so now. Both were measured against this compiler with the change stashed, so they are the baseline rather than this slice's doing -- the first version of the indented placement added 47 files to the non-idempotent count, and that was fixed by refusing to write a comment when the buffer cannot know what column the caller left the cursor in.
v0.1.510 — 2026-09-23
_Q-165: a caught failure can say why._
try_or could see THAT something failed and not WHY. What it handed back was the default; the message the raiser wrote went nowhere, and a library written here could not offer "branch on the kind of failure" to its callers -- contrib/url's parse gave up fail and returns ?url_parts for exactly that reason. try_or_msg : (unit -> 'a) -> (str -> 'a) -> 'a applies its handler to the message instead.
What the handler receives is the diagnostic line, the bytes the failure would have written to stderr had nobody caught it: fail "boom" arrives as fail: boom, tag included, and a failure raised inside a backend arrives under its own name. The tag belongs to the fail builtin rather than to the printer, which is why that string is the one every backend already had in hand at the catch -- taking it off would have been four different subtractions to keep equal.
Three of the four had the message and dropped it. The C backend has copied it into __lang_fail_msg since v0.1.67, for the --lib boundary's err buffer, and nothing in the language could read it; LLVM received the pointer in __lang_fail_impl and longjmped without it; Wasm received it and set a flag. They copy it now, into a fixed 256-byte buffer on C and LLVM and into reserved memory on Wasm, because the string the raiser built can be above a bump mark that a block rolls back on the way out. At the catch it is copied AGAIN, into the catcher's region: the buffer is one, and a second failure -- including one raised by the handler -- would rewrite what the handler is reading. try_or's 211 call sites across five repositories are untouched; this is a new name, not a new signature. The RISC-V backend refuses it by name.
The shape that only exists once there is a handler. A handler that fails -- "translate this into my own error", the ordinary reason to want one -- longjmped back into the catch that had just called it, which called it again, forever. try_or cannot have this bug: its default is evaluated before the _setjmp, so nothing on its failure path runs under its own jmpbuf. The catch comes down before the handler runs now. The parity suite found it, as a program that printed three lines fewer on one backend and then did not stop.
test/parity/fail/reason_*.mere is the gate: twelve programs, each writing the same failing expression twice -- once inside try_or_msg, whose handler prints what it was told, and once bare, so the process ends with that same failure. scripts/fail_reason_check.sh compares the two on every backend that can be built here, with a poison that prints a different reason.
_And a hole this arc found but did not fix._ A try_or catching a failure raised inside a region R { }, in a loop, segfaults on the LLVM backend at a hundred iterations where C runs twenty thousand. longjmp skips the block's exit, so @__lang_current_region still points at the region that was jumped over -- the C backend saves and restores it (v0.1.31) and releases what it jumped over (v0.1.301), and this backend has no __lang_region_unwind at all. try_or_msg restores the pointer on its own failure path, because the message it allocates has to land somewhere owned; try_or is unchanged and the hole is recorded rather than patched. It went unseen because nothing in the caught-side parity had ever failed from inside a region block.
v0.1.508 — 2026-09-22
_One rule written in one arm of four twins, and the spelling that never got checked._
`let rec` was skipping two of `let`'s safeguards. A top-level binding can be spelled let f = ..., let rec f = ..., as a member of a let rec ... and ... group, or inside a module. Four loops walk declarations — the interpreter's, type_of, region_param_report and infer_program_inner — and every one of them routed Top_let through infer_top_let and top_let_scheme, then inlined the Top_let_rec case by hand. Both of that pair's checks were therefore missing from all four:
let store = vec_new ();
let keep = fn (n: int) -> ... vec_push store (vec_new ()) ...; // refused under --lib
let rec keep = fn (n: int) -> ... same ...; // compiled
Under --lib each exported call runs in its own region, so the second program put a pointer into memory the next call reuses — the exact fault the boundary check exists to stop, reachable by changing one word. The value restriction had the same hole: let store = vec_new (); is monomorphic, and spelled let rec store = vec_new (); it was generalised, after which one Vec held an int and a str and neither use complained.
Both now go through one entry point, infer_top_rec, which the four loops share the way they already shared infer_top_let.
The gate is about the asymmetry, not the rule. Each property is run in all five spellings and must get the same verdict, so the next rule that lands in one arm fails here even when the arm it landed in is right. It includes a program that must be ACCEPTED in every spelling — a gate whose every check is "this is refused" passes on a compiler that refuses everything — and its poisons are a compiler that refuses only the let spelling, which is the shape of the bug, and one that refuses everything.
This was found sideways: an experiment that grouped adjacent let bindings made a regression test stop refusing what it refuses. The experiment was discarded; the hole it had routed ordinary lets into was older than it.
Counts
dune test 2817, parity 180, and one more gate script with poison runs in CI.
v0.1.507 — 2026-09-22
_The first layer a newcomer hits had twenty-six explanations everywhere else and none of its own._
A syntax error now says what to write instead. The typer, the exhaustiveness checker and the pipeline carry twenty-six help: lines between them; the lexer and the parser carried zero, although the syntax layer is the one a reader meets first. Seventeen spellings borrowed from other languages were tried against the compiler and sixteen came back with a generic message — several of them actively misleading, because var x = 1; and def f(n): both produced trailing input, and Python's a and b produced expected ';' or 'in' after let binding (and is Mere's keyword for a mutual-recursion group):
parse error: expected ';' or 'in' after let binding
--> x.mere:1:35
|
1 | let _ = if true then 1 elif false then 2 else 3;
| ^^^^
|
= help: `elif` — chain with `else if`
It is one table, not 149 edits. Parse_error is raised in 149 places and none of them were touched. A hint rides on the message, which already exists; what was missing is the answer to what did the reader actually write, and the token list answers that at the two entry points that parse. The typer reads the same table for the words that get past the parser as unbound names (def, return), where the fuzzy suggestion used to answer did you mean \e\?.
A word is only a keyword if this file does not bind it. var, case, val and mut are real identifiers in the Mere repositories — let var = list_sum ..., fn (case: int) -> case * 2, fn val ->, &mut R v, and as tuple pattern binders, let (val, j) = ... and let (classes, var, prefix, r2) = ... — so the table is given the names the file binds and stays silent about those. of is not in the table at all: it is Mere's own keyword (type t = A | B of int, measured 943 times across 989 source files) and never reaches the parser as an identifier.
The window is the failing line rather than the failing token, because the evidence is often a few tokens back: if c then 1 elif x < 0 then 2 fails at the second then, four tokens past the elif that explains it. Every hint names the token it keyed on, so a reader who was doing something else can see at once that the hint is not about them.
scripts/syntax_hint_check.sh holds all eighteen rows, and its poisons are the risk the design took on: != and let rec ... and are correct Mere and must stay silent, a file that binds one of the four words must not be told about another language, and with the help: lines stripped every row must go red.
Counts
dune test 2810, parity 180, and one more gate script with poison runs in CI.
v0.1.506 — 2026-09-22
_Two diagnostics that were right about what was wrong and silent, or wrong, about where._
A failure inside a prelude function points at the caller. assert, divmod and list_max are written in Mere, so their fail carried a position in <prelude> — a text nobody can open — and the line the person actually wrote appeared only in the stack below:
eval error: fail: assertion failed: boom
--> <prelude>:539:24 ← before
--> /tmp/x.mere:1:9 ← now
The innermost frame that is not the prelude's is where they can act, and the backtrace had computed it all along. A builtin implemented in OCaml (char_at) never had the problem, which is why this looked like one oddity rather than a class of them.
Declaring a type twice now says where, twice. The message was good and pointed at nothing at all (Top_type carries no position):
type error: type `t` is declared twice with different constructors
(`A | B` and `C | D`) (first declared on line 1)
--> x.mere:3:6
The parser's table records where every type name was declared, and a name declared twice has two entries in it — the caret goes on the second, the message names the first. One location is not enough when the error is about a pair; Gleam spends a whole secondary-label mechanism on this and uses it in fifteen places, which is worth remembering if a second case turns up here.
scripts/diagnostic_position_check.sh holds both, with poisons in the directions that would make it pass for the wrong reason: an OCaml builtin must still point at the user, and a type declared once must not be reported.
Counts
dune test 2800, parity 180, and one more gate script with poison runs in CI.
v0.1.505 — 2026-09-22
_One question that had two answers, and a formatter that deleted what the editor had just started reading._
A `let` that can fail is refused at compile time. let Some n = e; is a match with one arm, and until now the two constructs answered the same question differently: the match was refused with non-exhaustive match (missing None) and the let compiled, then failed at run time with top-level let pattern did not match. The witness comes from the same find_missing the match uses, so the error names the value that does not match:
error: this `let` pattern does not match every value (missing None)
= help: write it as a `match` with an arm for None
= note: a `let` binds; it has no other arm to take when the pattern does not
match, so this would fail at run time
What stays free: let (a, b) = ..., record patterns, a plain name, and a constructor pattern on a type that has only one constructor — that one is total, and refusing it would be refusing the ordinary way to open a wrapper. if let is untouched: the parser turns it into a two-armed match, which is the construct for a pattern that may not match.
⚠ THE SAME SPLIT REAPPEARED ONE LAYER IN. With the check written, mere check and every emit path refused the program while the interpreter still failed at run time — because it walks declarations and RUNS each one as it goes, so a refutable let reached its own failure before anything drained the findings. The findings are now drained before each declaration is evaluated, and all three paths give the same answer.
Measured before writing it: zero refutable lets in this repository, mere-ruby, mbrowse, medit2 and m3d. The sixteen places that look like one are all if let.
`mere fmt` keeps the comments in column 1. It used to delete every comment in the file. That was written down as an MVP limitation and it was survivable while nothing else depended on comments — and then v0.1.504 taught hover to read the block above a definition, so one half of the toolchain was reading what the other half deleted. textDocument/formatting is the same function, so format-on-save was the delete.
Column 1 only, and the reason is a measurement: of the 23,452 comment lines in this repository's .mere files, 19,338 (82%) start in column 1 — and those are exactly the blocks hover reads. The other two kinds need something the tree does not have: an indented comment belongs to an expression, and a trailing one belongs after a node whose extent no Loc.t records.
The material was already there: the lexer has been able to collect comments since semantic tokens needed them, and it records a position and a width, so the text is a slice of its own line. Nothing had asked for them.
⚠ Most top-level declarations carry no position at all — only Top_let, Top_let_rec and Top_forward do, which is why Parser.declared_types exists for the rest. A declaration that cannot be placed does not collect the comments above it; they go to the next one that can. Losing a comment is not an option, moving one down is the fallback, and it is written down here rather than discovered later.
scripts/fmt_comments_check.sh holds three things: a column-1 comment survives and stays above its declaration, formatting is idempotent, and over examples/ the count of column-1 comment lines does not go down — 287 files, 6,316 lines in, 6,316 out. It also prints what is still dropped (~612 indented lines in examples/) so that the remainder stays visible rather than forgotten.
Counts
dune test 2796 → 2800, and two more gate scripts with poison runs in CI (refutable_let_check, fmt_comments_check).
v0.1.504 — 2026-09-22
_The second pass over Gleam: four answers the compiler already had, a flag that makes yesterday's warnings mean something, a machine-readable API surface, one pattern, and pub._
_The previous pass's five landed in v0.1.503; this is the outside of that circle, chosen by the same rule — take what is a wire rather than a new analysis — plus one thing that was measured and dropped (below)._
Four LSP answers, none of them new work. The server advertised nine providers; Gleam enables twelve. Three of the four missing ones had their answer sitting in this repository already:
| what it needed | where the answer was | |
|---|---|---|
| document highlight | one handler | Query.references_at — the same list find-references returns, drawn in the file instead of in a panel |
| go to type definition | one handler | Parser.declared_types, which the parser fills because Top_type carries no position of its own |
| signature help | one handler | Ast.rv_spine, which the range-check pass already uses to split f a b into a head and its arguments |
The fourth, foldingRange, is not in this release and is not an oversight: a Loc.t is a start and a width, so nothing in the tree knows where a declaration ends. That is the same missing fact that stopped the redundant-arm quickfix in v0.1.503.
Two things the hover did not do, and now does. It answered nothing on a definition — Query.node_at looks for an expression and the inc in let inc = ... is a pattern, so hovering the place a name is introduced returned nothing at all. And it showed no documentation: the contiguous // lines above a definition are now part of the hover, found through the binding, so hovering a USE shows what was written at the DEFINITION. No /// marker, because Mere has no docs generator for one to feed — requiring a third slash would mean every comment already written shows nothing.
⚠ Signature help needs one guess and says so: without extents, "the call the cursor is inside" is not a question the tree can answer. It answers about the nearest call head at or before the cursor on that line, preferring a call that is still unfinished — which is what stops add3 (pick Red) from reporting about pick.
`--warnings-as-errors`. v0.1.503 gave the compiler warnings worth acting on (an unread binding, a deprecated name) and no way to make a machine act on them, so a CI could not hold the line against new ones. The rule is one rule: everything is printed exactly as it would have been, and then the status is 1. The count is of warnings produced, not printed — a file with twelve of them fails even though the terminal shows ten and a summary.
A machine-readable API surface. mere --decls --json prints what --decls prints, plus the type declarations and their constructors, plus the package's declared floor — the same pairing Gleam's package-interface has, where the version constraint sits beside the interface. What it is for is the diff between two versions of a package, which is the question mere fix raised in v0.1.503 and could not answer.
The two outputs are two renderings of one walk, and scripts/decls_json_check.sh is what keeps that true: it rebuilds the text from the JSON and compares byte for byte across the whole parity corpus — 179 files, 32 of them carrying a commented-out declaration, which are the entries nobody writes by hand.
`| "lit" <> rest ->`. The prefix test and the slice that always followed it, as one pattern. In contrib/ alone there were 38 str_starts_with calls, 13 of them followed within two lines by a slice; that shape is now:
match url with
| "https://" <> rest -> secure rest
| "http://" <> rest -> plain rest
| other -> none other
It is a node only as far as Pipeline's desugar, which rewrites the arm into a guard and a binding over str_starts_with / utf8_sub / utf8_len. No backend knows the syntax exists, and a guarded arm closes nothing, so the exhaustiveness checker already says what it should: a prefix arm does not make a match total.
⚠ It is an AST node rather than a rewrite inside the parser for one reason: the formatter. Which is how the next entry was found.
The formatter was rewriting the source it formatted. echo is lowered to echo_at "<where>" while parsing — and the formatter parses. For one release mere fmt on a file containing echo x printed echo_at "line 2" x: correct, equivalent, and not what the person wrote. Sugar lowering is now skipped on that one path (parse_program ~keep_sugar:true), with a test that says so.
`pub`, inside a module. module M { pub let get = ...; let helper = ...; } — M.get is callable from outside and M.helper is not, by name:
type error: `Store.secret` is internal to module `Store` (the module marks its exports with `pub`)
Opt-in per module, because it has to be: every module written before this marks nothing, and making unmarked mean private would have made all of them export nothing. A module that marks nothing is unchanged; one mark says the module has decided what its surface is. pub is not a keyword either — it is an identifier the module-body parser recognises in front of let, so nothing that uses pub as a name stops lexing.
⚠ This is not what Q-146 wanted, and reading the parser is what settled it. File-level visibility has nothing to enforce: import "path" is a SPLICE (parser.ml, the imported file's declarations are prepended to this one's), so after an import every name is in one top-level namespace. pub works inside module M { } because that boundary already exists. The file boundary does not, and cannot be added without the separate-compilation change the design notes have been holding since the last measurement of it. "Add pub and Q-146 closes" was half right.
Measured, and dropped
Gleam compiles patterns through a decision tree; Mere emits a linear chain of tag tests, and never emits `br_table` — zero occurrences in a whole Wasm module. That looks like something worth fixing. It is not, at the sizes anyone writes:
| same program, only the arm that matches differs | 16 arms | 64 arms |
|---|---|---|
C backend, -O2, 20M iterations | 0.02 / 0.02 s | the optimiser folded the loop away |
| Wasm, node, 2M iterations, five runs | — | first arm 0.21–0.33 s, last arm 0.21–0.26 s |
⚠ The first two attempts at this were not measurements: constant folding deleted the loop and printed 0.00 s. Only after the scrutinee came out of a 4,096-entry table did any work survive. The negative result is recorded in the note so the static fact does not produce the same proposal again.
Counts
parity 179 → 180, dune test 2784 → 2796, and three more gate scripts with poison runs in CI (warnings_as_errors_check, decls_json_check, module_privacy_check).
v0.1.503 — 2026-09-22
_Five things Gleam has, in Mere's spelling — and the four bugs that writing them found._
_Read gleam-lang/gleam at v1.18.0 next to this compiler and took the parts that are configuration of what Mere already computes rather than work Mere has not done. The plan, its measurements and what was deliberately NOT taken are in the internal note; what landed is below. Four of the five needed no new analysis at all._
Code actions (`textDocument/codeAction`). Gleam's language server builds 49 of them; this one advertised eight providers and none. It was not missing an analysis — it was missing a wire. Exhaustive has always known the arm to write (it prints it on the help: line), rename has always built a WorkspaceEdit, and Query has always resolved a position. What was in between was a string: classify flattened its findings to (loc, message) and Pipeline.diagnostic carried nothing else, so the arm died one function before anything could apply it. A finding now carries a fix — a title, a position, a width and the text — and three actions are built from answers that were already there: add the missing arm(s), declare a top-level binding's inferred type (let fn f: (shape -> int);, Mere's spelling of an annotation), and prefix an unread binding with `_`.
⚠ The plumbing changes NO output: the same bytes on stderr, the same exit codes, every gate green before and after. Its only witness is a unit test that reads the fix off a diagnostic, and scripts/lsp_smoke.sh now drives a whole round trip — ask for the actions, APPLY the edit that comes back, and require the file to compile where it did not before, with two poisons (apply nothing; apply at the wrong position) that must both leave it refused.
Unused bindings. Gleam has thirteen Unused* warnings. Mere had four warnings in total, all of them the same kind — a top-level name that collides with libc, an extern whose arity disagrees — which is to say the compile error you are about to get, and nothing about code that is simply dead. The check is built on Query.occurrences, the scope-resolving walker hover and rename already use, so shadowing is answered by the same code that answers it for go-to-definition rather than by a second opinion. Local `let` bindings and `match` binders only: a top-level name in a file that is imported has its readers in another file, and with no pub in the language there is nothing to tell a library's surface from dead code. One rule is borrowed wholesale from Gleam and worth naming — nothing is reported for a file that did not type-check, because "nothing reads this" in a half-inferred tree is usually "the line that reads it is the one being typed".
Pointed at mere-ruby's 32,832-line main.mere it answers 462, every one of them true and most of them one idiom: let (meths, sup, ocls, ivars) = world where the arm uses one of the four. Which produced the other half of this change — a terminal prints at most ten warning blocks and then says how many more there are. 462 blocks is four thousand lines of stderr in front of whatever the person ran the compiler to see; the editor draws all of them in the margin, where they cost nothing.
A compiler version in `mere.toml`. [package] mere = ">= 0.1.480" is now read, and checked against the compiler that is running — by mere install (for the package and for every dependency it fetches) and by every build of a file inside the package. Before this, a package needing newer syntax failed on an older compiler as a parse error in somebody else's file: true, useless, and unactionable without knowing the language's history.
The other half is mere fix, which is Gleam's gleam fix — a command that does exactly one thing: it makes the declared floor true. Feature.all maps a feature to the version it landed in (bytes 0.1.278, the 128-bit lane types 0.1.422, f32x4 0.1.445), the parser notes them where the names are already recognised, and mere fix writes the highest into the nearest manifest.
⚠ A table of versions is a claim about history, which is the kind of thing that is right the day it is written and never looked at again. scripts/version_floor_check.sh re-derives every row from `docs/changelog.md` — oldest mention of the feature's probe word, version of the section it sits in — and fails when a row disagrees with the record it came from. Rewriting one row's version by hand turns it red, which is how that sentence is known to be true.
Deprecation, without an attribute. Gleam marks a deprecation on the definition. Mere has no attribute syntax and adding one costs two parsers, this one and the self-hosted one, so the mechanism is a table keyed on the names the compiler itself ships, matched only where the name resolves to the builtin (a user who binds str_len themselves is not told about their own name). The table is empty, and that is a measurement rather than a placeholder: all 281 entries of initial_env were compared for the pairs a rename leaves behind and there are none. The tests install a row to drive the path end to end.
`echo`. echo x prints x to stderr with the line it was written on and answers x, so it drops into the middle of an expression and comes out again without moving anything. Two prelude functions and one pass — no backend knows about it, and all five have it.
⚠ The first version rewrote the token in the parser, which is simpler and wrong: test/parity/graphql_stack_portable.mere binds echo as a name of its own, and rewriting every occurrence turned its value into a partial application. The rewrite now runs over the parsed tree with scope tracked, and an echo the user bound means what they said.
The four bugs that only appeared when it ran
None of these is visible from reading. Each was found by the same program being asked to print the same thing on five backends.
| what was wrong | how long | |
|---|---|---|
| Wasm host | print_err has been an import of the emitted module since v0.1.259 and scripts/run_wasm.js never provided it, so any program calling `print_err` failed to instantiate | 243 versions |
| RISC-V | print_err wrote the bytes and not the trailing newline the other four backends write, so the same program's stderr differed by backend | since it existed |
| `-rv` positions | the RV prelude is glued in front as TEXT, so an echo on line 2 reported line 1925 — diagnostics are corrected on the way out, and nothing corrected a position the compiler puts INTO the program | new, and the same shape as the next row |
| `-rv` warnings | for the same reason, the unused check reported bindings inside the RV prelude against the user | new |
scripts/echo_check.sh is the gate, and it compares stderr across interp / C / LLVM / Wasm / RISC-V, because a debug print that reached stdout would be changing the answer it is watching. The RISC-V row has a stated exception: the emulator's write ignores the descriptor, so the guest's fd 2 arrives on fd 1 and the two streams are put back together before comparing — a fact about the emulator, written down rather than left as a filter.
Counts
parity 178 → 179 programs (the new case is echo, whose value must pass through unchanged on every backend), dune test 2767 → 2784, and four gate scripts with poison runs in CI: unused_check, version_floor_check, echo_check, plus lsp_smoke --poison, which did not exist.
v0.1.502 — 2026-09-22
_The rest of the class, and a check that keeps it out._
_v0.1.501 fixed the six harnesses that had gone RED when Q-136 removed the trailing unit line. It left the ones that were still GREEN, on the grounds that they were correct — their probes end with a bare 0, a real value the trim really does remove. That is true and it is not a good enough reason to keep them: a harness that drops its subject's last line can compare the wrong thing and pass, and the only thing standing between "correct" and "silently wrong" is whether someone remembers the sentinel when they edit the probe._
_So the pair is gone everywhere. Nine more gates print exactly what is compared. In url_parity the sentinel had leaked into the oracle — node was printing a matching 0 — which is how far this kind of thing travels._
_Three trims turned out to be doing something other than what they said:_
| what its comment claimed | what it actually did | |
|---|---|---|
selfhost_check | strips the () the CLI auto-prints | ate the blank line print leaves after a WAT that already ends in a newline |
rv_exec_check ×5 spellings | drops the auto-printed unit | would also drop a () a program printed |
rv_float_check ×3 | same | same |
_Four uses are left and each is named in the new gate with its reason: the diagnostic payload parity.sh READS off the last line, the generated source exhaustive_check.sh splices an arm into, the ---MARK--- line live_soundness_check.sh drops from psql, and rv_exec_check's documented "an RV binary never prints the program's own final value" — which is still load-bearing: run against the emulator, 19 of 96 programs match only after it._
_scripts/trailing_trim_check.sh keeps the class out, with two poisons. ⚠ It found three violations older than itself the first time it ran, in a file six hand-written greps had missed. ⚠ And its own first pattern used \? in a basic regexp, which BSD grep does not take, so on macOS it matched nothing and reported ok — the poison is the only reason that was caught._
v0.1.501 — 2026-09-21
_The workarounds v0.1.494 left behind, and the sixth place that decides what a program prints._
_CI had been red for seven commits. Every run failed the same eleven steps, and the first red one was v0.1.494 — Q-136, "a program whose value is unit prints nothing". That commit said the rule lived in five places and changed all five. It was wrong twice._
_Wrong about the harnesses. Six gates carried sed '$d' on the subject's output: until v0.1.494 every Mere program printed a trailing line for its own value, and these dropped it. Afterwards they dropped the answer. The clearest case is proto_parity, whose whole-message check reported_
FAIL message (interp)
protoc: 08ac0210ffffffffffffffffff01180322026869380142040102ac02
ours:
_— not a wrong encoding, an erased one. proto_gen_parity was worse: the same trim ran on the GENERATED SOURCE, so hello_pb.mere came out a line short of itself. render_agreement lost the last element of the server's tree and reported the client as having one extra._
_The probes that were still green show what the shape was: they end with a bare 0, a sentinel whose only job was to give the trim something to eat. Sentinel and trim are both gone now — the probes print exactly what is compared._
_Wrong about the count. There is a sixth place, and selfhost_check is what found it: the Mere compiler written in Mere still wrapped every program as let _ = print (show (main : main_ty)) in (). All seven of its equivalence cases were one line apart from the reference for seven commits._
_The rest were stale records — () sitting in test/boundary/EXPECTED, four test/vclock/*.expected, and the transcripts inside audio_check.sh and io_poll_check.sh._
_⚠ None of this was Linux-specific. All eleven reproduce on a development machine; they were simply not in the set anyone was running before pushing. Running the CI gate list locally — all 100 of them — is what this slice actually cost, and it found one more thing: width_check builds a Ruby comparison that read EastAsianWidth.txt in the caller's locale, so it died on a shell with LANG unset while CI, which sets UTF-8, never saw it. The encoding is stated at the read now._
_⚠ Six gates still carry a sentinel-and-trim pair. They are green and correct — their sentinel is a non-unit value that really is printed — and they are left alone, but a harness that drops its subject's last line can compare the wrong thing and pass, which is why the six that broke were rewritten rather than re-sentinelled._
v0.1.500 — 2026-09-21
_Floats stop being boxed at every node on the Wasm backend._
_Q-140 and v0.1.499 both came from asking the examples corpus a question nothing else asked. Sweeping it a third time — this time for allocation, not output — found examples/mandelbrot.mere dying with "out of memory" on Wasm and nowhere else: C reports alloc_total=6,291,736, Wasm used up the whole 64 MB linear memory._
_Isolating it took one probe and one control. 100,000 iterations of acc + (1.0 / (1.0 + (2.0 * 3.0))):_
Wasm C
float 6,400,036 B 32 B <- 64 bytes an iteration
int 30 B 0 B <- the control: the loop itself is free
_The control is what makes this a measurement rather than a guess: it is not the loop and it is not the Vec, it is the float. This backend keeps every value in an i64 slot, so a float is the address of an 8-byte box, and until now one was minted at every node — four literals and four results, eight boxes, to compute one accumulator._
_Three places the box is not needed, in the order they were removed:_
| probe | |
|---|---|
| before | 6,400,036 B |
| literals interned in the data segment (they are constants) | 3,200,024 B |
intermediates built on Wasm's f64 operand stack | 800,024 B |
let-bound floats in f64 locals | — (this probe has no let) |
_8×, and the 8 bytes left are the accumulator itself: the floor for a boxed value model. pi and e were allocating on every read too, and go the same way._
_No new mechanism — Q-109 had already built this shape for SIMD, so emit_float_f64 / float_locals / Ast.float_operand_only are emit_simd_v / simd_locals / Ast.simd_operand_only with the width changed._
_⚠ The predicate decides whether the unboxing pays, never whether the code is correct. An f64 local asked for its address gets boxed on the spot by its own arm in emit_expr. Resting correctness on an analysis that has to be conservative is how an analysis that is slightly wrong becomes a wrong answer._
_⚠ Nothing else in the tree could have caught this, and nothing else guards it now. The answers are bit-identical either way: examples/raytrace.mere writes the same 56,610-byte PPM and the same adler=48497-22482 before and after. The suite, parity, main_value_check and the size budgets are all green on the slow version. So the gate is the allocation meter — scripts/alloc_meter_check.sh check 6, with a ceiling and a floor: the ceiling catches boxing coming back, the floor says floats are still boxed at all, so the day that stops being true somebody has to come here and write it down._
_raytrace.wasm, where every vector is a (float, float, float), went 9,530 → 6,107 B (-36%) and broke its band's floor — which is the floor doing its job. Lowered deliberately, after checking the module still answers what the interpreter does, byte for byte._
_mandelbrot still does not run, and measuring why produced a different question. Reduced to 80×60 so it completes (the alloc_total of a program that hits the wall is how much fit, not how much it wanted), it went 22,133,389 → 13,702,001 B — 1.6×, not 8×. The same recursive loop at top level costs 16 B/iteration; nested inside a function that captures, 103 B/iteration, with an int control at 1.67. Counting bump sites per function in the WAT, the nested build alone carries $anon_1_fn and $anon_0_fn_fn2; the lifted body is identical. That is Q-139's __direct/fn2 not reaching a capturing inner function, and it is not a float question — filed separately._
_What remains genuinely float-shaped is the calling convention: an argument is a Mere value, so a float argument is an address, and two of them cost 16 bytes a call. Changing that is a different slice._
v0.1.499 — 2026-09-21
_The examples corpus is compared across backends, and the first run found a bug._
_Q-140 closed one un-gated path — what a program prints when it ends. The question it raised was how many others there are, and guessing at source was not going to answer it: the four backends hold 446 catch-all arms that produce a value rather than refusing._
_So the corpus was asked instead. examples/ holds 291 programs written for people, and nothing compared them across backends — check_cmd_check.sh sweeps the same tree but asks only whether a program is ACCEPTED, never what it prints. Running them on the interpreter and on C and diffing found, on the first pass:_
examples/vec_higher_order.mere
xs: Vec[1, 2, 3, 4, 5] (interpreter)
xs: <unknown> (C backend)
_`show` of a `Vec`, wrong on a shipped example, for as long as both had existed. And the same four-way split as Q-140 underneath it: interpreter Vec[1, 2], C <unknown>, LLVM a refusal, Wasm a function whose whole body was `(unreachable)` — a program that compiled, ran, and died with no message._
_C displays a Vec now, walking the elements and asking each one's own show, the way the list arm a few cases above it already did. ⚠ The arm had to go BEFORE the general `TyCon` one, which was swallowing it — OCaml's "this match case is unused" said so, which is the same check this compiler grew for Mere in v0.1.487._
_Wasm refuses instead of trapping. A type with no constructors is not a variant, and the tag chain collapses to a bare (unreachable); saying so at compile time is what LLVM already did and beats dying silently at run time. So the four now read: display, display, refuse, refuse — and none of them lies._
_scripts/examples_parity.sh is the gate. It selects by asking rather than by listing: an example is compared when the interpreter finishes it unattended AND its output is a function of the program. ⚠ That second question is asked after the C build, not beside the first run — back to back, two interpreter runs land in the same second and a per-second clock looks stable, which is exactly how log_levels_demo.mere slipped through the first version. The gap is the build, so nothing here is a fixed sleep guessing at the machine._
_186 compared, 89 skipped (input, ports, time), 1 not a function of the program, 15 refused by the C backend. The gate fails if fewer than 150 are compared: a corpus that quietly stops being swept passes forever._
_suite 2773, parity 196 PASS / 0 FAIL._
v0.1.498 — 2026-09-21
_Q-140: one rule for what a program prints when it ends._
_The same program, four answers:_
| main | interp | C | LLVM | Wasm |
|---|---|---|---|---|
true | true | 1 | 1 | true |
"abc" | "abc" | abc | abc | "abc" |
(1, 2) | (1, 2) | 1 | refused | (nothing) |
[1, 2] | [1, 2] | -872415215 | refused | (nothing) |
_That last cell is a POINTER printed as a decimal integer: main_format_of's catch-all was %d, so anything the table did not name became an int. The tuple row is the same thing printing its first field. Three backends had three different wrong answers and nothing could see it, because `parity.sh` runs programs that end with `print` — the path that displays the program's own value is one almost nothing exercises._
_`show` already agreed with the interpreter on every one of them. It was simply not on this path. So the rule is one line now, on all four: the program's value is displayed the way show displays it, which is what Eval.to_string has always done. Nine programs, four backends, one answer._
_⚠ A main whose type never resolved has no value to display. The common case is a program ending in exit 0, whose type is a variable because the call does not return — asking show for a 'a is a refusal, and it took down the whole suite the first time. The interpreter exits inside the evaluation and prints nothing, so printing nothing is what agrees with it._
_Seven assertions moved. Every one pinned a lowering (printf("%lld\n", 42LL), zext i1, @printf(ptr @.fmt_s, ...)) rather than an answer, and each was re-pinned at the new lowering with its question intact — except the zext, whose question stopped existing when printf did: the value now crosses to show_bool as an i1, so there is nothing to widen, and what is asked instead is that it crosses at its own width._
_scripts/main_value_check.sh is the gate, with a poison that rewrites one backend's answer and requires the comparison to catch it. It is in CI, because this is precisely the kind of drift that survives for years when no gate runs the shape._
_suite 2773, parity 196 PASS / 0 FAIL._
v0.1.497 — 2026-09-21
_Q-139 closed: every cell of the table is zero, on all three backends._
| B/call | C | LLVM | Wasm |
|---|---|---|---|
f a b (named top-level, concrete) | 0 | 0 | 0 |
f a b (f is a function parameter) | 0 | 0 | 0 |
| trait method | 0 | 0 | 0 |
vec_fold / map_iter / vec_map | 0 | 0 | 0 |
vec_sort, per element | 0 | 0 | 0 |
_Two rows were left, and NEITHER was the closure mechanism any more. Reading each one before touching it is the whole of why this is one slice._
_A capture-free top-level fn used as a value is a CONSTANT. ap add acc i in a loop built the same twelve-byte record every iteration — env 0 and two table indices known at emit time. It goes in the data segment once now, and the expression is a constant. Twelve bytes a call to nothing._
_And the last eight bytes an element were the merge sort's scratch, not a closure at all: n slots taken from the bump heap and left there. The C backend malloc/frees its scratch and its comment says exactly why — "a region alloc would leave n slots behind on every call" — and the Wasm backend had no such escape, so it simply left them._
_⚠ It is only given back when nothing was allocated after it. A comparator can allocate: the two-step path builds an environment, and a user's comparator can push to a Vec it captured, whose buffer may have been reallocated above the scratch. Resetting the bump past any of that hands out memory something still holds. The test is the honest one — the scratch is reclaimable exactly when the bump is still where the sort left it — and a probe whose comparator pushes on every comparison answers the same on all four backends._
_⚠ `$src` and `$dst` are swapped once per pass, and at the end $dst has been pointed at the vec's own buffer for the copy-back, so neither still names the block that was allocated. The first version compared against $dst and reclaimed nothing at all, silently._
_Q-139 took six slices: three on LLVM (v0.1.491-493) and three on Wasm (v0.1.495-497). It started because the allocation meter (Q-138, v0.1.488) made the question askable, and the first answer it gave was that Q-135's fix had been C-only — which nothing could have known before._
_suite 2773, parity 196 PASS / 0 FAIL, wasm 558 KB against 616 KB before the arc, every gate green._
v0.1.496 — 2026-09-21
_Q-139 on Wasm, second half: a closure value carries the uncurried entry._
| B/call | C | LLVM | Wasm v0.1.494 | v0.1.495 | now |
|---|---|---|---|---|---|
f a b (named top-level, concrete) | 0 | 0 | 32 | 0 | 0 |
f a b (f is a function parameter) | 0 | 0 | 67 | 12 | 12 |
| trait method | 0 | 0 | 32 | 20 | 0 |
vec_fold (2-arg callback) | 0 | 0 | 16 | 20 | 0 |
map_iter (2-arg callback) | 0 | 0 | 15 | 15 | 0 |
vec_sort, per element | 0 | 0 | 230 | 286 | 8 |
_Three pieces, the same three the LLVM backend needed: a two-argument adapter for an anonymous lambda whose body is immediately another fn (registered in the table beside the one-argument one, so the value and the adapter are decided by the same test); the fn2 index written into the closure record; and the branch in each of the three higher-order helpers — vec_fold, map_iter (BOTH variants, the linear one and the tombstoned one) and vec_sort._
_The comparator in a merge sort is asked n log n times — 1.7 million for the 100,000 the board measures — and every ask built an environment to carry the first element: 28.6 MB to 1.85 MB. The order and the count are unchanged, and test/parity/vec_sort_stable and closure_fn2_order both say so on this backend now as well as on the others._
_The trait row came free with the anonymous adapter: a dictionary field holds a lambda, so once lambdas carry fn2 the dictionary does, and the saturated closure-call path already admitted a field read as a head._
_Size: 522 KB to 558 KB — the adapters and branches are real code — against 616 KB before the arc. The bands are set to this settled state, measured once rather than chased slice by slice._
_What is left of Q-139 is two rows at 12 and 8 bytes, both on Wasm, and neither is the closure-value mechanism any more: something else allocates on those paths. The table is the gate, and it now reads zero in nineteen of twenty-four cells._
_suite 2773, parity 196 PASS / 0 FAIL, size and budget gates green._
v0.1.495 — 2026-09-21
_Q-139 on Wasm: the uncurried entry, and the shipped modules get 15% smaller._
_The third backend gets the mechanism the other two have. On Wasm it also turned out to be a size question, and the answer was the opposite of what the first attempt measured._
| B/call | C | LLVM | Wasm before | Wasm now |
|---|---|---|---|---|
f a b (named top-level, concrete) | 0 | 0 | 32 | 0 |
f a b (f is a function parameter) | 0 | 0 | 67 | 12 |
| trait method | 0 | 0 | 32 | 20 |
_The closure record is twelve bytes now, { i32 env, i32 fn_idx, i32 fn2 }, with the uncurried entry LAST so the host is untouched — scripts/mere_host.js reads env at +0 and fn_idx at +4 and neither moved. The absent value is -1, not 0: fn2 is a table index and index 0 is a real function._
_⚠ Six hand-written closure records in the runtime WAT kept the old width, and the eight bytes after them were read as a table index and jumped through. test/parity/logger_metrics_caps found it: the program printed three log lines and then "out of memory". Widening a record means every construction, including the ones written by hand in a string._
_And then the size gate refused it. The twins took the shipped modules from 616 KB to 769 KB, six of fifteen over their ceilings. The cause was not the twins: a `_closure` adapter was emitted for every top-level fn, the elem table it sits in is a root the pruner cannot see past, and so every curried body shipped whether or not anything could reach it. Emit the adapter only for fns actually used as values, and the curried body only when something can still enter it — a fn with a twin, never used as a value and never applied to fewer arguments than the twin takes, has no way in — and the pruner does the rest._
_Result: 522 KB, down from 616 KB before any of this. gameboy −29%, raytrace −26%, chip8 −22%, selfhost-tyck −19%. Every band in scripts/wasm_size_budget.txt and test/budget/BUDGETS moved DOWN, which is the floor half of those gates doing exactly its job: "a demo this much smaller than recorded probably stopped being built properly; if it is a real improvement, lower the band". It was checked before the bands moved — suite 2773, parity 196/0, the bootstrap tests that RUN the self-hosted compiler under node, and the size gate's own behaviour check._
_⚠ Two things had to be excluded, and both were found by breakage, not by reading. A monomorphised instance gets no twin: its calls arrive under a mangled name the collectors cannot see, so giving it one dropped a curried body that the multi-instance dispatch still called, and wat2wasm named the missing function. And the saturated closure-call path only takes a head that is already a VALUE — a local or a field — because a partially applied multi-instantiated fn has no value form on this backend, and asking for one is a refusal that surfaced as the self-hosting bootstrap failing to emit._
_⚠ A measurement taken while the tree was being edited is not a measurement. The first verification of this ran in the background while the next slice was still landing, and reported three parity failures that did not exist. Re-run on a settled tree: 2773 / 196 / 0._
_What remains of Q-139 is the Wasm closure-value half — fn2 is carried by top-level fn values but not yet by anonymous lambdas or dictionary fields, so the callback rows still allocate. The table is the gate._
v0.1.494 — 2026-09-21
_Q-136: a program whose value is unit prints nothing._
_() was printed by all four backends from Phase 25.11 / 27.0, to keep the interpreter and the compiled backends saying the same thing. They still say the same thing; they say nothing._
_What it cost. Anything comparing output exactly — a judge, a diff, a golden file — saw an unconditional mismatch on every program whose main is unit, which is every program that ends by printing its answer. All twelve solutions on the mjudge board carry let _ = exit 0; as their last line for no other reason, and one of them carries a comment explaining why. Everybody writing the same line to switch a default off is the evidence that the default points the wrong way._
_Five places decide this and all five changed, which is why it is one slice: Pipeline for the interpreter, main_format_of in the C and LLVM backends, the TyUnit arm of the Wasm epilogue, and the LLVM format global that fed the call. Checked end to end on all four: print "answer" now writes answer\n and nothing else, byte for byte, on interp / C / LLVM / Wasm._
_⚠ `Eval.to_string` is NOT where this changed. show () goes through it and has to keep answering (). The decision belongs to "what does a program print when it ends", so it sits at the end of Pipeline.process — which now has an option-shaped twin, process_opt, because "nothing to print" and "print an empty line" are different answers and a str main whose value is "" still gets its newline. process stays a string for its six hundred callers; one rule, two spellings._
_Four golden files and twenty assertions moved with it. The two that pinned the OLD behaviour were re-pinned at the new one rather than deleted — the question "does a unit main print anything" is still worth asking, and now it has the other answer._
_The `()` was there for a parity that does not actually hold, which came out while measuring this: a str-typed main prints "abc" — with the quotes — on the interpreter and abc on the compiled backends, and no gate notices, because parity programs end with print rather than with a bare value. Recorded as its own question rather than fixed here._
_suite 2773, parity 196 PASS / 0 FAIL, ten gates green._
v0.1.493 — 2026-09-21
_Q-139, third slice: the LLVM backend allocates nothing where the C backend allocates nothing._
| B/call | C | LLVM v0.1.490 | v0.1.492 | v0.1.493 |
|---|---|---|---|---|
f a b (named top-level, concrete) | 0 | 16 | 0 | 0 |
f a b (f is a function parameter) | 0 | 56 | 0 | 0 |
| trait method | 0 | 16 | 8 | 0 |
vec_fold (2-arg callback) | 0 | 8 | 8 | 0 |
map_iter (2-arg callback) | 0 | 8 | 8 | 0 |
vec_sort, per element | 0 | 111 | 111 | 0 |
_The three rows had three different causes, and reading each one before touching it is the only reason this took one slice instead of three._
_The builtin higher-order helpers are emitted IR text, and their loops did the two-step call — vec_fold, map_iter and vec_sort, which is the whole surface: three sites. A merge sort asks its comparator n log n times, 1.7 million for the 100,000 the board measures, and every ask built an environment to carry the first element._
_Then the numbers did not move, because the callback in vec_fold v 0 (fn a -> fn b -> a + b) is an ANONYMOUS LAMBDA and those carried null. The C backend added __anon_N_fn2 for this in v0.1.482; this adds the LLVM equivalent. The condition is "the body is immediately another fn", which is the same question the order pin asks: if anything happens before the inner fn is returned, that work belongs between the two arguments and an entry taking both at once cannot express it._
_And the trait row was not about the value at all. The dictionary already carried fn2 — @anon_23_fn_fn2 sat in field 2 of @mu_Ord2__int__dict — and nothing was willing to look, because the call-site guard admitted only a Var head and a trait method is exactly what a Var is not: it elaborates to a read from the instance's dictionary. A field read is always a value and never a builtin, so those heads are admitted now. One row that looked like the other two was a missing line in a guard._
_suite 2773, parity 196 PASS / 0 FAIL, twelve gates green. Wasm is what remains of Q-139: it has none of this, and its column is unchanged at 32 / 67 / 230._
v0.1.492 — 2026-09-21
_Q-139, second slice: the uncurried entry becomes N-ary, and a closure value carries one._
_Two changes, and the second was measured into existence rather than assumed._
_N-ary, not exactly two. The first slice took exactly two parameters, because that is where the meter had pointed. ap add acc i is a THREE-argument spine, so its first application still built an environment per iteration — the whole of that row. Opening the twin to two or more took it from 56 to 8 bytes a call._
_And a closure value now carries `fn2`, the uncurried entry, the way the C backend's does since Q-135. The layout is { ptr env, ptr fn, ptr fn2 }, and null is a real case, not a defensive one: a partial application, a polymorphic value, an eta adapter and anything with a region parameter all carry null and take the two-step path, which is the path every closure call took before this. So the choice is a branch at runtime, not a decision at emit time. Every construction starts from zeroinitializer rather than undef for that reason — an undef field here is a pointer that would be CALLED._
_Worth its width, and checked by removing it: with the fn2 path disabled and everything else in place, the same row reads 8; with it, 0._
| B/call | C | LLVM v0.1.490 | v0.1.491 | v0.1.492 |
|---|---|---|---|---|
f a b (named top-level, concrete) | 0 | 16 | 0 | 0 |
f a b (f is a function parameter) | 0 | 56 | 48 | 0 |
| trait method | 0 | 16 | 8 | 8 |
vec_fold / map_iter | 0 | 8 | 8 | 8 |
_⚠ Two things went wrong, and the tree caught both._
_The arm took over calls that were not the user's. emit_user_app is also the fallback for shapes the builtin arms declined, so a partially applied BUILTIN — channel_recv_timeout c ms — came through it and produced a different refusal than the one the suite pins. The arm was right about the shape and wrong about whose call it was; it now requires a user-bound head._
_And it changed evaluation order. The second argument was evaluated before the branch, but on the two-step path the FIRST application runs first, and a body that is not immediately another fn can print before the second argument is reached. test/parity/closure_fn2_order.mere holds noisy 1 (loud 2), whose prints must read "mid" then "arg", and this printed them the other way round. That file was written for the C backend's own version of this change and it caught this one on its first run. The second argument is now evaluated inside each branch._
_suite 2773, parity 196 PASS / 0 FAIL, ten gates green._
v0.1.491 — 2026-09-21
_Q-139, first slice: the uncurried entry reaches the LLVM backend._
_`__direct` existed in `codegen_c.ml` and nowhere else, and nothing could see that until this backend had an allocation meter (v0.1.488). What the meter said, for a named top-level function at a concrete type applied to both its arguments — the row that is zero on C by construction:_
B/call C LLVM Wasm
f a b (named top-level, concrete) 0 16 32
_A two-argument call compiled here to "allocate an environment holding the first argument, return a closure, apply it to the second". So on this backend it was never abstraction that allocated; it was the call._
_Same shape as the C table: for an eligible f = fn p1 -> fn p2 -> body, emit @mu_f__direct(p1, p2) holding the body, and send exactly-saturated call sites straight to it. The curried definition is still emitted beside it — this adds an entry rather than replacing one, so partial application and first-class use are untouched._
_Conservative on purpose for a first slice: exactly two parameters, no region parameters, concrete throughout, not an inner lift, head not shadowed by a local, arity matching exactly._
_⚠ Rerouting a call also reroutes it away from `musttail`. The first version emitted a bare call, and the two v0.1.451 assertions went red immediately: they ask a two-argument let rec for musttail call %w and got a bare call. That is not a cosmetic loss — a tail-recursive function without the guarantee grows the stack once per iteration, which is the defect the C backend records against its own single-argument case. The twin now publishes its own prototype (two i64 arguments, not the curried (ptr, T)), which is what lets the self tail call inside it match and become musttail; wide aggregate returns still take notail, because tail on an sret call is the Q-129 miscompile. A two-argument tail-recursive function had no constant-space form on this backend before this._
| B/call | C | LLVM before | LLVM after |
|---|---|---|---|
f a b (named top-level, concrete) | 0 | 16 | 0 |
| trait method | 0 | 16 | 8 |
f a b (f is a function parameter) | 0 | 56 | 48 |
_The rest of Q-139 is the closure-value half — the equivalent of Q-135's fn2 slices — and the Wasm backend, which has none of this yet. Both stay open, with the table as the gate._
_suite 2773, parity 196 PASS / 0 FAIL, musttail_budget / debug_info / determinism / stack_overflow / vectorize green._
v0.1.490 — 2026-09-21
_The Wasm meter was silent for every program that ends by exiting._
_`exit n` reaches the Wasm host as `exit_proc`, which calls `process.exit` and never comes back, so the line that reported allocation was never reached. Not an approximation and not a wrong number: no output at all, on exactly the shape every judged program has — exit 0 at the end is how the trailing () is suppressed (Q-136). The meter was checked on programs that fall off the end and those all worked._
_Reported from the exit event instead, which fires for process.exit as well as for a normal return. The LLVM backend reached the same conclusion two days earlier with atexit; Wasm was the one that did not have it._ _test/allocmeter/exits.mere pins it, and the gate compares the number against the same work written without the exit — silence would otherwise read as a program that allocates nothing._
_And then the meter was used for what it was built for. Q-138 existed so that Q-135 — 24 bytes per saturated application, fixed for C in v0.1.481-482 — could be asked of the other backends. The answer, measured rather than assumed:_
| B/call | C | LLVM | Wasm |
|---|---|---|---|
f a b (named top-level, concrete) | 0 | 16 | 32 |
f a b (f is a function parameter) | 0 | 56 | 67 |
| trait method | 0 | 16 | 32 |
vec_fold (2-arg callback) | 0 | 8 | 16 |
vec_map (1-arg callback) | 0 | 0 | 0 |
map_iter (2-arg callback) | 0 | 8 | 15 |
vec_sort, per element | 0 | 111 | 230 |
_The fix was C-only, which the changelog for v0.1.481-482 could not have said at the time because nothing could measure it. And the first row says the problem there is wider than Q-135: that row is zero on C by construction — a named top-level function at a concrete type, applied saturated, written out at the call site — and it costs 16 and 32 bytes on the others. On those backends it is not abstraction that allocates; it is the two-argument call._
_No change to those backends here. The instrument now exists and has been read; what it found is written down in the project's open questions, and fixing it is its own slice._
v0.1.489 — 2026-09-20
_The compiled backtrace is withdrawn: it answered on macOS and said nothing on Linux._
_v0.1.486 printed the frames under a compiled failure itself, reading them back from backtrace() + dladdr. Every gate was green — on macOS. CI is where it broke, and the interpreter half (five cases, position and frames) passed there untouched:_
runtime_loc: 5/5 ok (position + frames on every case)
compiled frames: got [] want [a__direct b__direct c__direct d__direct]
_`dladdr` consults the DYNAMIC symbol table on glibc, and the frames a Mere program actually stands on are its __direct twins, which are emitted static. A static function has no entry in .dynsym — -rdynamic exports globals and does not change that. Measured in the CI image with a four-frame probe: plain link resolves nothing, -rdynamic resolves mu_a, mu_b and main, and mu_a__direct resolves under neither. So the design could not have worked there; it was not a flag away._
_The retraction rather than a platform carve-out: a diagnostic that answers on one platform and is silent on the other is worse than one that is silent on both, because the silence reads as "nothing to say" instead of "this does not work here". The compiled leg of the gate now pins the SAMENESS — stderr is the message and one newline, byte-identical at -O0 and -O2 — and MERE_FAIL_TRAP=1 is the way to the frames, handed to a debugger whose unwinder reads DWARF and can see a static function._
_The interpreter side is untouched and is where this mattered: the position and the call stack, which is what a 30,000-line file needed._
_⚠ The gate was run on macOS only before it was pushed. The repo already knows to run a new gate once on the CI image — ocaml/opam:ubuntu-24.04 builds mere.exe in about 33 seconds — and that step was skipped. "Tested" is true about the machine it was tested on._
_suite 2773, parity 178, and the three gates plus their poisons green on both macOS and the CI image._
v0.1.488 — 2026-09-20
_Q-138: the allocation meter, on the other two backends._
_`MERE_REGION_STATS` existed in `codegen_c.ml` and nowhere else, so "this change cut allocation" was a sentence only one backend could be asked about. Q-135 — 24 bytes per saturated application, v0.1.481-482 — was verified on C alone for exactly that reason, and the other three were taken on trust._
_Three implementations now answer, and they agree to within 18 bytes. The same program (a 20,000-element Vec, no region block):_
| backend | alloc_total | where it counts |
|---|---|---|
| C | 262,176 | a field on the region struct (already there) |
| LLVM | 262,184 | @__lang_region_alloc's use AND growdone, plus in-place growth |
| Wasm | 262,194 | the bump pointer, plus $__lang_reclaimed |
_Wasm counts on the way DOWN. $__lang_bump goes up at an allocation — twenty-odd sites — and down at exactly one, the region release. Instrumenting the one place that goes down is a ninth of the work and avoids writing one rule in twenty places. The total handed out is the bump plus everything a release took back, which also means the bump alone is a lower bound — the reason the second number has to exist._
_And that second number is one the C backend does not have: its regions free whole blocks and never report how much was in them. On the region version of the same program, 99% of 524,306 B goes back._
_⚠ The first number this meter produced was 34 bytes, for a program that allocates 262,194. The release subtracted from mark — where the block STARTED — instead of from the bump as it stood. mark is below the release target, so the difference clamped to zero every time. One name, two questions. It was visible as wrong only because C had already answered for the same program, which is the argument for building the third meter rather than trusting the first._
_LLVM needed both exits of its one allocation function, as the investigation predicted: instrumenting only the path that reads like the main one would undercount exactly the programs that allocate enough to need a second block. In-place growth moves the top pointer without going through the allocator and needs its own line. getenv moved to the unconditional declarations, because env_var_runtime_llvm is emitted only when the program calls env_var and two conditional declarations of one symbol in one module is an IR error — the only program that would have shown it reads an environment variable AND is measured._
_Opt-in through MERE_REGION_STATS, one line on stderr, stdout byte-identical with it on or off. The cost on Wasm is hello.wasm 1,946 → 1,970 B (+24 B, +1.2%, against a band of 1,751-2,141). RV32I stays out of scope: --bare has no stderr._
_scripts/alloc_meter_check.sh is the gate — three-way agreement, the identity alloc_total = live + reclaimed, silence by default measured in BYTES, and twice the work reporting twice the bytes (a meter stuck on a constant passes everything else). Two poisons: the agreement band pointed at the pair that legitimately differs by 2x, and the doubling band pointed at one program against itself._
_parity 178, suite 2773, wasm_size_check 15 files all in band._
v0.1.487 — 2026-09-20
_Which arms no value reaches._
_The other side of the question this compiler has asked since Phase 1. Exhaustiveness says which values no arm answers for. Nothing said which arms no value reaches, and a duplicated constructor arm went through mere, mere check and mere check -c in silence:_
| Rock -> "r" | Paper -> "p" | Rock -> "r2" | _ -> "x" | Scissors -> "s"
_— accepted, exit 0, nothing printed. Two arms there are dead. The defect this answers is a recorded one: when a second arm for one constructor never runs, what the reader eventually sees is an error raised from the arm they did NOT write the code in, and there is nothing in the build to connect the two._
_Exhaustive.redundant_findings already had everything it needed — the arms in order, with their positions. An arm is reported when every alternative it offers was already closed by an earlier UNGUARDED arm, closed meaning an irrefutable pattern, a bare constructor, a constructor whose payload pattern is itself irrefutable, or a literal._
_Only when certain, because the cost of a false positive is a deleted live arm. Four shapes are left alone, and all four are asserted as silent in the suite and in the gate: a guarded arm above the same constructor (a guard can be false, so the arm below it is the one that answers), a constructor with a refutable payload (Link (0, _) is some Links, not all), an or-pattern only half of which is closed (Paper | Scissors under a Paper is still reached by Scissors), and the defensive _ written after every constructor is already named — dead, reportable, and deliberately not reported, because it is the shape people write so a match keeps compiling when the type gains a case, and the noise would bury the findings that matter._
_A warning, not an error. It is a check being added to trees that already compile. Escalating is a separate decision, to be made when the count is zero and stays there._
_It found nothing. 836 sources here, plus mere-ruby's main.mere (30,392 lines and seven modules), m3d and mbrowse: zero dead arms. That is a measured zero and not a blind one — a Nil arm duplicated inside m_strutil.mere, which main.mere reaches through an import, is reported with its file, line and column, so the check does run over a program that size and through its imports._
_scripts/unreachable_arm_check.sh is the gate: seven cases (three reported, four silent), then a sweep asserting the shipped tree stays clean, and a --poison that raises one case's expected count and moves another's expected LINE, requiring exactly one red each. The line half matters on its own — a check that warns about the wrong arm is the failure that costs the most, and a count alone cannot see it._
_suite 2773._
v0.1.486 — 2026-09-20
_A runtime failure says where it happened, and how the program got there._
_Static errors in this compiler have carried a position and a code frame for years; runtime ones carried neither. map_get: key not found in Map was the entire report. In mere-ruby's main.mere — 30,392 lines — that names nothing, and the first thing anyone did with it was go looking for which of the callers reached it._
_Three parts, and the materials for all three were already here._
_The position comes from the application node. The fifty-odd builtins that raise do it with Loc.dummy, because a builtin is handed no position; the App case lends the call site's one to any Eval_error that arrives without one. A builtin that called back into user code (vec_sort's comparator, map_iter's block) re-raises a failure that already knows a better line, and that one is left alone._
eval error: map_get: key not found in Map (use map_has to check first)
--> lookup.mere:7:29
|
6 | let m = map_new ();
7 | let lookup = fn (k: str) -> map_get m k;
| ^^^^^^^ map_get: key not found in Map
8 | let go = fn (n: int) -> lookup "missing";
call stack (innermost first):
lookup at lookup.mere:8:25
go at lookup.mere:9:8
_The frames are the interpreter's own call stack. Frames are popped on the way BACK and deliberately not on the way OUT, so an unwinding failure leaves them standing for the driver to read — the bargain call_depth right beside it already makes. The first version took a snapshot in an exception handler on every call instead: fib 32 ran 0.66 s with that handler and 0.62 s without, against 0.62 s for the build with none of this. The handler, not the frame and not the allocation, was the whole cost. A frame is the application node, not a (name, position) pair, so a call costs one cons cell and no walk down the spine for a name only a printed frame ever needs._
_⚠ The symmetric-looking version of that is wrong. Restoring call_depth per frame on the way out — the obvious counterpart to the success path — makes the suite fail on the spawn tests, because `call_depth` is one global shared by every domain. A child domain running its own closures writes the same counter, so a main-domain unwind that also writes it lands between the child's +1 and -1. Measured drift: 0 → 1 → 59, and it stuck there for every later program in the process._
_The compiled backend gets the frames from the machine, through dladdr. It works because a user's Mere function is emitted with external linkage (mu_lookup) while the prelude's helpers are static, so the filter is the program's own functions with no list to maintain._
_⚠ A return address is never a function's first instruction, so a pc that resolves to offset 0 is not a frame — it is a pc the walk could not place being handed the nearest symbol that starts there. Without that rule the report named a real function in the program that was not on the path to the failure. Measured on a four-deep chain: at -O0 one spurious +0 frame, and dropping it leaves exactly the chain; at -O2 every remaining frame is +0, and dropping them leaves none — which is the truth, because inlining and tail calls mean those frames are not on the stack to be found._
_MERE_FAIL_TRAP=1 raises SIGTRAP at an uncaught failure instead of exiting, so a debugger holds the program AT the failure with -g mapping it back to the .mere line. Opt-in, because without a debugger attached SIGTRAP is fatal and the exit status stops being 1. A caught fail is control flow and does none of this: silent, no frames, no trap, exit 0._
_MERE_BACKTRACE=0 turns the frames off on both backends; MERE_BACKTRACE_FRAMES caps them (default 10). Runaway recursion arrives with as many frames as the depth limit allows, so repeats are collapsed — down at f.mere:8:35 (x 39)._
_scripts/runtime_loc_check.sh is the gate, and it carries its own poison (--poison): it switches the frames off and requires every case that declares frames to go red, then moves one expected position by a column and requires exactly that case to go red, with the clean run asserted green first so that "red" means the poison and not a broken harness. Both run in CI. Five interpreter cases plus a compiled leg (frames at -O0, none at -O2, both switches, and a caught fail that stays silent with the trap armed)._
_parity 178 unchanged, suite 2767._
v0.1.485 — 2026-09-15
_str_of_int stops going through printf._
_`show_int` was `asprintf("%lld")` followed by `__lang_str_of_cstr` and a `free` -- the whole printf formatting machinery, a malloc, then a strlen, a REGION allocation and a memcpy, and finally a free: two allocations and a copy to spell a number. It writes the digits into a 24-byte stack buffer and makes one region allocation now._
_Measured in C alone, 2,000,000 conversions: asprintf 135-145 ms, digits by hand 52-54 ms. In mere-ruby, where anything that keys a table by an id calls it:_
| before | after | |
|---|---|---|
| CSV.parse quoted, 1,600 rows | 4.29-4.42 s | 3.33-3.40 s |
| 120k Array / Hash / String / Integer method rounds | 9.60 s | 9.18 s |
| 200k method calls | 4.23-4.29 s | 4.08-4.24 s |
_Three runs each for the first two, five for the last -- which is the honest one to read carefully: its ranges nearly touch, so call it 2-3% and not more. The CSV row is 23%._
_⚠ The input that breaks a hand-rolled itoa is the most negative integer, where -v overflows -- and signed overflow is UB the optimiser may delete a guard for. The form here computes (unsigned long long)(-(v + 1)) + 1 and never negates the value at all. Two tests pin it: the minimum, and the zero / negative / maximum trio._
_Checked on the interpreter, the C backend and LLVM -- eleven lines each, identical, edges included -- and the emitted C was compiled and run on the Linux CI image for the same answer. (LLVM has its own show_int through mint_show_format and is untouched; it agrees because it always did.)_
_2767 tests, parity 178/178._
v0.1.484 — 2026-09-14
_The underscore that was worth 100x._
_`try_or` compiles to `_setjmp` / `_longjmp` instead of `setjmp` / `longjmp`. On macOS and the BSDs the plain pair SAVES AND RESTORES THE SIGNAL MASK, and that is a sigprocmask syscall on every entry. Measured here with a three-line C program, 3,000,000 iterations:_
| time | per call | |
|---|---|---|
setjmp | 673-687 ms | 229 ns |
_setjmp | 6.7-7.0 ms | 2.3 ns |
_~100x, and it is paid on ENTRY -- not on the rare unwind. Anything that wraps a call in try_or pays it per call. In mere-ruby, an interpreter written in Mere, try_or is entered once per Ruby method call, and the plain pair was 7% of the whole profile (225 of 3,305 samples, with sigprocmask and __sigaltstack sitting under it)._
| before | after | |
|---|---|---|
| mere-ruby, 200k method calls | 5.20-5.31 s | 4.91-5.02 s |
_Three runs each, no overlap between the two sets._
_Nothing depends on the mask being restored. The runtime never blocks a signal, and no signal handler longjmps out -- the SIGSEGV handler writes one line and _exits. On glibc setjmp already does not save the mask, so this changes nothing on Linux; the emitted C was compiled and run on the CI image to confirm _setjmp is declared there too (same answer, clean build)._
_Both backends that emit it: the C one (4 sites) and the LLVM one (declare i32 @_setjmp(ptr) returns_twice, and the _longjmp call in the fail path). A try_or program gives the same answer on the interpreter, the C backend and LLVM._
_A test refuses the plain pair, because dropping the underscore is a 100x regression on one platform and invisible on the other -- exactly the kind of change no ordinary test notices. It emits C for a try_or program and fails if a bare setjmp( or longjmp( appears; poisoned by putting one back, and it went red._
_2765 tests, parity 178/178._
v0.1.483 — 2026-09-14
_A name can be promised before it is defined._
_`let fn <name>: <type>;` declares a top-level name ahead of its definition, so two functions can call each other without sharing one let rec ... and ... chain. The chain was the only tool for mutual recursion, and it is also the unit of MONOMORPHISM: everything inside one is instantiated at a single type. So in a large program "these two call each other" and "these two must have the same type" were one statement, with no way to say only the first._
let fn even: int -> bool;
let odd = fn (n: int) -> if n == 0 then false else even (n - 1);
let even = fn (n: int) -> if n == 0 then true else odd (n - 1);
_A declaration is checked, not trusted. The definition is unified against the declared type where it appears, and a promise still unkept at the end of the program is an error naming the declaration's line._
_mere --decls <file> prints the declarations for a program's own top-level names, so an existing chain can be cut without transcribing types by hand._
_The feature landed with four defects, and its seven unit tests were green for all four. Every one of them needed a real program to see, and what saw them was pasting --decls output back into the file it came from and requiring byte-identical output — now scripts/decls_roundtrip.sh, over the 178-program parity corpus. It found:_
| what it did | what saw it | |
|---|---|---|
| a promise on a name that shadows a builtin | never kept — let fn odd; + let odd = ... was refused as undefined | round-trip |
| a definition more specific than its promise | accepted; a caller above it could pass a type the definition has no body for. Typer passed, C backend failed to compile | round-trip |
| declaring a type | made the name LESS general than not declaring it: 'a -> 'a was fixed by its first call site | a two-call-site program |
--decls | printed the prelude's ~70 names, under post-uniquify spellings, and RAN the program while doing it | round-trip |
_The first is uniquify_toplevel_shadows, which renames a top-level binding that shadows a builtin and did not know about forward declarations: the promise kept the name odd and the definition became odd__v2, so the promise was registered under one name and kept under another. A declaration and its definition are one binding, so the rename happens at the declaration now and the definition inherits it._
_The second and third are one root: the declaration and the definition were related by INSTANTIATION rather than SUBSUMPTION. Unifying the definition against a fresh instance of the declared scheme lets the definition be narrower than the promise, and the instance's variables — made at the outer level — drag the definition's own variables down out of reach of generalize. Unifying against the declaration AS WRITTEN fixes both: the parser makes 'a a TyParam, which unifies only with itself or an unbound variable, and that is exactly a skolem. The declared scheme is then what the name means both above and below its definition, which is the whole point of writing one._
_A fifth thing the round-trip found is not a defect: a declaration for a name that shadows a builtin moves the shadow up to the declaration, so a caller written above the definition stops seeing the builtin. That is the feature working — test/parity/shadow_builtin.mere exists to hold exactly that ordering — but it is not what someone pasting a generated file expects, so --decls prints those lines commented, with the reason. Names bound twice at top level get the same treatment, for a plainer reason: one declaration cannot name two bindings._
_let fn also accepts a module-qualified name (let fn M.f: int -> int;). The lexer makes each . its own token and the parser matched a single identifier, so every module member was undeclarable — 625 of the declarations --decls printed for the corpus._
_167 of the 168 corpus programs round-trip. The one that does not names a record type declared inside a module, which cannot be named in an annotation from outside it at all — let use = fn (r: M.t) -> r.a fails the same way, with no declaration involved. It is exempted BY NAME in the gate, and the gate fails if it ever starts passing, so the exemption cannot outlive its reason._
_The feature came from a program that had outgrown the chain: an interpreter whose let rec eval_e = ... and ... had reached 38,856 lines in one file, because everything reachable from the evaluator had to live in it. Six generated declarations were enough to cut it. That program does not ship the cut — it peeled off everything the chain did not need instead, which is the better answer for that codebase — but the constraint it was working around is gone either way._
_16 tests (2764 total), one new gate. Documented in docs/language-reference.md._
v0.1.482 — 2026-09-14
_And the lambda written at the call site._
_An anonymous closure gets an uncurried twin too. v0.1.481 gave named fns and their values an fn2, which left vec_sort v (fn a -> fn b -> a - b) paying an environment per comparison: an anonymous closure had no __direct twin to point at. It has one now, and the generic call path reaches it through a dictionary field as well, so a trait method is no longer the expensive spelling of the same thing._
_Every row of the table this arc started from is zero:_
| spelling | before | after |
|---|---|---|
f a b, f a named top-level fn | 0 | 0 |
f a b, f taken as a PARAMETER | 24 B/call | 0 |
a trait method, cmp a b | 24 B/call | 0 |
vec_sort v (fn a -> fn b -> …) | 333.5 B/elem | 0 |
vec_sort v cmp | 333.5 B/elem | 0 |
vec_fold, map_iter | 24 B/call | 0 |
_The condition for peeling is the one peel_lifted_direct already used, and it is not a detail: the body must be IMMEDIATELY another fn. fn a -> { print a; fn b -> … } does work when it takes its first argument, a program can see whether that happened, and it keeps the two-step path. The cost of the twin is that a two-argument lambda's body is emitted twice._
_The generic path now also accepts a head that is a FIELD rather than a name. A name can be something the arms above call directly with no closure anywhere — that is what broke the first version of this guard — but a field read is already a value, and it is how a trait method arrives._
_`scripts/closure_alloc_pin.sh` pins both directions, and the floor is the one that matters: a partial application must still allocate. A gate that only checked that saturated calls are free would stay green the day the two-step fallback was deleted, and a callback with no fn2 would then call through a null pointer. It was poisoned before being believed — with fn2 disabled it reports 5,034,136 bytes and fails; with it, 8,224 and passes._
_New parity program closure_fn2_order.mere holds the three things a program can see across all four backends: partial application still returns a closure, work between the parameters still happens before the second argument is evaluated, and a comparator is called the same number of times in the same order._
_2748 tests, parity 178 + 18, ctest 18, contrib 79 of 87._
v0.1.481 — 2026-09-14
_The uncurried twin was already in the file; the value did not carry it._
_A closure whose return type is itself an arrow now carries `fn2`, an entry point taking both arguments at once. Applying a curried closure to two arguments used to build the intermediate closure's environment just to pass the second one: 24 bytes and about 13 ns, on every spelling that abstracts over the function. On a 200,000-point sort that was 1,855,936 environments and 60.4% of everything the program allocated._
_What the search turned up is that half the fix was already emitted. A named top-level fn gets a __direct twin, and for the sort in question mu_cmp__direct(P, P) and mu_cmp_as_value were in the same file — the value simply did not point at the twin. Same shape as Q-066 (v0.1.325), where the twin had been emitted for years and the call site looked it up under a name it was never keyed by._
| before | after | |
|---|---|---|
vec_sort v cmp, n = 100,000 | 33,349,392 B | 0 |
| a function taken as a parameter, 1e6 calls | 24,000,000 B | 0 |
| that sort as a judge problem: allocation | 134.8 MB | 39.2 MB |
| …peak RSS | 140.6 MB | 45.1 MB |
_fn2 is NULL wherever a twin cannot be named — a polymorphic callee, a twin that takes leading region parameters, an anonymous lambda (which has no twin at all yet) — and the two-step path is still emitted right below it. That is the fallback, and both codegen assertions were written to keep BOTH paths in the text: a test that only saw the fast path would go quiet the day the fallback was dropped._
_Two things had to be narrowed after they broke something. The first guard checked only the TYPE of the call's head, so an inner-lifted fn — which the arm below calls directly, with no closure anywhere — was turned into a compound literal of a closure typedef the program had never needed and therefore never emitted. The C did not compile. scripts/parity.sh caught it, by failing to BUILD rather than by answering wrong, which is the cheaper of the two ways to find out. The head must be a local holding a closure._
_The second is evaluation order. A curried closure with no fn2 may do work when it takes its first argument, and the program can see the order, so the partial application still happens before the second argument is evaluated. The emitted shape is awkward on purpose: the inner closure is declared without evaluating the call (__typeof__), assigned only on the slow path, and the second argument appears once. fn2 exists only for a callee that does nothing between its two parameters, so skipping the step there is unobservable._
_Only arrow-returning closure structs grow, by one pointer: unit -> unit (spawn's) and str -> unit (a Logger field) are untouched. 2748 tests, parity 177 + 18, ctest 18._
v0.1.480 — 2026-09-14
_The output was already assembled, and building it bought nothing._
_`print_no_nl` and `print_bytes` hand the whole buffer to one `write(2)` instead of fwrite-ing it onto stdio. main sets stdout line buffered, which is right for a program someone is watching -- a server's log should not stop existing because it was redirected -- and wrong for one that has already finished its answer: an fwrite of a 15 MB buffer onto an _IOLBF stream flushes at every newline. So the idiom the docs recommend, accumulate into a StrBuf and print once, still paid one syscall per line._
_Measured against a judge's test data, twelve programs, worst case each:_
| before | after | C++ reference | |
|---|---|---|---|
| 1,000,000 lines out | 2.40 s | 0.23 s | 0.37 s |
| 500,000 | 1.25 s | 0.10 s | 0.16 s |
| 500,000, graph | 1.31 s | 0.18 s | 0.30 s |
| 1 line out | 0.08 s | 0.10 s | 0.05 s |
_The constant that disappeared was about 1.6 microseconds per output line, and it was the same number in all ten programs that print more than one line -- 1.50 to 1.99 µs, median 1.6. The two whose answer is a single line did not move, which is how the attribution was confirmed rather than assumed: a fix that also changed the control would have been measuring something else._
_fflush(stdout) leads, because print still goes through stdio and both share fd 1: the order has to be the program's, not the buffer's. A short write is a loop and EINTR is a retry -- dropping the tail of an answer is the one failure mode worth writing three lines to prevent. print is unchanged and still writes a line at a time, which is what a log wants._
_The two codegen assertions this moved were updated to assert the same property they always did: that the LENGTH crosses the boundary, so a zero byte in the middle of a bytes does not end the output._
v0.1.479 — 2026-09-13
_A 13x search, found by pointing a benchmark at an editor rather than at the compiler._
_`str_index_of` compared the needle at every offset. For a needle that is not in the haystack that is one memcmp call per byte of haystack, which is the case a user waits through: "search the whole file and find nothing". memchr finds the candidate first bytes and memcmp only confirms them, which hands the scanning to the one libc function that is certainly vectorised. Measured over a 1 GB file (medit2/bench/big.sh, which is what asked):_
| before | 548 MB/s — 1,868 ms for 1 GB |
| after | 7,211-7,428 MB/s — 140 ms |
_memchr and not strchr: a Mere string carries its length and may hold NUL bytes, so stopping at the first one cuts the haystack short._
_The search window is derived from last — the final offset a needle can start at — rather than kept in a counter alongside the pointer. A counter that drifts reads past the end of the haystack WITHOUT CHANGING THE ANSWER, because the extra candidates all fail the comparison; that bug is invisible to a differential test and invisible to AddressSanitizer too, since these strings live inside a region arena and the over-read lands in a block ASAN considers live. The possibility is removed rather than tested for._
_test/ctests/str_index_of_scan.mere is new, and is the first behavioural test this helper has ever had on a compiled backend: every existing one went through Pipeline.process (the interpreter) plus an assert_contains that the emitted C mentions the helper's name._
_`scripts/ctest.sh` identified its subjects by basename. With the default set — one flat directory — a basename is unique and this never mattered. Pointed at contrib/, five basenames appear in two libraries each (ast, build, gen, parser, path): FAIL gen named two different files and their temp files collided on one path. Labels are now <dir>/<stem>._
_Two gates are new, both of them turning something that was written down into something that is asked every run:_
_`scripts/contrib_ctest.sh` runs every contrib library's self-tests on the compiled backend and diffs against the interpreter — 79 of 87 pass, 8 pinned with reasons (7 refused by Wasm emission, 1 where the INTERPRETER overflows its stack and the native binary completes). This is the gate that was missing when contrib/toml spent the project's whole history unable to produce C that compiles. The subject list is explicit and its length is asserted, because a glob that skips what will not compile reports green over a set that shrank._
_`scripts/alloc_region_pin.sh` pins both directions of the trade an allocation makes when it is not visible in a function's type: it falls back to the program-lifetime region, which is conservative and correct and not free. The direction nobody intends to change is pinned too — region changes have shipped three versions of this compiler that segfaulted a 3D viewer on its second frame._
v0.1.478 — 2026-09-13
_Six things the medit2 dogfood found by finishing the editor — and four of them were in code no program had used before._
_`mere lsp` was answering in BYTES where the protocol counts UTF-16 code units. They agree on ASCII and on nothing else, and every test this server had was ASCII, so the whole class was invisible. One line with three kanji on it puts the question six columns to the left, and the server answers about the token it finds there — correctly, and about the wrong thing, with no error anywhere. Converted in both directions: incoming positions become byte columns against the buffer the server is holding, and every outgoing range, diagnostic and semantic token converts back. The token LENGTH converts too, or one kanji gets three columns of highlight. Two discriminating inputs are needed and neither substitutes for the other: kanji separate bytes from characters, an astral character separates characters from UTF-16 units._
_Completion could not offer a builtin. Query.completions_at walks the tree for what is in scope, and str_len and print are not declarations, so typing str_l and asking for a completion offered list_iter. Query.builtin_bindings reads Typer.initial_env; they go last, so a name the file defines still shadows them._
_`tty_raw` was not clearing IEXTEN — the same bug as v0.1.476's IXON, one letter along, found the same way: by binding a key to it. IEXTEN enables VDISCARD (Ctrl-O on BSD and macOS), which eats the byte AND throws away the program's next output, so the symptom is "the key did nothing and the screen froze". ICANON, ECHO and IEXTEN are now cleared together because the reason is one reason; ISIG stays with tty_no_signal_keys, because taking Ctrl-C away is a trade rather than a fix, and both sides are pinned in scripts/tty_raw_check.sh (now six checks)._
_A local `let` bound to a polymorphic VALUE was defaulting to int. let init = ("", Nil) in ... list_fold lines init step emitted a tuple_str_list_int and handed it to a function wanting a list of pairs. The v0.1.99 pass that fixes this for local polymorphic FUNCTIONS looked for uses whose type is a concrete ARROW, and a tuple is not one, so values fell through to the defaulting. `contrib/toml` is written this way and has never compiled on any compiled backend — its self-tests run under the interpreter, which has no such type to get wrong._
_`extern fn ... -> str` was broken. This backend's strings carry a length header before byte 0; a foreign function returns a bare char* that has none, so ++ read whatever preceded the constant as the length:_
extern fn getenv: str -> str;
print ("[" ++ getenv "HOME" ++ "]") => out of memory
_The declaration was right and a comment said the difference was "absorbed by the implicit const conversion" — const was the only part of it that was absorbed. The adoption helper already existed and its own comment named getenv as a caller; the extern path was the one place that did not call it. Fixed on both call shapes (direct, and through the _as_value closure), with NULL — an unset variable — becoming the empty string rather than a crash. contrib/log reads LOG_LEVEL this way and could not be built natively either._
_`now_ms` had no interpreter mock, so any program with a timeout in it was compile-only._
_And one feature, because the editor could not use what was there: semantic tokens now carry keyword, string, number and comment, from the compiler's own lexer, merged with the four name kinds from the tree. An editor that coloured the identifiers and left let and the comments plain does not look like it is highlighting anything, and the alternative was a second lexer in every client. Operators and interpolated-string literals are deliberately left uncoloured: the lexer rewrites an interpolation into several tokens that all claim the literal's own position and width._
_Measured while fixing the toml bug: of the 87 extern-free contrib libraries, 79 compile and agree with the interpreter, 7 fail only at Wasm emission, and one "mismatch" is the interpreter stack-overflowing where the native binary completes. toml was the only genuine C-backend failure._
v0.1.477 — 2026-09-13
_Two CI failures from the last two slices, and the reason both were invisible here: one gate was measuring an allocator rather than the compiler, and one number in the README is not derived from anything._
_region_reclaim's perbigvec legs were comparing peak RSS, which cannot answer the question they asked. "Held but reused" and "freed and re-obtained" have the SAME peak; they differ in what the process is sitting on afterwards. macOS hands large frees back to the OS, so the accumulating case happened to show up as a bigger peak and the gate appeared to work. glibc reuses the block, so the identical binary read flat on Linux and CI went red for a property that was never being measured._
_MERE_REGION_STATS now reports the thing itself — region-stats cache: regions=N retained=B, how many regions are cached and how many bytes they are carrying. A function of the program, not of the allocator. Both directions are pinned with it: under __LANG_REGION_KEEP_MAX a released region KEEPS its grown block (measured: 8,388,608 bytes retained after a 5 MiB value), over it the block goes back and the region re-seeds at 1 MiB (1,048,576). Poisoned in both directions._
_The README's parity count said 176 and test/parity holds 177. first_run_check.sh re-derives it, which is why it was caught; the test count beside it is not derived and was stale too._
_`contrib/unicode/width.mere` is not the first display width in this project, and v0.1.476's entry implied it was. utf8_width has been in the prelude since v0.1.45 and in the stdlib reference all along; the search that missed it looked in contrib/ and in the builtin matrix and never in the prelude. Measuring the two is what settles which to use, so that number is now in both docs: over 17,661 code points they disagree on 2,083 (11.8%) —_
| 1,488 | the table says 0, the prelude says 1: combining marks, format characters and conjoining jamo outside its single U+0300..036F range, U+200B ZERO WIDTH SPACE among them |
| 380 | the table says 1, the prelude says 2: narrow characters inside its coarse CJK block |
| 209 | the table says 2, the prelude says 1: wide characters and emoji outside its two hardcoded emoji blocks |
_utf8_width needs no import and is enough for lining up a column of a table, which is what it was written for. The generated table is for when a CURSOR has to land where the glyph ends — an editor, a pager, anything that draws over what it drew before — because there the error does not stay in one cell._
_And the two gates added in v0.1.476 declare their dependencies in the CI toolchain preflight, where this project already asserts them: a gate that skips in CI is a gate that passes without running. width_check needs ruby and reline; tty_raw_check needs the python3 already asserted there. All four gates run green under dash, which is the shell CI uses._
_dune runtest 2737/0, parity 195/195._
v0.1.476 — 2026-09-13
_Six things a text editor asked for, five of them small and one of them embarrassing. The embarrassing one is that raw mode was not delivering the keys a program had asked for, and every test in this project that could have seen it was a pipe._
_`tty_raw` clears IXON. A fix rather than a choice: no full-screen program wants software flow control. With it on, Ctrl-S is XOFF and Ctrl-Q is XON — the line discipline consumes both and the program never sees either byte. The medit dogfood documents Ctrl-S as save and Ctrl-Q as quit; driven under a real pty it drew zero bytes after each, stayed alive, and never wrote its file. It had been unusable in a real terminal since July. Rebuilt against this compiler, with no change to its own source, it saves and quits._
_`tty_no_signal_keys` is a second call and not part of tty_raw, because ISIG is a trade rather than a bug. Clearing it delivers Ctrl-C / Ctrl-Z / Ctrl-\ as bytes, which an editor needs — with ISIG set, Ctrl-Z is SUSP and an undo bound to it silently does nothing — and takes away the interrupt key, which a game that quits on q should keep. Folding it into tty_raw would have taken the escape hatch from every existing TUI to serve the one that asked._
_Neither is visible to a piped test, in either direction: a pipe has no line discipline, so 0x13 and 0x1a both arrive and everything looks finished. `scripts/tty_raw_check.sh` drives a Mere program through a real pty and asks whether the bytes it was sent reached it — including the leg that pins Ctrl-C still interrupting under plain tty_raw, so that trade cannot be quietly reversed later. It is poisoned in both directions._
_`Grapheme.clusters` builds each cluster as a `str` rather than a StrBuf, and the difference is where it lives. A container is allocated in the program-lifetime region when the allocation is not lexically inside the caller's region block — and a library function never is — so a StrBuf per cluster was a StrBuf that was never reclaimed. Two thousand frames of forty lines of Japanese came to 127.8 MB, growing linearly, inside a region block that was doing its job. As a str the same run is 1.6 MB and flat, and faster: 0.45 s against 0.84 s. The StrBuf was not buying anything — a cluster is one to a handful of code points, so the quadratic that concatenation would pay is bounded by the length of one cluster and not of the text. Still agrees with ICU on all 8,509 conformance inputs._
_That made clustering affordable per keystroke, so `Width` now sums over clusters: 👩👩👦 is 2 columns rather than 6, 🇯🇵 is 2 rather than 4, and nothing else moves. Width.of_cp is still the per-code-point answer, which is what a table comparison wants._
_`\}` is a literal brace. A } in a string never needed an escape — only { starts an interpolation — so \} was "unknown escape", and a brace PAIR had to be written with its two halves spelled differently: "\{\"id\":1}". This project's own test suite is written around it. The unescaped form still works._
_An `extern fn` for a name this compiler implements is checked for arity. It was not before: the declaration was taken at face value, the call emitted with that many arguments, and clang reported too few arguments to function call, expected 3, have 2 about a line of generated C — the user's mistake, named somewhere the user did not write. The expected arity is derived by scanning the C this backend emits rather than listed beside it, so it cannot drift from the runtime. Arity only: comparing types needs a compatibility notion that does not exist here, since tcp_close : int -> unit is what every contrib declares and the runtime returns int. Swept over 400 .mere files in this tree: zero false positives._
_A match's unreachable fall-through is cast to the match's type. The placeholder was a bare 0, and a statement expression is typed by its last expression — so ({ ...; 0; }) is an int, and putting that in the else-arm of a conditional whose other arm is a pointer made clang warn on every match over a user variant. Nothing was broken (__lang_fail_impl is noreturn, so the value is unreachable) but a warning that fires on correct code on every build is one nobody reads, and it stops a -Werror build dead. The medit2 dogfood's emitted C carried eleven._
_Measured and not done: incremental LSP sync. mere lsp advertises textDocumentSync: 1, so every change sends the whole buffer, and the obvious next step was 2. The measurement says no. Per-change time against file size, median of five, didChange to publishDiagnostics:_
| lines | bytes | ms | growth for 2x lines |
|---|---|---|---|
| 100 | 1,981 | 4.2 | |
| 800 | 18,081 | 9.0 | x1.55 |
| 3,200 | 79,881 | 27.4 | x1.66 |
_Linear, not quadratic. And incremental sync would reduce the bytes SENT while the server still re-checks the whole buffer — 80 KB over a socketpair is under a tenth of a millisecond against 27 ms of checking. It would buy almost nothing; what costs is the check, and that is a different piece of work._
_dune runtest 2737/0, parity 195/195, and every gate above._
v0.1.475 — 2026-09-12
_A positioned read that builds a bytes, the region a positioned read lives in, and what a region hands back when it is released. All three came out of one dogfood — a text editor that opens files it cannot hold in memory — and the third one is the one nobody had measured._
_`file_pread_bytes : File -> int -> int -> bytes`, on all four backends. The read half of file_pwrite_bytes, which has existed since v0.1.222; only the write side had a bytes version. file_pread returns Vec[R, int] -- eight bytes per byte, built one fgetc at a time -- which is what mbtree wants for a page it is about to index as numbers, and not what a reader streaming a file wants. Measured on 208 MB in 256 KiB pages, C backend: 3.64 s at 10.1 MB peak RSS becomes 0.36 s at 1.8 MB. Ten times faster and a fifth of the memory, and it removes a choice: before it, read_bytes was 0.33 s at 210 MB, so paging a file meant picking between memory and speed._
_On Wasm this is the cheap one: the host import already HANDS BACK a bytes pointer and file_pread adds a vec_of_bytes after it, so the new builtin is the same call with the conversion removed._
_`file_pread` was the last container constructor missing from the typer's region-marker list. vec_new, read_file_bytes, strbuf_new, map_new, bytebuf_new and the rest bind their result to the region open at the call site; file_pread's region-quantified scheme instantiated a fresh marker that nothing unified with the block around it, and it settled on the default region. A region R { let v = file_pread f off len in .. } paging through a file retained every page it had read -- 200 pages of 64 KiB came to 107 MB, linear in the iteration count, inside a block that was acquiring and releasing correctly. Writing (file_pread f off len : Vec[R, int]) bound it by hand and cut it to 2.0 MB; it now needs no annotation. The identical hole, with the identical cause, is written up in this file for bytebuf_new (Q-127 / m3d Q-10) -- it survived in the one constructor that gets called in a loop._
_A released region keeps its largest block. This is the one that had to be measured rather than read, and the reason is that the bookkeeping was already right. Releasing a grown region freed its whole chain and re-seeded at 1 MiB; every block WAS freed, and an instrumented run agreed -- 40 releases, 40 chain frees, 40 of 40. The process grew anyway, because the next iteration asks the allocator for those megabytes again and malloc does not hand the same pages back. Forty iterations of a 5 MiB Vec inside a region block reached 216 MB of resident memory. Keeping the largest block and dropping the rest reuses the same memory: 13.7 MB, flat._
_Across the sweep that found it (20 iterations, by the size of the Vec's data):_
| Vec data | before | after |
|---|---|---|
| 4 MiB | 8.6 MB | 11.3 MB |
| 5 MiB | 107 MB | 13.4 MB |
| 8 MiB | 167 MB | 19.5 MB |
| 16 MiB | 327 MB | 35.9 MB |
_Resident memory now scales with ONE iteration rather than with the iteration count. The 4 MiB row is the cost: a region that grew no longer shrinks back to 1 MiB. The retained block is capped (__LANG_REGION_KEEP_MAX, 16 MiB), because a program that builds one enormous value in a region and then stops using regions should not hold it for the rest of the run -- eight regions are cached per thread, so an uncapped keep is eight times the largest value the program ever built. scripts/region_reclaim_check.sh pins both directions: under the cap the footprint must not follow the iteration count, over it it must. A trade with one side measured drifts._
_io_poll_new / add / mod / del / wait / get and io_set_nonblocking are documented (v0.1.313, stdlib-reference until now only in the changelog). They are poll(2) over a registered interest set and take ANY fd -- fd 0, or a socketpair to a child process from a shim of your own -- which is what lets one loop wait on the keyboard and a subprocess with no busy-wait and no second thread. The dogfood that found this had designed a busy-polling loop around stdin_byte first. Same shape as file_pwrite_bytes, which the mraft dogfood missed for a whole slice for the same reason._
_dune runtest 2733/0, parity 195/195, region_reclaim 11 checks._
v0.1.474 — 2026-09-10
A type name declared twice with different constructors is refused. This is the last item the exhaustiveness arc left open, and investigating it moved it out of the language's DEFERRED list entirely: it is not a module problem and it does not need module types namespaced.
What was measured first, because the framing decided the size of the change:
type t = A | B;
type t = X | Y;
match A with | A -> 1 | B -> 2 // exit 0, answers 1
Both declarations are accepted, the SECOND wins for the name t, and the FIRST's constructors stay usable — they are keyed by constructor name, not by type. So the compiler holds one model for two types, and everything downstream answers from whichever came last: a match over the first type is checked against the second's cases, == compares values of two types as one, and a function annotated with the name accepts either. All three were measured, at top level and inside modules, and they behave the same. Module types are registered globally and unqualified on purpose — the parser drops the prefix from a qualified annotation, and says so — which is why module A { type t } beside module B { type t } is exactly the top-level collision and not a scoping bug. DEFERRED §4.1 (namespacing module types) is a different and larger change that this does not need.
Restating a type identically stays fine, because it is ordinary here: twelve files in this tree restate 'a list or 'a opt for self-containment, and both declarations describe the same type. Only a CONFLICTING redeclaration is refused, with both constructor sets in the message and a help: that says which half is allowed.
The count, measured before it became an error
Zero conflicting redeclarations across 945 files here and six downstream repositories. Fifteen type names are declared with different shapes across the ecosystem — tree in three shapes across ten files, expr in four — but each shape is a different program, and the registries are per-compilation, so none of them is a collision. The refusal cost nothing.
Which is how it found a real bug of its own
The first sweep reported conflicts that were not there: a type shape in one test program read as a prior declaration of the type shape in the next. `Exhaustive`'s three registries were never cleared between compilations. Typer.reset_type_registries has restored the typer's own tables per program since v0.1.291 for exactly this reason, and the checker's were not in it — so a process that compiles more than one program (the test binary, and the language server on every keystroke) could consult the previous program's constructor list for a type the current one does not declare. They are cleared now.
What this closes
test_basic.ml carried a row asserting today's wrong answer for the half the exhaustiveness checker could not decline — a subset collision (X | Y beside X | Y | Z), where every arm is found in the entry and the extra constructor is reported as missing. That program is refused at its declaration now, before any match is looked at, and the row is a refusal test. The checker's entry_describes_column decline stays as a backstop; the cause it guarded against cannot reach it.
Two modules may still share a CONSTRUCTOR name (Traffic.Red and Mood.Red), which is what module scoping is for and has its own test beside this one.
Unit 2733/0, parity 194/0, exhaustive_check 31, check_cmd_check 883, refusal verified on all five paths.
v0.1.473 — 2026-09-10
The four real holes are fixed, and a nested missing case is an error. v0.1.472 shipped the usefulness algorithm with its new findings as warnings and called that a migration switch. Measuring is what made a better line available: the whole ecosystem had four holes, they were all the same one, and fixing them took four lines.
All four are the same idiom — reading an optional second command-line argument:
let live = match argv with | Cons (_, Cons (b, _)) -> int_of_str b | Nil -> 2000;
With exactly ONE argument neither arm matches. The default is what an absent second argument means, and one argument is as absent as none, so the arm is | _ ->. Two benchmarks (churn/bench.mere, churn/bench_regionloop.mere) and two region-reclaim tests. benchmarks/churn/MANIFEST passes two arguments always, so the measured path is unchanged; what changed is that the one-argument invocation now returns the default instead of falling through.
The line, which is now a property rather than a phase
Not "what the old checker could express". Two classes of witness, and they differ in kind:
| witness | what it names | fix | verdict | |
|---|---|---|---|---|
Cons (_, Nil), Some false, Triangle (_, _), Flag { on = false, off = false } | a shape, from finite signatures only | the arm the error prints | error | |
missing 2, missing "bbx", missing (1, _), missing _ | a value out of an infinite domain | | _ -> … | warning |
The second row's only fix is the catch-all the language reference already asks for at every match over a scalar, so refusing the program adds nothing the warning did not say. The first row's finding carries information the person did not have. witness_has_open_literal is the whole predicate.
It also keeps test/parity/nonexhaustive_caught.mere and its fail/ twin compiling, which matters more than it sounds: they hold the RUNTIME behaviour of a fallthrough, and if every named miss were refused, no compiling program could reach one. The per-file flag hook scripts/parity.sh would have needed is not needed.
After both
Zero findings of either severity across 838 files here and 118 across sixteen downstream repositories, mere-ruby's 51,000 lines included. The promotion cost nothing because the measurement came first.
Unit 2732/0, parity 194/0, exhaustive_check 31, check_cmd_check 883, bench_check and region_reclaim_check green on the four edited files.
v0.1.472 — 2026-09-10
The exhaustiveness check looks inside patterns now. It compared TOP-LEVEL constructors, which answers the common question and is blind to a whole class: | Cons (TId nm, r) and | Cons (TP lp, r) between them "cover" Cons, so a third token kind in that position was never reported. What replaces the four hand-written branches is the standard usefulness algorithm (Maranget), which this pattern language fits without ceremony — no ranges, no array patterns, no lazy patterns, a constructor carries at most one sub-pattern.
Exhaustiveness is usefulness of an all-wildcard row against the matrix of arms, and asking it that way produces a witness: an actual value the match does not handle. The message prints the witness and the help: arm is built from it, so the arm comes out of the same computation that found the problem instead of being assembled next to it.
warning: non-exhaustive match (missing Cons (_, Nil))
= help: | Cons (a1, Nil) -> ...
That is this repository's own benchmarks/churn/bench.mere: an arm for two arguments, an arm for none, and nothing for one.
The count, and a correction to v0.1.470's account of it
v0.1.470 reported 269 candidate sites here from a syntactic sweep and said the first two read by hand were both real holes. Running the actual algorithm over the same 838 files finds four — all of them the same one-element-list hole, in two benchmarks and two region-reclaim tests — and every downstream repository clean, mere-ruby's 51,000 lines included. 269 was a loose upper bound on a necessary-but-not-sufficient shape.
And one of the two hand-read sites was a misreading, corrected in that entry: contrib/db/ redis_cluster.mere has a third arm, | Cons (_, rest) ->, that the first pass did not read. Its match is exhaustive. Which is the argument for computing the answer rather than eyeballing the shape — twice over, since the eyeballing is what produced both wrong numbers.
What it can say that the old check could not
- Inside a constructor.
| Cons (0, Cons (b, _))with| Nilis missing more than a
Cons; the witness says which shape.
- A value, where it used to report an absence.
match n with | 0 -> … | 1 -> …was
"no wildcard arm for int" and is now missing 2. A str witness is one character longer than the longest the column holds; a tuple with a literal component reads (1, _) instead of (_, _), which claimed the pair the arm does handle.
- Records by declared field, with an omitted field read as a wildcard — both forms are
legal source — and bool and unit as the finite types they are, at any depth.
The arm is not the witness verbatim. A wildcard in the witness is a position the arm should BIND, so the headline reads Triangle (_, _) and the arm reads | Triangle (a1, a2) -> ...; a literal in the witness becomes a binder too, because | 1 -> ... for "missing 1" leaves everything except 1 uncovered, and an arm that is a bare binder prints as | _ -> .... scripts/exhaustive_check.sh pastes the arm back into the program and uses every name it introduces, which is what caught the first version of this printing | Triangle _ -> ....
What stayed the same, deliberately
The error is still a single top-level constructor with an unconstrained payload, which is exactly what v0.1.468 refused. Nested findings and open-signature findings are warnings. That is a migration switch rather than a principle, and it is a property of the witness, so there is no second checker kept around to ask: witness_is_shallow.
Two things it buys. The four real holes can be fixed before the class is promoted. And the two parity cases that hold the RUNTIME behaviour of a fallthrough are compiled by exactly this permission — with a complete checker and no permission, no compiling program can reach a fallthrough, so its behaviour becomes untestable without a flag scripts/parity.sh has no hook for. Promotion is: fix four sites, give the harness a per-file flag, then flip.
The same-named-type guard is generalised to any column rather than the top level. Two modules that each declare type t share one entry in the variant registry, so a match over the first type would be judged against the second's constructors; an arm naming a constructor the entry does not have is the evidence that the entry is about another type, and the match is DECLINED — v0.1.468 chose silence there and it is still the right answer.
The dependencies it needed were two pushes the registry never made: a type's declared parameters, without which a polymorphic payload cannot be instantiated (Some of an int opt is int), and the record field list. Both come from Typer, whose dependency on this file is one-way, so they arrive the way the variant list already did.
The search is bounded at 20000 constructor specialisations and declines rather than hangs.
Unit 2732/0, parity 194/0, exhaustive_check 31, check_cmd_check 883.
v0.1.471 — 2026-09-10
A producer that outruns its consumer past 65536 queued messages succeeds on the interpreter and on C, and died on LLVM with no output at all. Found by sweeping the remaining bare abort() sites after v0.1.470 rather than by hitting it: six were left in the LLVM backend, four of them show / eq_ / cmp_ walking variants they generated the arms for — internal invariants — and two reachable from a program.
let c = channel_new ();
let rec fill = fn (i: int) -> if i > 70000 then i else let _ = channel_send c i in fill (i + 1);
| interpreter | queues 70000, exit 0 |
| C | queues 70000, exit 0 — its buffer doubles and copies, carrying each message's region |
| LLVM was | exit 134, nothing on stdout or stderr |
| LLVM now | channel_send: this backend's channel holds 65536 messages and it is full, exit 1, caught by try_or |
The buffer here is fixed at 65536 slots and its own comment says so. The divergence stays: growing it needs the per-message region array the C channel has and this struct does not, which is a different slice. What is fixed is that it no longer dies mute — the label said oom when the condition is capacity, and abort() took the buffered stdout with it, so a program that had queued 70000 messages printed nothing. Named, exit 1, and catchable, which is what every other refusal in that file already was.
try_or around the fill now returns -1 on LLVM and 70001 on C. Both are legible; one of them used to be a signal.
v0.1.470 — 2026-09-10
A `match` that falls through gave four answers, and the note in v0.1.468 said the wrong thing about all of them. Measured this time, on a program that actually reaches the missing arm rather than one that merely has it:
| exit | stderr | try_or around it | |
|---|---|---|---|
| interpreter | 1 | no matching arm in match, with the line | catches it |
| C | 134 | nothing | killed — a signal is not something a jmpbuf sees |
| LLVM | 134 | nothing | killed |
| Wasm | 1 | nothing | trapped |
| RV32IM | 0 | nothing | never failed at all |
So the claim v0.1.468 wrote into six places — "each backend invented a value, and the answer came back wrong instead of refused" — was true of exactly one backend. C and LLVM ran a bare abort(), Wasm executed unreachable, and all three failed late and mute rather than wrongly. RV32IM was the one that really did carry on with a value: the arm chain's last mismatch falls straight into the merge label with the saved stack pointer still in a0, so the match evaluated to a stack address and the program used it. Its comment said "typer guarantees exhaustiveness, so some arm matched" — which it does not, because a match over an int with no wildcard arm is a warning and compiles.
Every copy of the wrong sentence is corrected in place, with what measuring found: lib/pipeline.ml, scripts/exhaustive_check.sh, .github/workflows/ci.yml, benchmarks/json/bench.mere, both places in the v0.1.468 entry, and one comment in mere-ruby. The commit messages are history and stay as they are; this entry is the correction they point at.
One answer, on all five paths
Nothing new was written to do it — each backend already had the mechanism, applied to fail and not to this:
- C:
__lang_fail_impl("no matching arm in match")in place ofabort(). Its own note
explains the choice made for fail: "an uncaught fail is a program error the language defines, not a crash. abort() made the shell report 134 / SIGABRT and could dump core, while the interpreter and the Wasm backend both exited 1 for the same program."
- LLVM: the same helper with a length-prefixed message constant, beside the
lb_push
ones. The other @abort() sites in that file are show / eq_ / cmp_ walking variants they generated the arms for, which is a different claim and left alone.
- Wasm:
call $__lang_fail, which is what v0.1.275 gavevec_getfor the reason
written there — "no message, nothing for try_or to catch, where the interpreter raises a catchable failure". A match that falls through was the remaining one.
- RV32IM:
emit_abort, this file's own compile-time-known message, whose comment
records the same lesson a third time: the copy that used to live there was the write-and-exit half only, "which is why every abort this backend emitted was uncatchable".
The message carries no fail: tag, because that tag belongs to the fail builtin and this is the backend's own failure — the convention the other internal failures already follow, and the one scripts/parity.sh's payload() is written around.
Two parity cases, because one shape cannot hold it
test/parity/fail/uncaught_nonexhaustive.mere pins the message and the status across four backends. test/parity/nonexhaustive_caught.mere pins catchability, and it is in the passing corpus rather than fail/ because it succeeds: every backend exits 0 and the answer differs, which the failure section cannot compare. It is also the only one of the two that covers RV32IM, because scripts/rv_exec_check.sh sweeps the passing corpus and the failure section has no RV leg.
Both use an int scrutinee on purpose. Since v0.1.468 a missing case the checker can NAME is a compile error, so a program that reaches a fallthrough is one the compiler refuses — except where the checker admits it cannot name what is missing, which is exactly the scalar case that stays a warning. That decision is what keeps these two files compiling without a flag; if it is ever revisited, they are what says so.
Poisoned three ways. Reverting C to abort() gives c:EXIT(134) on the uncaught case and fails the caught one; reverting Wasm to unreachable gives wasm:OUT. Parity 194/0 (176 + 18), unit 2724/0, exhaustive_check 31, check_cmd_check 883. RV32IM checked on both cases against the C backend under the emulator, which is the comparison rv_exec_check.sh makes.
What is still open, and why it is not here
The checker's blind spot. It compares TOP-LEVEL constructors only, so Cons (TId nm, r) and Cons (TP lp, r) between them "cover" Cons and a third token kind falls through. Swept: 269 such sites in this repository (of 38,088 matches), 162 in mere-ruby, 127 in mbrowse.
v0.1.472 corrected the two sentences that followed. They said the first two sites examined by hand were both real holes. One was —benchmarks/churn/bench.merehas no arm for a one-element argument list. The other was a misreading:contrib/db/redis_cluster.merehas a third arm,| Cons (_, rest) ->, that the first pass did not read, and its match is exhaustive. And 269 was a loose upper bound on a necessary-but-not-sufficient shape: running the real algorithm over the same 838 files finds four, all the same hole, and every downstream repository clean.
Closing it is the standard usefulness algorithm, which this pattern language fits (no ranges, no arrays, no lazy patterns) and which needs two things the registry does not push yet — a type's declared params, to instantiate a polymorphic payload, and the record field types. That is implementable; making it an ERROR is a migration of several hundred sites, so it wants to land as a warning first. This release makes the runtime half honest in the meantime.
Two types with the same bare name. Not a checker bug: module-internal types are registered globally unqualified and the parser drops the Module. prefix from a qualified annotation, on purpose. So Foo.t and Bar.t are the same type — measured three ways: a match over one takes the other's arms, Foo.X == Bar.X is true, and a function annotated Foo.t accepts Bar.Z. The checker is being asked about a type with two conflicting declarations and has no right answer available, which is why it declines (v0.1.468). Fixing it means namespacing module types — DEFERRED §4.1, a breaking change across the parser, three typer registries and four backends — and it is not an exhaustiveness item.
v0.1.469 — 2026-09-10
`mere check <file>` — accept or refuse the program, emit nothing, say nothing when the answer is yes. There was no way to ask that question. The only cheap thing that looked like an answer was mere -t, which runs the declaration loop and NOT the borrow, spawn- capture or exhaustiveness checks — its own help has said so for a while, and examples/borrow_conflict.mere is the standing proof: -t exits 0 on it and every compiled backend refuses it. So the honest way to ask was to compile the program and throw the output away.
On this repository's largest Mere program — mere-ruby, 51,014 lines — that costs:
mere check | 5.4 s |
mere -c > /dev/null | 33.6 s |
mere -w > /dev/null | 32.0 s |
Codegen is 28 of those 34 seconds, so on a file that size the answer was six times more expensive than the question. On a small file the two are the same 30 ms of process startup: the win is entirely where it is needed and nowhere else.
It runs exactly what the four backends are handed — Pipeline.infer_program: inference over the declarations and over the desugared program, the channel-element Send obligations, the borrow conflicts, the spawn-capture move analysis, and the exhaustiveness findings. Silent on success, because the exit status is the whole interface and a check that prints on the good path cannot go in a loop; that is how mere fmt --check already reads.
The gap it has is stated rather than hidden. Codegen is not run, so a backend that refuses what it was handed is invisible to the bare form — test/parity/bytebuf_edges.mere type-checks and both the LLVM and Wasm emitters refuse it. mere check -c | -ll | -w | -rv runs that backend's emit as well and discards the bytes, which is the only way to ask "will this build" rather than "is this a valid program". The same four flags the compile paths use, rather than a second spelling (--target c) of the same four things.
The gate is a differential, because the hazard is not a bug
What mere check is exposed to is becoming a second -t: a fast answer to a different question. Three hand-picked cases cannot hold that, because the failure would be a program nobody thought to pick. So scripts/check_cmd_check.sh sweeps all 291 files in `examples/` three ways and holds:
check -cand-cagree exactly — they do the same work and differ only in whether
the bytes are printed, so a disagreement is the subcommand having grown its own opinion;
checkaccepts everything-caccepts — a check that refuses what builds sends people
hunting a bug that is in the checker;
- where
checkaccepts and-crefuses,check -cmust refuse it too. 13 of the 291
are such rows, so this is exercised rather than vacuous.
Plus the named cases the sweep cannot supply: -t accepts borrow_conflict.mere and check refuses it (if that inverts, the subcommand has become -t); bare check accepts bytebuf_edges.mere while -ll and -w refuse it (without this, the third invariant could pass on a corpus where no backend ever refuses anything); silence on the good path; and the two usage errors.
Three poisons, and one of them found the gate lying. Making check behave like -t is caught by the named row. Making it silently run codegen is caught both by the accepts-what-builds sweep and by the bytebuf_edges row. Dropping its ~quiet:true was caught by nothing: the silence check read $(...), which strips trailing newlines, and check returns the empty string — so printing it produced exactly one newline that command substitution ate. It counts bytes now, and the poison reports wrote 1 byte(s) ... \n. That is the same trap the puts work two commits earlier used od -c to avoid, walked into one file over.
v0.1.468 — 2026-09-10
A `match` missing a case was a warning, and the warning was printed from inside the interpreter's path — so `-c`, `-ll`, `-w` and `-rv` said nothing at all and exited 0. The four backends that produce the artifact were the four that never mentioned it.
type shape = Circle of float | Square of float | Triangle of float * float; // the edit
let area = fn (s: shape) ->
match s with
| Circle r -> 3.14159 * r * r
| Square w -> w * w; // untouched
Before: mere shape.mere printed one line of warning above the answer and exited 0; mere -c shape.mere emitted C and exited 0. This is the ordinary edit — a case added to a type, every match over it left as it was — which is exactly why the check has to be on the path the person or the agent is actually using.
v0.1.470 corrected the two sentences that used to be here. They said a fallthrough has no value to return, so each backend filled it in with one of its own and the answer came back wrong instead of refused. Measured, on a program that actually reaches the missing arm: the interpreter names the failure and points at the line, C and LLVM ran a bareabort()(exit 134, no output), Wasm executedunreachable(exit 1, no output), and only RV32IM carried on with a value — the saved stack pointer. So on three of the four the failure was LATE AND MUTE, not wrong. That is a smaller claim than the one written here, and still reason enough to move the check to compile time; the runtime half became one answer on all five paths in v0.1.470.
After, on all five:
error: non-exhaustive match (missing Triangle _)
--> shape.mere:4:3
|
3 | let area = fn (s: shape) ->
4 | match s with
| ^^^^^ non-exhaustive match (missing Triangle _)
5 | | Circle r -> 3.14159 * r * r
|
= help: | Triangle (a1, a2) -> ...
= note: or `| _ -> fail "todo"` to compile before writing them
The arm is the message. The edit that provokes this check is almost always the same one, and the next thing anyone does with the error is write the arm it names — so the error writes it, one help: line per missing case, and then the hole: fail is typed 'a, so | _ -> fail "todo" satisfies any match and the rest of the file keeps compiling while the arms are filled in. One finding per match, not one per case: a match missing seven constructors used to be seven separate lines that had to be read together to discover they were one edit.
What stayed a warning, deliberately. no wildcard arm for int is not promoted. That arm of the checker is an admitted approximation — it fires because the scrutinee's cases are unknown, so it cannot tell a match that is missing one from a match over a type it has no model of, and making it an error would demand | _ -> on every match over a scalar rather than point at anything. The split is exactly "can the checker name the case": if it can, the build stops; if it cannot, it says so and carries on. --allow-nonexhaustive downgrades the errors too, for a tree mid-port.
Warnings now go through the same renderer the errors use, which they did not: a warning was a bare line 4, col 3: warning: ... while a parse error two lines above it got the file, the source line and the caret. And warning is yellow where error is red — they were the same colour, which is how a terminal showing both made them one class of thing.
What promoting it found
Three latent fallthroughs in this repository, and one of them was an arm the parser had given to the wrong `match`. contrib/regex's match_re ends with
| _ -> match_seq (Cons (pat, Nil)) s 0;
written for match pat with — and parsed as a third arm of the nested match atoms with, which already had a _ arm above it. So the arm was dead code, the outer match had one arm for seven constructors, and match_re (RChar "a") "a" reached a fallthrough instead of the answer written for it four lines below. Parenthesising the inner match is the whole fix; the RSeq path is unchanged and examples/regex_demo.mere still reports all 21 cases ok. benchmarks/json/bench.mere had no arm for JFloat in a walk whose job is to count every node — unreachable for the document the MANIFEST pins, and for any other one an abort with no message (this sentence said "a number invented"; see the correction above).
A false positive that had been hiding behind the warning. Two modules that each declare type t share one entry in the variant registry, because it keys on the bare name — so a complete match over the first type was judged against the second's constructors and told to add Z. Harmless while nobody read the warning; a rejected correct program the moment it became an error. An arm naming a constructor the entry does not have is the evidence that the entry belongs to another type, so the check now declines instead of naming a case out of the wrong declaration. v0.1.284 fixed the other half of this (a qualified constructor M.Leaf normalised to its bare segment); this is the type-name side of the same collision. The half still answered wrongly — one type's constructors a subset of the other's, where every arm is found and the extra one is reported — is a row in test_basic.ml asserting today's wrong answer, so the day the registry is keyed on the qualified name, that check goes red and says where to look.
Outside this repository: four sites in `mere-ruby`, across three files — eval_e has no arm for EKwSplat (the parser builds it in ten places) and parse_x_target1 none for Nil. Each is a Ruby-semantics question rather than a mechanical fix, and each is a silent wrong answer today on all five backends. --allow-nonexhaustive builds all three files meanwhile.
Gate
scripts/exhaustive_check.sh (31 checks). It holds the refusal on interp / C / LLVM / Wasm / RV32IM — the shape of the bug was "one path checks, four do not", so a gate on one path would have been the bug again — and it reads the arm out of the compiler's own output, splices it back into the program, and runs it. A hint nobody can paste is worse than no hint, and a remembered string here would not notice the day the pattern syntax it prints stops parsing: dropping the parens from the tuple-payload arm turns this red with the parse error to prove it. It also carries the negatives — a match with every arm written must build and be silent on all five paths, | _ -> fail "todo" must build and run on the compiled backends too, and a match over an int must not stop a build.
Swept: 838 .mere files here plus 113 across six downstream repositories. With and without --allow-nonexhaustive the mere tree gives the same 769 ok / 69 pre-existing failures, so nothing in it is refused by this change that was not refused before. Unit suite 2724/0.
v0.1.467 — 2026-09-10
A closure written inside a `region` block that calls an inner function which allocates did not compile, on both compiled backends, and had not since v0.1.453. Not "compiled wrongly" — the emitted C and the emitted IR named values that were not there:
error: use of undeclared identifier '__region_SC' (C)
error: use of undefined value '%v' (LLVM)
v0.1.453 gave a lifted inner function its enclosing region as a leading argument, so a body called from a nested block allocates where it was written (Q-131). A closure that CALLS such a function has to carry that argument in its env. The C backend already pulled a lifted callee's captures into the closure's env — and then filtered them by `List.mem_assoc n current_var_types`, which a region name never satisfies, because it is not a variable. The one capture that mattered was the one dropped. The LLVM backend had no such pulling at all, so it lost ordinary captures too.
Three parts, and only the first is obvious:
- a region capture is not a variable, so it must be let through that filter and typed
as the region marker rather than looked up;
- the lifted function may be defined inside the closure rather than free in it
(m3d's let rec prims inside its walk callback), so the scan is over every name the body mentions or binds, not over its free variables;
- a region the body opens must not be captured —
one_frame_intocontains the
region SC { }, and the curried closure form of it is built at the top level where SC does not exist. Widening the scan without this emitted __env->__region_SC = __region_SC; with nothing to read.
Why it took thirteen versions
It needs all three at once: a region block, a closure written inside it, and an inner function the closure calls that allocates. Nothing in 175 parity programs, 286 examples or 13 downstream repositories had written that. m3d did, the first time it put a block around its frame's working set — which only became worth doing once v0.1.458 and v0.1.464/466 made a block reclaim what a callee allocates. The feature's first real user found the hole its own machinery had left.
test/parity/region_closure_calls_lifted.mere is the shape, in eleven lines. It fails on a build of v0.1.457 with the same message, so this is not something today's region-passing introduced.
parity 175 / 0. dune runtest 2717 / 0. escape 20 routes, 0 holes. 13 downstream repositories type-check. region_reclaim 9 checks, region_params 11.
v0.1.466 — 2026-09-09
The LLVM backend passes the region too, so Q-127's remaining half is closed on both compiled backends.
define ptr @mu_deep__int__Vec_int(ptr %__rp1352, i64 %n)
call ptr @mu_deep__int__Vec_int(ptr %t6, i64 2) inside `region A { }`
call ptr @mu_deep__int__Vec_int(ptr @__lang_default_region, ..) outside it
test/regionreclaim/percall.mere, 10 → 100 iterations | before | after |
|---|---|---|
| C (v0.1.464) | 9 → 80 MB | 2.7 → 4.3 MB |
| LLVM (here) | 9 → 80 MB | 3.3 → 3.3 MB |
Both legs of region_reclaim_check assert FLAT now. Between the two versions the LLVM leg asserted the OLD behaviour and went red the moment it changed, which is how the table got updated instead of quietly rotting — and it is what told me each half had landed.
It is not the same hook, because this backend has no uncurried twin
The C backend hangs the region arguments on __direct, the N-ary form of a curried top-level function. LLVM has no such form: a top-level function is define T @f(P %param), one parameter, and a saturated multi-argument call goes through closures. What it does have is a direct call for the first application, so the region arguments lead there — and the innermost body of a curried function is a separate define that cannot see them, so region_var_of's fallback answers the default region for it. Sound, and it reclaims less for multi-argument functions than C does.
The closure adapter hands over the default region explicitly, because a closure value has nowhere to carry one — which is also what any call site reaching a function that way will have settled on.
The bug that made it look like the region was not reaching the callee
The first version emitted nothing at all, and the reason is worth keeping: every decl name in this backend carries a uniform `mu_` prefix, and the table that made the mangled part does not. So source_name_of_llvm missed every name, region_params_for answered the empty list, and the output was byte-identical to before — indistinguishable from the region genuinely not reaching the callee. A lookup that fails silently and a mechanism that does not apply produce the same emitted code.
parity 174 / 0. dune runtest 2717 / 0. escape 20 routes, 0 holes. 13 downstream repositories type-check. m3d: linalg, gltf, raster, shade, questions, northstar (50 images) and bench all pass — and its own 0.7 MB a frame is unchanged, for the reason v0.1.465 recorded: it has one region block and no call site inside one.
v0.1.465 — 2026-09-09
A correction to yesterday's own measurement: `__heap` is a `TyRef` and is not a block, and the report was counting it as one. v0.1.462 said m3d had eight call sites inside a region block. It has zero. So does the 286-program examples corpus.
The classifier's first arm matched any TyRef (_, r, TyUnit) and called it named, which __heap — the default region — satisfies. The emitted C is what disagreed: 116 region parameters declared and forwarded in m3d, and not one call site passing a real block.
| corrected | named | forwarded | undecided |
|---|---|---|---|
| m3d | 0 | 142 | 190 |
| 286 examples | 0 | 10 | 289 |
m3d has one `region` block in the whole program (src/view.mere:197, in the window path), and the bench that measures its 0.7 MB a frame renders through --out, which does not go through it. So v0.1.464's closing note — "a precision gap to measure next" — was wrong about what the gap is: the region-passing machinery is not missing anything on m3d, m3d is not asking it for anything. The 0.7 MB a frame has been waiting on a region block that the measured path never opens.
That is a better problem to have, and it is m3d's to try: put the frame loop in a block and see whether the number moves. What this repository can say is that the mechanism works — test/regionparams/chain.mere carries a block's arena three frames down, and test/regionreclaim/percall.mere went from 80 MB to 4.3 MB on the strength of it.
A measurement that cannot tell "the default region" from "a region" answers the question it exists to ask with the wrong number, and it did so for three versions.
dune runtest 2717 / 0. The chain fixture's line is unchanged (1 named, 2 forwarded, 1 undecided), which is what says the correction did not move the case that matters.
v0.1.464 — 2026-09-09
Q-127's remaining half, closed on the C backend: a function allocates where its caller decided, because the caller hands it the region.
mu_deep__direct(__region_A, __da0) the call written inside `region A { }`
mu_wrap__direct(__rp1352, __da0) deep forwards its own parameter
mu_build__direct(__rp1350, __da0) wrap forwards its own
mere_vec_int_new(__rp1345) and build allocates in what it was handed
Three frames down, and no rule for chains: wrap instantiates build at `wrap`'s own quantified variable, so reading the call site inside wrap gives wrap's parameter. Unification did the propagation; the backend only reads it.
This is the change that broke m3d for three versions, done the other way round. v0.1.453 let the callee's BODY reach for whatever region was current at run time, which in a chain is the block around the outermost call while every type said the default region. Here the CALL SITE reads what it bound and passes it, so the type and the value say the same thing — and where nothing bound it, what gets passed is the default region, explicitly.
test/regionreclaim/percall.mere, 10 → 100 iterations | before | after |
|---|---|---|
| C | 9 → 80 MB | 2.7 → 4.3 MB, flat |
| LLVM | 9 → 80 MB | unchanged — not taught yet |
Where it had to be keyed, which is the part that took two attempts
A first attempt read the region parameters out of the fn_decl's types and emitted nothing at all: `Monomorph.erase_container_regions` replaces every container's region with `__heap` before those types are stored, deliberately, so a region never splits an instance and Typer.unify at a monomorphisation site does not fail on two different region names. Preserving __rp there instead just moves the failure into unify.
So the region rides beside the type, keyed by the source name — which is what it is a property of anyway. A call site already has that name; Monomorph.instance_of is handed it before it mangles anything. A declaration only has the mangled name, so the map is rebuilt from the table that produced it, rather than by adding a field to fn_decl that a second place would have to keep agreeing with.
And the schemes are recorded at the inner `Let`, not at Pipeline's top-level one: the compiling path has no top-level let, because desugar_program turns every one into a nested Let. The first version recorded in the place that reads like the right place and recorded nothing.
What was checked, and what did not move
- m3d, the program that killed v0.1.453: linalg, gltf, raster, shade, questions,
render_props, northstar (50 images), bench, and reference_check (49 models against three.js) all pass.
- The escape check still refuses the dangerous shape, and nothing new was written for
it: test/escape/callee_built_into_vec.mere is GATED on all four backends and the interpreter, because the call site binds build's region to A and storing into something that outlives the block is a type error by name. That is the whole argument for passing the region rather than guessing it — the type follows the value.
- The self-tail-call loop is intact. Region parameters lead, so a goto that reassigns
only the tracked parameters leaves them alone; the call site declines the goto if a call would hand one back changed. 20M iterations at -O0 still run in constant stack.
- 13 downstream repositories still type-check. parity 174 / 0.
dune runtest2717 / 0. - m3d's own 0.7 MB a frame is unchanged, and the emitted C says why: 116
__rp
occurrences, and not one call site passes a real block — all 57 pass the default region. The eight call sites its own --dump-region-params reports as inside a block do not reach the __direct path. That is a precision gap to measure next, not a correctness one, and it is stated here rather than left for someone to discover from the bench number not moving.
region_reclaim_check now asserts C reclaims and LLVM still does not, so the leg goes red when LLVM is taught too. region_params_check asserts the C output contains a region argument — "no __rp anywhere" was true when nothing passed one and would stay true if the region stopped reaching the callee.
v0.1.463 — 2026-09-09
The region parameters have names now, and nothing passes one — which is the whole point of shipping it separately. The change that closes Q-127's remaining half moves a function's allocations from the default region into whatever region its caller decided. That is the change that broke m3d for three versions. So the plumbing goes first, on its own, where "did anything move?" has the answer no and can be checked.
A quantified allocation region that appears in a scheme's type — one a call site could decide — is linked at emit time to __rpN, where N is the variable's own id. Everything still undecided after that becomes the default region, exactly as before. And __rpN also lowers to the default region, so the emitted code is unchanged.
Why the name comes at emit time and not when the variable is found: instantiate_with_map reads through Ast.walk, so a linked variable stops being copied per use, and every call site would end up sharing one region. It has to be after all inference is over.
Why the name is the variable's own id and not a per-function index: the declaration and the call site then do not have to agree on a numbering — they agree because it is the same variable. Order, where it will matter, comes from walking the type, which both ends already do.
What is checked, since nothing observable changed
- `#bound 3` on the chain fixture. "The pass ran" and "the pass did anything" are
different claims and only the second is worth a gate; poisoned by making the pass link nothing, and it says so.
- Zero occurrences of `__rp` in the emitted C, LLVM IR and WAT. A name in a type that
reaches the output as an identifier is a compile error, or an LLVM use of undefined value — which is precisely the bug v0.1.459 fixed, arrived at by assuming a name in a type is a name in scope. Three backends, asserted separately.
- The table is reset per program, beside
reset_send_constraints, for the same
reason: the LSP re-checks the same process on every keystroke, and one emit would otherwise leave every later run's variables already linked and stop new ones being recorded.
dune runtest 2717 / 0. parity 174 / 0. region_reclaim 8 checks — still recording that a callee-built container is NOT reclaimed on C and LLVM, which is the thing the next slice is supposed to change.
v0.1.462 — 2026-09-09
The mechanism the rest of Q-127 turns on, validated before anything is built on it — and it needs no new machinery, because unification already did the hard part.
The remaining half of Q-127 is closed by passing a function's allocation region in as a hidden argument. The question that decides whether that is writable at all is: at a call site, what would you pass? mere --dump-region-params now answers it, per call:
@wrap -> build : ^param forward the caller's own region parameter
@deep -> wrap : ^param again, one level out
@<top> -> deep : ? no block here -- the default region, which is today
@<top> -> deep : A inside `region A { }` -- pass __region_A
No side table, and no rule for chains. instantiate_with_map makes a fresh copy of each quantified variable per use and infer stores the instantiated type on the Var node, so the scheme's body and that node's type have the same shape: walking them in step reads off what each region variable became. And wrap instantiates build at `wrap`'s own quantified variable, so the call site inside wrap reads wrap's parameter. The chain propagates because unification propagated it.
That is the whole difference from v0.1.453, in one line: there the BODY guessed the runtime current region, which is wrong for a chain and cost three released versions and m3d's second frame. Here the CALL SITE reads what it bound and hands it down.
The numbers, on programs that exist
| named (a real block) | forwarded (down a chain) | undecided (default region) | |
|---|---|---|---|
| m3d | 8 | 142 | 182 |
| 286 examples | 0 | 10 | 289 |
m3d is the program that killed v0.1.453, and Acc.floats — the function whose allocation did it — is in its list with one region parameter and status ok. Eight call sites would actually change behaviour; 142 exist to carry the answer to them. The examples corpus has no named sites at all, because almost none of it opens a region block: for most programs this change is inert, which the earlier "110 functions, all ok" number could not have told you on its own.
test/regionparams/chain.mere pins those four lines exactly, and the gate was poisoned by inlining one link — the two ^param lines vanish and it says so.
And a counting bug in the gate, found by its own new output
The corpus sweep read the first # line of the report to get its totals. The report now emits #sites ... before the summary, so the sweep started counting 289 call sites as 289 value-used functions and printed it without blinking. It reads ^# now. A summary line will state whatever it is handed.
dune runtest 2717 / 0. parity 174 / 0.
v0.1.461 — 2026-09-09
A correction to v0.1.460's gate, and the thing it got wrong is the oldest one in the book: its known-failure list was a photograph of this machine.
region_params_check sweeps examples/ and requires every program that does not type-check to be on a list, by name, so that a newly-broken example cannot quietly replace a previously-broken one. The list was written from what failed here. Three http_* examples import github.com/284km/mere-markdown/..., which resolves out of .mere_modules/ — git-ignored, populated by mere install, present on this machine and absent on the runner. Green here, red on CI, naming three files.
Failures are classified by REASON now. An unresolvable import is a fact about the checkout, so it is skipped and counted; anything else is a fact about the program and must be on the list, which now holds only the four deliberate ones (a type error, two lexer limits, a parse limit). stream_lines, which imports contrib/ relative to the repo root, stops needing a name — it is an import skip like the others.
Verified by reproducing the runner rather than by reasoning about it: with .mere_modules/ moved aside the gate still passes and reports 283 examples, 4 skipped, 102 functions, against 286, 1, 110 with it. Both are true; neither is a failure.
And an unexpected failure now prints the compiler's first line. "stopped type-checking" with no reason cost a thirty-three-minute CI round trip to ask what the reason was — a refusal has to name what it refused, and that applies to a gate's refusals too.
The measurement itself is unchanged: nothing is value-used or partial in either corpus.
v0.1.460 — 2026-09-09
Q-127 stage 1: the size of the change, measured instead of guessed. Nothing compiles differently. mere --dump-region-params lists, per function, how many hidden region arguments closing Q-127's remaining half would give it — and whether it can take one at all.
The remaining half is that a function's body allocates through the region variable in its own scheme while the call site binds a different copy: the type says the block, the value is in the default region. Passing the region IN is what makes those the same variable. v0.1.453 tried letting the body GUESS the runtime current region instead, and m3d's second frame is what said that is unsound for a chain, three released versions later. So the next attempt starts from a number, not an impression.
What it reports, and what it costs to be wrong about
| status | meaning |
|---|---|
ok | every occurrence is a saturated call, so an argument has somewhere to go |
value-used | the name is used as a value, becomes a closure, and a closure's ABI has no room for a region |
partial | every occurrence is a call, but one passes too few arguments; the region belongs to the saturated call |
Disqualified is not an error — those keep today's default region, which is over-strict and never unsound. So the ok fraction is the fraction of the problem the change would actually solve.
The numbers
Over 286 example programs: 110 functions would take a region parameter, all 110 `ok`. On m3d — the program that killed v0.1.453, and the only one here with a call chain deep enough to have done it — 40, all 40 `ok`. memo_fib reports zero, which is the other half of the answer: for most programs this change is inert.
That nothing is disqualified in either corpus is a result and not an absence of a detector: test/regionparams/valueused.mere and partial.mere exist so that both ways of saying no are seen to fire, and both were poisoned.
Also
--dump-region-params resolves imports relative to the FILE, the way running a program does rather than the way -t does. Two examples that -t cannot read type-check under it; one that imports contrib/ relative to the repo root cannot, and is on the gate's known list by name with the reason. A count would have let a newly-broken example replace a previously-broken one in silence.
scripts/region_params_check.sh is in CI. Three fixtures are the check; the corpus sweep is the measurement, with a floor on how many files it examined — a sweep that examined nothing agrees with every claim.
dune runtest 2717 / 0. parity 174 / 0.
v0.1.459 — 2026-09-09
A poison written to test a gate indicted the LLVM backend instead, and the gate it was testing is the one that pins the half of Q-127 that is still open.
The measurement first
region_reclaim_check asks whether a region block returns memory. It has always asked about a TREE built inside the block — a variant, which comes from the current region and is reclaimed. test/regionreclaim/percall.mere asks the other question, the one Q-127 is about: a container built by a function and returned, with only a scalar leaving the block, exactly as before.
| 10 iterations | 100 iterations | |
|---|---|---|
| C | 9 MB | 80 MB |
| LLVM | 9 MB | 80 MB |
| Wasm | — | completes inside its fixed 64 MiB |
C and LLVM hold every one of them: a container's region is in its TYPE, and a function's body allocates through the region variable in its own scheme, which the call site does not bind because the two hold different copies. The type says the block; the value is in the default region. Over-strict, never unsound — and it is the 0.7 MB a frame m3d pays. The leg asserts that this is STILL SO, so that whoever closes it finds out from a gate rather than from a paragraph. v0.1.443's LLVM leg is the precedent for writing it that way round.
Wasm reclaims it, and not by being cleverer: one bump for every region, nothing stores this vector into anything older than the block, so nothing raises the high-water mark (v0.1.458) and the rollback takes it. Right answer, unrelated reason — so the three legs assert three separate claims rather than that the backends agree.
What the poison found
Poisoning that leg means making the allocation lexical, which is what a fix would look like. The poisoned program did not build on LLVM:
use of undefined value '%__region_R'
%t5 = call i64 @__lifted_f_0(ptr %__region_R, i64 %d, ptr %t4, i64 0)
v0.1.453 gave a lifted inner function its enclosing region as a leading parameter, so a body called from a nested block allocates where it was WRITTEN rather than wherever is current (Q-131). The parameter is named __region_R — and on the C backend the block's arena local is also called __region_R, so emitting the capture's own name as the argument happens to name the right thing. LLVM calls the arena %t4. Emitting %__region_R there is invalid IR.
`mere -ll` exited 0 the whole time. Only the assembler objected, and only if somebody ran it. The shape needs a let rec inside a region block inside a FUNCTION: every existing region/inner-fn parity case puts its block at the top level, where the two naming conventions coincide. Five versions.
Fixed by resolving a __region_* capture through current_regions — the same place region_ptr_for looks — at both sites that build capture arguments, the direct lifted call and the closure env. test/parity/region_inner_rec_in_fn.mere is the nine-line shape, confirmed to fail on a build of 2f1e180 with exactly that message.
And the gate's own bug, which the same poison showed
When the percall legs could not build, the section skipped and the PASS line then read its own unset variables: set -u killed the run with a shell error instead of a sentence. A leg that quietly does not run reports the question answered. Build failures are named failures now, the C and LLVM legs are separated so the message says which, and the run needs 7 checks rather than 4.
parity 174 passed / 0 failed. dune runtest 2717 / 0. region_reclaim 8 checks.
v0.1.458 — 2026-09-09
Q-132 closed, one day after the witness that found it — and it was twice as big as that witness showed. Every region in the Wasm backend shares ONE bump pointer, so an allocation made anywhere during a block is inside the block's range and the rollback takes it, whatever the value's type says about where it lives. Three shapes, all measured against v0.1.457:
| written as | interp / C | Wasm at v0.1.457 |
|---|---|---|
a callee builds a Vec and stores it in an outer container | 700 | -424242 |
| a callee builds a string and stores it | v700 | raw memory, pages of it |
vec_push outer <int> written inside the block, past its capacity | 1199 | -7 |
The third is the one that matters for how this was missed. The backend did refuse escaping stores — but the refusal was syntactic, so it never saw the first two (the store is three frames down, not in the block), and it exempted unboxed elements on the ground that an int cannot dangle. The int cannot. The buffer can: vec_push grows a full container by allocating, that allocation comes off the same bump, and the rollback took it. The exemption was true about the value and false about the container.
The fix is to stop reclaiming, not to keep refusing
A store into a container that predates the innermost open block raises a high-water mark to the current bump; the block's exit restores max(its mark, hwm). Sound because the value stored, and any buffer the store reallocated, were allocated before that point and lie below the bump. Conservative because the block keeps its other garbage too — which is exactly what the C backend does with a __heap value: never frees it. So this makes Wasm agree with the semantics the types already claimed, rather than inventing new ones. One global compare when no block is open; nothing for a block that touches only its own containers. region_reclaim is unchanged: C flat at 2 MB, LLVM flat at 5 MB, Wasm inside 64 MiB.
And the refusal is gone, because its stated reason — "the backend reclaims the whole block on exit and has no per-container storage to copy into" — is no longer true, and because a guard that has to be right about every store in the program cannot be syntactic. What is still refused is what a mark cannot help with: channel_send, spawn, and closure-registering externs hand the value to another thread or to the host, on no schedule ordered with the block's exit.
Five things the gates said, in the order they said them
- `test/parity/region_chain_callee_alloc.wasm.expected` went stale within the day.
It was installed the previous version to pin the wrong answer; parity's reply to the fix was wasm matches now — delete the stale pin. Deleted.
- `test/escape/ROUTES` has no open holes again —
callee_alloc_chainis promoted
from HOLE to SAFE. Its row read ACCEPT everywhere the whole time, which is the lesson kept in the header: when four backends agree on a verdict, that is not four backends agreeing.
- Three routes stopped being unreadable on Wasm.
str_out,variant_of_strand
closure_out had all been refused by the store guard too, so their -w column was recording a subset limit rather than a verdict. The gate caught the stale listings.
- Two parity programs gained a fourth backend:
region_closure_value_nestedand
region_inner_fn_nested — the Q-131 witnesses — were Wasm-UNSUP only because of the guard. Wasm's blind spot is 12 of 173, down from 14 of 172.
- A unit test asserted the opposite and now asserts the store compiles, with the
refusal it used to pin re-pointed at channel_send, which still needs it.
test/parity/region_outer_push_realloc.mere is new and is the third row of that table: a separate program from the chain witness because it is a separate mechanism — that one is about the value stored, this one about the buffer the store moves. Both were run against a build of 5fca98b to confirm they fail there (-424242, -7) and pass here.
parity 173 passed / 0 failed, no declared divergences. dune runtest 2717 / 0. escape_check 19 routes, 0 open holes.
v0.1.457 — 2026-09-09
The witness this repository did not have — and it found a live bug on its first run. v0.1.456's lesson was that three released versions shipped a memory-safety change whose only witness lived in another repository. dune runtest, parity, every gate and all 29 dogfood type-checks were green while m3d could not render a second frame. The obvious follow-up is not a note; it is the program.
test/parity/region_chain_callee_alloc.mere is that program, and it is small:
let cache = vec_new (); -- top-level, so its region
let leaf = fn (n: int) -> .. vec_new () ..; -- variable is not quantified
let fill_cache = fn (n: int) -> vec_push cache (leaf n);
let mid = fn (n: int) -> fill_cache n;
let top = fn (n: int) -> mid n;
let _ = region A { top 700 }; -- the block holds ONE call
let _ = region B { .. 4096 pushes of -424242 .. }; -- reuse A's arena
print_int (vec_get (vec_get cache 0) 0)
The allocation happens three frames below the block, where no region is lexically active, and it reaches a container that outlives the block. Built at 62e1858 (v0.1.455) this prints 700 on the interpreter and -424242 on C — the value the second block scribbled, read through a pointer into the first block's reclaimed arena. At HEAD both print 700. The type says __heap either way, so no escape check can fire on it; only running it separates the two answers, which is exactly why nothing here separated them for three versions.
region B is not decoration. An arena is REUSED rather than freed, so with only one block the stale pointer still reads 700 and the bug looks absent — and ASan is blind to it, because the arena is one live allocation.
And Wasm answers -424242 today. That is Q-132.
The Wasm backend has one bump pointer for every region, so an allocation made anywhere during a block — by a callee, three frames down, with a type that says the default region — sits inside the block's range and is rolled back with it. The backend does reject escaping stores, loudly and by name, but its guard reads the block's own body: it cannot see an escape a callee performs.
Not fixed here, because a fix is either per-region storage in Wasm or a call-graph-aware guard, and both are a slice of their own. It is PINNED twice instead, in the two places that can each detect it being fixed:
test/parity/region_chain_callee_alloc.wasm.expectedholds-424242exactly. Any
other output is a DIFF, and the right output makes the pin stale, which parity reports as a failure telling you to delete it.
test/escape/ROUTESgainscallee_alloc_chain HOLE ACCEPT ACCEPT - Q-132— the
first open hole that table has held. It is unlike the others: the row reads ACCEPT everywhere because the TYPE rule is right (nothing escapes, on the reading the types have) and three of four backends implement that reading. The hole is in one backend's storage model, which a verdict column cannot show. The header now says so.
Five comments that claimed the withdrawn behaviour
emit_program in all four compiled backends said an undecided allocation region "becomes __caller (the runtime current region ...)". It does not, since v0.1.456. So did the shared-pass note in codegen_llvm and the Q-127 note in test_basic. All five now state what the code does and, briefly, what it did and why that was withdrawn — a comment describing behaviour the code no longer has is the failure mode m3d's own "this would stop compiling" note had just been caught in.
parity 172 passed / 0 failed (one declared divergence: wasm/region_chain_callee_alloc). dune runtest 2715 / 0. escape_check 19 routes, 1 open hole.
v0.1.456 — 2026-09-09
A CORRECTION to v0.1.453, and the claim it withdraws is the headline one. Q-127 said a container a function builds should live in the region the CALLER is standing in, and settled an undecided allocation region on __caller, which the backends lowered to the runtime current region. The argument was that a call does not change the current region, so inside the callee it must be the region open around the call. That is false for a chain of calls, and m3d is the witness:
region FR { render_at o } -- render_at -> one_frame_into -> attr
-- -> Acache.floats -> Acc.floats
Only the outermost of those is written inside a block, so only its copy of the region variable is bound. Acc.floats's BODY allocates through the scheme's own variable, which nothing bound — and lowering that to the current region put the value in the frame's arena while every type involved said the default one. m3d stored it in a cache that outlives the frame and read it back after the arena was reused: a segfault from the second frame on, in v0.1.453, 454 and 455.
The body and the call site hold DIFFERENT COPIES of the variable. The only ways to make a body allocate where its caller decided are to pass the region in or to specialise per region, and neither exists yet. So an undecided region is the default region again — what it has always meant, and safe. `region R { let v = build .. }` reclaims nothing again: the 822 MB → 10 MB measurement is withdrawn.
What survives, and it is not nothing
The types still name the block, so the escape check still fires. Binding at the call site was doing two jobs and only one of them was unsound. Carrying a callee-built container out of a region block is a type error, reported by the check that was already there — being typed to a region the value does not actually live in is over-strict and never unsound, which is the direction this now errs in.
Every GATED route in `test/escape/ROUTES` is REJECT on the interpreter too. Ten of them read ACCEPT before, with the honest explanation that the interpreter has no arenas so nothing dangles there — but the escape rule is a TYPE rule, and one the toolchain only half applies is a rule with a hole. The hole was at the top level: let o = vec_new () was generalised over its region, so a store inside a block bound a COPY and o's own type never mentioned the block. Giving the top-level let the value restriction the inner one has had since Phase 36 makes it one variable and the check fires. One program, one answer.
And that restriction now exists in one place. Pipeline had the top-level let written out THREE times — process_decls, infer_program_inner, type_of — and they had drifted; the interpreter and -t disagreed about the same program for a while today because two of them had it and one did not.
The lesson, since it cost three versions
The premise was checked against a witness that did not distinguish it: a call made directly inside a block, where the runtime current region and the caller's binding coincide. The case that separates them needs a CHAIN, and the corpus had one — m3d — that nothing in this repository runs. dune runtest, parity, every gate and all 29 dogfood programs' type-checks were green while three released versions could not render a second frame. The bench that would have caught it is in m3d's repository, not this one.
One cell in the host matrix stopped being about something else
bytebuf_new's LLVM cell has read unattributed since 2026-08-30, when that value was invented for a probe that cannot reach its subject. It was one: let _ = bytebuf_new 1; 0 never got as far as the backend's host-builtin check, because the ByteBuf's region slot was still a type variable and ty_tag refused THAT first — unsupported LLVM codegen type element: 'a. The message matched unsupported and named no builtin, so the harness correctly said the cell was about something else.
Defaulting an undecided region settles the slot, the probe reaches its subject, and LLVM refuses in its own name: bytebuf_new has no LLVM lowering yet (host builtin). refused. Fourteen rows still have such a cell; this one was hiding behind a type the backend could not spell.
parity 171 passed / 0 failed. dune runtest 2715 / 0. m3d's four bench models run.
v0.1.455 — 2026-09-09
An exported call *is* a region, and the typer knows it now. mere -c --lib compiles a shared library whose boundary opens a region per call and releases it at return — that is what makes a call a transaction, and it is why v0.1.311 stopped pinning containers to the default region there. Nothing enforced the other half of that bargain: a host that stored a callee-built container into module state and then made one more call read back garbage, because the arena had been reused. Pinned as a known gap in v0.1.452, with the instruction that whoever closed it should delete the pin and say so.
It is closed, and not by choosing a region. A top-level function's body is now typed inside the call's region, so the store is the same mistake as carrying a value out of a region block and gets the same treatment:
region escape across the library boundary: `store` now holds a value built during a
call, which is freed when that call returns (its type became `Vec['a, Vec[__call, int]]`).
Each exported call runs in its own region -- build it at module init, or copy its
contents out
lib_check keeps its section and checks the new claim in both directions: the store is refused and names the binding, while a call that builds containers and keeps none is still accepted with them in the call's own arena. A gate that refused everything would pass the first half and be useless.
And the mode-dependent answer is gone. heap_container_region said "the current region" in --lib mode and "the default region" otherwise; Q-127 makes that unnecessary and, with it, wrong. A container a CALL builds now carries __caller, which is the current region already, so the leak that special case existed for stays fixed by the general rule. What is left as __heap belongs to something outliving the call — module state, or a pass-through — and sending those to the per-call arena was the defect itself.
Two things fell out that were on m3d's list, not on this one
`bytebuf_new` never bound its region to the block it was written in. Unlike vec_new, map_new and strbuf_new it was a plain polymorphic scheme, so the marker stayed open and settled on the default region — which is never freed. m3d measured that as row 1 of its Q-10 table and paid for it twice, allocating its render target once and clearing it per frame because a region around a frame reclaimed none of it. It binds like every other container constructor now: 676 MB → 268 MB on Q-10's own measurement (200 iterations of a 4 MB buffer inside a block). test/parity/region_bytebuf_reclaimed.mere.
A record field can hold a `Vec` now (m3d's Q-8). type box = { v: Vec[R, float] } answered expected &R unit, got &__heap unit: the field's declared region was the rigid name R and vec_new () produced the rigid name __heap, and two rigid names do not unify. A function taking or returning one worked, because a signature can be generalised over the region and a field cannot — so the shape a record wanted was the one shape that could not be written. With the marker a variable until something decides it, the field's declaration is what decides it. ByteBuf always worked and was the exception that gave it away: its region is erased from the tag, so the mismatch never arose. test/parity/record_holds_containers.mere.
The top-level let, typed once
Pipeline had the same eight lines twice — process_decls, which the interpreter walks, and infer_program_inner, which every backend starts from — and they had drifted: neither applied the value restriction Typer's inner let has had since Phase 36. While every region settled on __heap that did not show. Q-127 made it show: let store = vec_new () was generalised over its REGION, so each use instantiated a different one for a container there is only one of. One function now, called from both.
parity 171 passed / 0 failed. dune runtest 2715 / 0. All 29 Mere programs in the dogfood corpus type-check unchanged.
v0.1.454 — 2026-09-09
A function written inside one region and called from a region nested inside it allocated in the wrong arena, and the value died with the inner block. Q-131, found while designing Q-127's fix and confirmed with a witness:
region R {
let mk = fn (n) -> let v = vec_new () in .. v in // typed Vec[R, int]
let cache = vec_new () in // Vec[R, Vec[R, int]]
let _ = region S { vec_push cache (mk 700) } in // mk's body ran with S current
.. // interp 1 / 700, C 4096 / -1
}
Every type here is consistent and the escape check is right to say nothing: the value is typed R and stored in R. What was wrong was the ARENA. mk is lifted out of the block into its own function, so the block's __region_R local does not reach it, and v0.1.433 made it fall back on the RUNTIME CURRENT REGION — "the answer a lexical name cannot give". At the call the current region is S, not R.
The region travels with the function now. A lifted body takes each region block it was written inside as an ordinary capture, named __region_R — which is already the C local the block binds, so the parameter, the argument the call site passes and what region_var_of answers are one string. Typed as the region marker, which c_type_of lowers to __lang_region*; the env struct, the closure adapter, the __direct twin and the transitive capture fixpoint all treat it as a capture and need to know nothing about regions. A lexical name CAN give the answer, provided it is carried rather than looked up.
Two paths, two witnesses, because fixing either leaves the other. The C backend lifts let mk = fn .. inside a block to a function with a leading region parameter; the LLVM backend makes it an anonymous closure and carries the region in its environment. test/parity/region_inner_fn_nested.mere is the first, region_closure_value_nested.mere — the same function passed as a VALUE — is the second. Poisoned separately: removing the C capture makes the first read 4096, removing the LLVM one makes the second read 4294967295.
Both witnesses end with a third block that writes 4096 words over the arena. Without it the stale bytes read back correctly and the bug looks absent — the arena is REUSED, not freed, which is also why a sanitiser cannot see it: it is one live allocation.
Also
- The Wasm backend's refusal for a store into an outer container from inside a block now
says "unsupported in Wasm codegen subset". The phrase is not decoration: parity.sh reads it to tell a documented limit from a backend that fell over, and worded as "not supported yet" this deliberate refusal was tallied as an EMITFAIL — the first parity case to exercise it made the harness red for a program it had correctly refused.
parity 169 passed / 0 failed. dune runtest 2715 / 0.
v0.1.453 — 2026-09-09
A container a function built landed in a region nothing frees, and a `region` block around the call reclaimed none of it. let build = fn n -> let v = vec_new () in .. v came out as int -> Vec[__heap, int]: the container's region was decided when the FUNCTION BODY was checked, where no region is open, so every call site got the same answer — the default region, which is never freed. Two hundred iterations of a half-million-element Vec built through a call inside region R { } reached 822 MB; m3d measured the same thing as 0.7 MB a frame and had to work around it twice. Recorded as Q-127.
The region is now decided by the CALL SITE. Outside a region block the marker starts as a variable rather than the name __heap, the binding generalises over it, and each call binds its own copy to whichever region is open around that call. The same program is 10 MB, and the other half arrives with it: carrying such a container out of the block is a type error, reported by the escape check that was already there —
region escape: `cache` now holds a value from region `A`, which is freed at the end
of this block (its type became `Vec['a, Vec[A, int]]`)
Nothing new checks anything. The type stopped lying, and a check written for region R { let v = vec_new () in .. } in v0.1.291 started seeing the case it was always about.
Quantified AND allocation, which the first attempt got wrong
Only a region a call ALLOCATES INTO is the call site's to decide. let id = fn (v) -> v passes its argument's region through and touching it would be wrong, and the difference is not always visible in the type: mere-ruby's frame pool returns either a pooled map or a fresh one, so the fresh map_new's variable is UNIFIED with the global pool's, and which of the two survives is unification's business. Binding it retyped the global as living in the block, and the escape check reported it — correctly by its own lights.
Level discipline already separates them: a variable shared with the environment is not local to the binding, so generalize declines to quantify it. Every marked variable generalisation declines is unmarked there, once, and non-local never goes back. With that, all 29 Mere programs in the dogfood corpus type-check unchanged — including m3d, whose frame loop had already been written to this discipline ("a value created in this region could not be stored in a cache that outlives it. Hence the warm pass above, out here, and a lookup inside that never writes").
No hidden argument, because a call does not change the current region
A region left undecided settles on __caller, which the backends lower to the RUNTIME CURRENT REGION. That is not an approximation: the current region changes at a region block and nowhere else, so inside the callee it is exactly the region that was open around the call — the same one the typer bound the caller's copy to. The two agree by construction rather than by a rule written twice, and a function called from three regions needs one body, not three.
The alternative was measured and rejected. Specialising per region needs the region to distinguish instances, and the region is deliberately not part of the C type (Vec[R, T] and Vec[__heap, T] are both mere_vec_<T>*, the region a pointer inside the struct). Those two requirements are not compatible: every attempt fixed one and broke the other.
The tag stopped naming the region, in all four backends
ty_tag kept the region for Vec and Map while already dropping it for StrBuf and ByteBuf — and the note on that exclusion said why: it "removes the whole class of mismatch where one spelling resolved the marker and the other did not". With three spellings in play the class came back immediately, as %tuple_Vec___caller_str.. against %tuple_Vec___heap_str.. in the same LLVM function. The region is out of the tag everywhere now, and arrows are compared and unified modulo container regions for the same reason.
That cuts both ways on size, and both directions are the same cause. mere-ruby's emitted C is 2.8% smaller (633 KB) because instances differing only by a region collapse into one; examples/live is 3.9% larger and m3d 1.5%, because a generic helper whose only unresolved variable was a region used to be unnameable and therefore silently dropped — an accidental dead-code pass that this removes. The wasm_live band moved for that reason; there is no reachability pruning for top-level functions, and until now an accident was standing in for one.
Also
test/escape/ROUTES:via_fn_arg's interp column moves REJECT → ACCEPT, which is the
eight other store-into-an-outer-container rows' answer. It was the odd one out because a rigid __heap made its annotation fail to unify — the interpreter rejected it with a message about nothing. The COMPILED column improved: that confused unification error is now the escape check naming the binding and the region.
- The LLVM backend's private
close_open_regionsis the shared pass now. It answered the
same question with one fewer case.
parity 184 passed / 0 failed. dune runtest 2715 / 0. Every gate green.
v0.1.452 — 2026-09-08
A container the callee built does not survive its call in `--lib` mode, and the comment that said it did was wrong. Investigating Q-127 turned up a defect rather than a design question. A host that calls an exported function which stores a returned Vec into module state, then makes one more call, reads back garbage:
before 700 (stored, read straight back — fine)
after -1 (one more call reused the arena)
v0.1.311 stopped pinning __heap containers to the default region in --lib mode because the process outlives every call and the pin was a 512 B/call leak; the boundary's rule is that containers cannot cross it. For a container created in a NAMED region the compiler already enforces exactly that, by name:
region escape: `cache` now holds a value from region `A`, which is freed at the end of
this block ... build it outside the block, or copy its contents out
It cannot say that here, and the reason is the one Q-127 named: `__heap` means two things. It is the default region during module init and the per-call region during a call, so the escape check reads it as "outlives everything" while the lowering puts the value in memory the call reclaims.
The store does not rescue it either, and this is the part that was written down wrongly. The comment on heap_container_region claimed copy-on-store carried the contents to safety. Copy-on-store does not copy CONTAINERS: containers are shared by identity in this language — mutate one through either name and both see it, in all four backends, which is what the interpreter says too — so __mcopy_Vec_<T> is generated as the identity function by design. It copies the strings and records inside; the container itself is a pointer, and the pointer is into the arena that just went away.
Nothing is silently changed here. Both sides are now pinned in scripts/lib_check.sh, because they are mutually exclusive as the design stands and the gate should say so: section 8 already required that repeated calls do not grow the default region (pinning the leak fix), and section 9 now requires that the callee-built container still reads back wrong — and fails, loudly, with instructions, the moment it reads back right. Poisoned both ways: making __heap default-region again (the obvious "fix") is caught by section 8 as 123 MB over 50,000 calls, and a probe that stores an int instead of a container is caught by section 9 as "the gap looks closed".
That pair is the finding. You cannot close this by choosing a region; the two requirements meet only where `__heap` stops meaning both things, which is Q-127's call-site instantiation and a larger change than a slice.
v0.1.451 — 2026-09-08
`mere -ll` emitted a `musttail` that clang would not build — and only at `-O0`. LLVM has to keep musttail at every optimisation level, and past a return width the target cannot forward it stops with
fatal error: error in backend: failed to perform tail call elimination
on a call site marked musttail
-O1 and above optimise the call away before that check, so the failure lived exactly where the build is unoptimised: parity.sh's LLVM column, the first thing a reader types (clang out.ll), and any sanitiser run — which is why ASan was given up on during the Q-129 hunt rather than pointed at the miscompile it would have named.
The width is the ABI's, not the IR's. Swept 1..20 8-byte fields, self tail call returning the record:
| forwards up to | first refusal | |
|---|---|---|
| arm64 (Apple clang 21) | 8 leaves (64 B) | 9 |
| x86-64 (Ubuntu clang 18.1.3) | 4 leaves (32 B) | 5 |
So the emitter now marks musttail only when the return is at most four leaves, the smaller of the two — and the size it uses is an upper bound rather than a layout model: every scalar this backend emits is at most 8 bytes and at most 8-aligned, so leaves × 8 can never come out under the real size, padding included. Over-counting costs a musttail; under-counting would emit IR the target refuses, so anything the counter does not recognise counts as too big.
The price, stated. A self tail call whose return is wider than 32 bytes is no longer constant-space on the LLVM backend: a 12-double accumulator that looped forever at -O2 now reports stack overflow (recursion too deep). Three things make that the better side of the trade. The C backend — the default, and what every gate builds — turns self tail calls into a goto and is untouched. Q-129 had already taken the tail call away from every aggregate return that did not go through this one branch (tail on an sret call is what produced NaN on x86-64), so the guarantee was never general. And a named stack overflow is a worse day than an unbounded loop but a much better one than a compiler that dies.
Nothing in the corpus was returning one. Every aggregate-returning musttail in all 167 parity programs — 32 of 2066 sites — is a two-leaf { i32, ptr } or { ptr, i64 }, all of which keep it. The whole class of wide returns had no witness at all, which is why a compiler-killing IR shape survived: test/parity/wide_record_tail_return.mere is that witness now, and on v0.1.450 it fails to build.
scripts/musttail_budget_check.sh asks in both directions, because a bound with only one is a number that outlives its reason: with the shipped budget every return width from 1 to 20 must build at -O0, and with MERE_MUSTTAIL_LEAF_BUDGET raised — which makes the emitter produce the very musttail the bound is refusing — the target is asked where its own line is and compared against the pinned table above. If LLVM learns to forward more, this fails and says the budget can go up. Both directions were poisoned: raising the shipped budget past the line reproduces the refusal for widths 9..20, and a wrong table entry is reported as stale.
Who could not be built, concretely: three of m3d's LLVM-capable tests — linalg_dump, linalg_props and `shade_props` — were refused at -O0 on v0.1.450, all three on musttail call %mat4 (16 doubles). shade_props is the file whose x86-64 NaN was Q-129: the sanitiser that would have named that miscompile could not be pointed at it, because the build it needs was the one that did not exist. All three now build at -O0 under ASan + UBSan, run clean, and agree with the interpreter.
Swept the whole parity corpus the same way: 148 programs built and ran under ASan + UBSan with output identical to the unsanitised `-O0` build, 0 differences, 0 build failures (19 do not emit LLVM at all — the backend's documented refusals). dune runtest 2713 passed / 0 failed, parity 184 passed / 0 failed.
v0.1.450 — 2026-09-08
The C compiler was choosing the order the operands ran in, and gcc chose the opposite of everything else. p "one" ++ p "two" ++ p "three", where each operand prints, printed one/two/three under the interpreter, the LLVM backend and clang — and three/two/one from the same emitted C built with gcc. Argument evaluation order is unspecified in C; the backend had already decided this for calls (every argument is bound to a __da temporary before the call) and had simply never decided it for operators. parity.sh builds with clang unless CC says otherwise, so nothing here could see it.
And the same nesting was the rest of the bracket-depth tax. a ++ b ++ c is concat(concat(a, b), c): one bracket per link. So is (((a && b) && c) …), and so is if … else if …, which becomes (c1 ? x : (c2 ? y : …)). Clang stops at 256 and mere-ruby's prelude writes all three in the hundreds — which is why v0.1.449 fixed the let chain and mere-ruby still needed -fbracket-depth=4096.
One change answers both. A chain is emitted as one statement expression with a temporary per step, and a long else if cascade is emitted as statements, where each if closes before the next opens:
/* a ++ b ++ c */ ({ __auto_type t0 = A; __auto_type t1 = B;
__auto_type t2 = concat(t0,t1); … t4; })
/* a && b && … (10 links) */ ({ int t0 = A; int t1 = t0 ? (B) : 0; … t9; })
/* if … else if … (10) */ ({ T r; if (c1) { r = x; } else if (c2) { r = y; } … r; })
Each rule fires only where it is owed, so ordinary code keeps its shape: n * n * n is still ((n * n) * n), a chain with one effectful operand is still nested, and a two-arm if is still a ternary. Sequencing costs a temporary; a chain gets one when it has two or more operands that can have effects (nothing to order against, with one) or when it is eight links or longer (the depth reason, which holds even for a chain of string literals). && and || are never sequenced — the whole point of them is that the right side may not run — they are rewritten into conditionals that keep the skip, and logic_chain_short_circuit.mere prints from inside every operand to say so.
Measured, all by asking the tool that imposes the limit rather than counting brackets:
| before | after | |
|---|---|---|
gcc, p"one" ++ p"two" ++ p"three" | three two one | one two three |
300-link ++ chain, clang -fbracket-depth=256 | refused | accepted |
| mere-ruby, 154k lines of emitted C, same question | refused (one line, 524 kB) | accepted |
So mere-ruby no longer needs `-fbracket-depth` either — the flag that started as m3d's and was then rediscovered per repository is now unnecessary in both. Built at HEAD with no flag, its 204-program corpus matches ruby 4.0.6 on all 204, and bootstraptest reads pass=1628 fail=27 err=41 against pass=1627 fail=27 err=41 for the same tree built with v0.1.449 — the same failures, the extra pass being one pair the reference reproduced on this run and not the last (the harness counts those as drift and leaves them out of the denominator, which moved 1696 → 1697).
dune runtest 2710 passed / 0 failed. parity 183 passed / 0 failed, including the three new cases. m3d renders all 50 models byte-identically to the v0.1.449 build, which is also what says its reference gate's one red model was the browser and not this change.
The bug the hand-written tests could not find is worth naming: conditions arrive already parenthesised, so the cascade trimmed the redundant pair — and a condition that is a CALL is emitted as ({ … }), where in if (X) the paren belongs to the if. The brace then lands where C wants an expression and the file stops parsing. Every synthetic cascade in the test suite compared integers, so every one of them passed; mere-ruby's prelude, which dispatches on predicates, did not. test/parity/if_cascade_call_conditions.mere is that shape, written down.
scripts/bracket_depth_check.sh is the new gate and it asks rather than counts: clang -fsyntax-only -fbracket-depth=256 on three generated 300-link shapes, with the same files at -fbracket-depth=4 required to be REFUSED so a probe that failed to emit cannot read as a pass — and the order half builds the same C with every compiler on the machine, because one compiler cannot disagree with itself. Recorded as Q-130, and it closes Q-128.
v0.1.449 — 2026-09-08
A chain of `let`s no longer nests in the emitted C. Each one used to become its own statement expression — ({ a; ({ b; ({ c; ... }) }) }) — so the bracket depth of the result grew with the LENGTH OF THE CHAIN, about two levels per binding. Clang's default limit is 256 and Ubuntu's clang 18 enforces it where Apple's clang does not, which is why programs built here and failed there: m3d's main sat at 533 with its CI red at the first gate for the life of that project, and mere-ruby's CI carries -fbracket-depth=4096 for the same reason. A tax rediscovered per repository is one the compiler should not be charging.
The chain now collects into one ({ s1 s2 s3 ... final; }). Three things keep it honest: a name already bound in the chain ends the run (two declarations of mu_x in one C scope is not the same program), anything that is not a plain binding — a tuple or constructor pattern, an owned-vec binding that wants a scope-end free, a global whose initializer is not the recorded one — keeps its old emission with the chain resuming underneath it, and the guard only fires when the FIRST link is flattenable, so the two paths are mutually exclusive rather than two spellings of one thing.
Measured: m3d's emitted C went 533 → 259 deep, a fifty-binding recursive probe went 108 → 22, and the frame did not grow — same sub $0x10,%rsp, same 32,000 recursion levels at a 1 MB stack, which is the thing that would have made this a bad trade. All 50 of m3d's pictures are byte-identical; mere-ruby builds and its 204-program corpus still matches ruby 4.0.6 exactly. parity 163 passed, 0 failed.
*Correction, same day: the numbers above are from a hand-written bracket counter, and clang does not count the way it does. Asked properly — `clang -fsyntax-only -fbracket-depth=256`, which is what a stock clang enforces — m3d's emitted C now compiles, and did not before, and it builds on the CI image (Ubuntu clang 18) with no flag at all. So for m3d this is not a reduction, it is the end of the tax: `scripts/ccflags.sh` there keeps the flag only as a guard.
mere-ruby still needs it, and the cause is named: the prelude's "..." ++ "..." is a right-nested chain of __lang_str_concat calls, the same shape as the let chain and untouched here. A 64-element list literal (right-nested Cons) is the other one, and it is now under the limit rather than over it. What would fix those generally is a depth-limited hoist of right-nested applications, which fixes an evaluation order C leaves unspecified — a decision rather than a repair.
And the measurement lesson is the reusable part: a metric you wrote yourself is not the one the tool enforces. The counter said 533 → 259 and read "still over 256"; the compiler said "was an error, is not". Ask the thing that will refuse you.
v0.1.448 — 2026-09-08
A closure call that returns a struct was miscompiled on x86-64, and only there. Nine calls of the same expression: the first eight answered correctly and every one after that answered NaN. Right at -O0, wrong at -O1 and above; right on arm64 (macOS and Linux both), wrong on x86-64 (Linux and Rosetta both); the C backend right everywhere. That shape — same input, same code, a cliff after the eighth call — reads as state rather than arithmetic, and it was neither.
LLVM's `tailcallelim` marks ordinary calls `tail`. On x86-64 a three-double struct is returned through memory (sret) rather than in registers, and a tail-marked indirect call returning one is where it broke. Bisected to that single pass: opt -passes=tailcallelim ALONE on unoptimised IR reproduces it, and inline,tailcallelim does not. The fix is to emit `notail` on indirect calls whose return type is an aggregate — the two places this backend calls through a closure pointer. It costs nothing: those are calls in value position, and the ones that must stay constant-space go through the musttail branch, which is unchanged.
It was found from outside. m3d's shading properties agree across four backends on the developer's machine and failed on CI, which is the first thing that had ever compiled this language's LLVM output for x86-64 and run it. The minimal case is 17 lines and is now test/parity/tailcall_aggregate_return.mere — a regression test for the CI machine rather than for the laptop: it passes trivially on arm64 and is the whole point on x86-64.
What this does not fix, said plainly: mere -ll output still fails to build at -O0 on x86-64 with "failed to perform tail call elimination on a call site marked musttail", which is the same aggregate-return ABI meeting musttail's guarantee. That one is a trade-off rather than a bug fix — dropping musttail for aggregate returns would cost the constant-space guarantee on arm64, where it works — so it is recorded and not guessed at.
v0.1.447 — 2026-09-07
A `type` declared inside a `module` produced C that did not compile, in two different places, and mere -c emitted both of them happily. A dot is not an identifier in C, in LLVM IR or in a Wasm name; the record's own typedef was mangled to M__t and these two were not.
The first is ty_tag, which names a closure from the tags of its parameter and result: a function value of type int -> M.t was named closure_int_M.t. The same three lines were written in Monomorph, codegen_wasm and codegen_llvm and all three were wrong.
The second needed a different shape to reach. A match whose arms cover every case still emits an unreachable fallback, and that fallback names the result type to build a zero of it — (M.t){0}. So a match RETURNING such a record was broken where one merely mentioning it was fine.
Neither shows up unless the function is used as a VALUE. Calling M.mk directly never materialises a closure type, so a program can use module-qualified records all day without meeting the first; the second needs a total match returning one. Both were found by pointing a glTF reader at the language: a type inside a module, handed to a higher-order function. test/parity/module_qualified_record_closure.mere pins all three tag-building paths (closure, container, tuple) plus the match fallback.
And the self-hosted Wasm codegen had to be told about `float_of_str`, which v0.1.446 made contrib/json depend on. That codegen keeps a hand-written list of the builtins it knows and lowers every value as an i32, so a float has nowhere to go; the list going stale the moment the language gained a use for one is the shape a hand-maintained list has.
It lowers to a TRAP rather than to a placeholder, and the difference matters. str_of_float next to it borrows show_int and prints a bit pattern, which is wrong in a way you can see. A float_of_str that borrowed int_of_str would read 0.8 as 0 and hand back a document that is silently different. The module still validates, so contrib/json stays inside that codegen's coverage — everything else in the file is compiled and checked, and only this one operation is an unreachable.
_(A third thing was noticed and not fixed: from outside the module, that type cannot be named in an annotation at all — M.t and t are both rejected against what M.mk returns, and the error says they are different named types. show prints it as M.t. That is a typer question and not a codegen one.)_
v0.1.446 — 2026-09-07
`contrib/json` could not parse a number with a decimal point. Found by pointing a glTF loader at it: every colour, every node translation and every accessor bound in that format is fractional, so the format was simply unreadable. The limitation was already known — examples/graphql_server.mere carried a comment saying JNum is an int and that a fractional variable "would need a json that can hold one, which is that library's problem" — and had been left there.
JFloat of float is a SECOND number constructor, not a replacement. Which one a number becomes is decided by how it is WRITTEN, a point or an exponent, and not by its value: 1.0 is a JFloat and 1 is a JNum. Three reasons, and the third decided it — a single float type would change every existing reader, to_json_str would stop round-tripping (12 coming back as 12.0), and the documents this is pointed at genuinely distinguish the two: an array index, a byte offset and a glTF count have to stay exact where a colour component does not. Json.as_float widens either for readers that do not care.
It must not have got more permissive, and the first version had. 1. parsed as 1.0, where the integer-only parser it replaced left the . behind and failed. JSON wants a digit on both sides of the point and after the exponent; five malformed literals are now pinned as refused, next to the ones pinned as accepted.
Adding a constructor made every existing `match` over `json` non-exhaustive, and the three in this tree were wrong in two different ways. contrib/schema/reflect.mere had a wildcard returning "", so a fractional value silently became the empty string — a catch-all answers wrongly for every kind it has not been told about, and the kind arrived without it having to change. examples/graphql_server.mere had no wildcard and would have failed at runtime; it now maps JFloat onto the GFloat that GraphQL's own value type already had.
Which surfaced something worth a question of its own: the non-exhaustive-match warning is printed by the interpreter and by nothing else. mere -c, -ll and -t are silent. The downstream gate — mere -c over twelve dogfood repositories — therefore cannot see a match that a new constructor just broke, in either of the two ways above. The three in this tree were found by reading, not by a gate. Recorded rather than fixed here, because making the warning reach the other entry points is a change to those entry points and not to a JSON library.
All thirteen downstream repositories still compile, and that is a weaker statement than it sounds. The two that use this library — mq and mere-blog — VENDOR it at a pinned revision, so the change has not reached them: they are green because they are still reading the old copy. mq is a jq-lite whose serialiser has no wildcard arm, so when it updates the pin it will fail at runtime on a fractional number, and by the paragraph above nothing will warn it. Deferred, not avoided, and written down here so the next person to bump that pin meets this sentence.
Leading zeros are still accepted (01 parses as 1) where the grammar forbids them. Pre-existing, and deliberately not changed as a side effect: a parser getting stricter is a separate decision from a parser getting a feature.
v0.1.445 — 2026-09-07
`f32x4`, the third 128-bit SIMD type. Four single-precision lanes, on all four backends that have a floating-point unit: the C backend's vector extension, LLVM's <4 x float>, Wasm's v128, and the interpreter as the oracle. RV32IM / RV64IM refuse it by name, the same way they refuse f64x2 and for the same reason.
Eight builtins — splat, make, extract, add, sub, mul, div, reduce_add. No `load` or `store`, deliberately. For f64x2 a load is a load, because Vec[R, float] already holds doubles; the same spelling on f32 lanes would be a narrowing conversion of four doubles wearing the name of a load, and the caller that actually wants four f32 lanes off a buffer wants them from bytes. Two different right answers is not a thing to guess at before a caller exists, so the name is simply unbound.
The lanes are single precision and the language's scalar `float` is a double, so this is the one SIMD type whose boundary converts: narrowed going in, widened coming out, both the backend's own round-to-nearest-even — the delegation f32_bits already makes (Q-038). A value with no float32 becomes an infinity rather than wrapping. f32x4_reduce_add is specified left to right at LANE precision, (((l0 + l1) + l2) + l3), because a pairwise tree rounds differently and four backends have to agree.
test/parity/simd_f32x4.mere holds all of that across interp, C, LLVM and Wasm in IEEE-754 bit patterns. Its two discriminating lines were both WRONG when first written and were fixed against an independent float32 model before being committed: (1, 1e-8, 1e-8, 1) gives 2.0 whether the reduction is left-to-right or pairwise, and the chain that was supposed to separate lane-precision arithmetic from "compute in doubles, narrow once" did not separate them either. The file now uses (1, 3e-8, 3e-8, 3e-8) and (1 + 0.1) - 1, which do — and a poison (the interpreter's reduction switched to pairwise) was run to confirm the gate goes red.
`benchmarks/mat4xvec4_f32`: the mat4xvec4 kernel at single precision. A vec4 fits in one f32x4, so a vertex costs eight vector operations against the f64 row's sixteen — a ceiling of 2x, of which 1.37x arrives: 100 ms against the f64x2 row's 137 ms, 1.9x the scalar double row, and ahead of C's scalar float at 131 ms. Its answer is single precision and therefore not the f64 row's answer, so it is a separate directory: run.py's same-answer rule is per directory and is correct in both. The f32 accumulator does not saturate, which was measured (the checksum still moves between 4999 and 5000 passes) rather than assumed — a saturated one would have let a wrong implementation agree with a right one by both running out of mantissa.
v0.1.444 — 2026-09-06
`region_reclaim_check` measured peak RSS with a tool the CI runner does not have. It shelled out to /usr/bin/time -l, which is the BSD spelling; GNU wants -v and labels the line differently, and the Ubuntu image the gates run on does not ship the binary at all. Green here, red there, on the first push after the gate was added.
It now compiles its own wrapper around getrusage(RUSAGE_CHILDREN) — POSIX, no package — and the one platform difference that remains, whether ru_maxrss is bytes or kilobytes, is a compile-time question the C file answers instead of something the shell guesses. Checked on both: the same program reports 18.2 / 68.5 / 269.8 MB for a 16 / 64 / 256 MB working set on macOS and on the Linux CI image, within a few KB of each other.
The gate's own "only 2 checks ran" guard is what turned this into a red CI line rather than a pass that measured nothing.
_(The first attempt to verify the Linux side reported 3.2 MB for a 256 MB program, which looked like a broken wrapper and was a broken test: the memset into a malloc nothing read had been optimised away. The instrument was right and the subject was not.)_
v0.1.443 — 2026-09-06
The LLVM backend reclaims region blocks: 316 MB becomes 5.8 MB (Q-116, Q-115). v0.1.438 measured the gap and pinned it; this closes it.
Three pieces, in the order they had to happen:
A current region. @__lang_current_region is a global, and @__lang_alloc(i64) reads it -- so the rule about where a value goes is written once instead of at each of the nineteen allocation sites that named @__lang_default_region directly. A region R { } stores itself there for its body and puts the old value back afterwards.
A copy-out. With values living in the block, the block's result has to leave before the block is released. @__mcopy_<tag> is the sibling of the existing @eq_<tag>: the same structural walk, allocating into a named region instead of comparing. Scalars pass through; a str is one allocation and one memcpy (its length is the eight bytes before its pointer); bytes likewise; tuples and records rebuild the aggregate over copied fields; a variant gets a new node and a new payload box, which is where the first version was wrong -- both representations keep the payload behind a pointer, and inserting the value directly is a type error LLVM catches.
Two results are refused, by the same rule the C backend uses: a container (identity -- a copy would be a different object; C has refused this since v0.1.31, and examples/todo_app.mere now gets the same answer from both backends instead of compiling here and not there) and a function (its captured environment is in the block, and this backend has no environment copier). A result whose type never resolved is not copied: the only values that reach a region boundary without a concrete type are the shared nullary nodes of a boxed variant (v0.1.322), which live outside every arena. region A { Nil } is that case.
One rule for containers (Q-115). Where a container whose typer region is R goes was answered two different ways in seven places here: three fell back to the default region, four refused outright. Now all seven ask region_ptr_for, which gives the block's pointer when the block is lexically present and the runtime current region when it is not -- which is the case inside a function defined in a region body, since that function is emitted on its own. test/parity/region_inner_fn_container.mere was UNSUP on LLVM since v0.1.433 and now matches; &R v and view literals still refuse an out-of-scope region, because those name their arena explicitly, exactly as on C.
Evidence. region_reclaim_check flat at 5.8 MB and red again if the values go back to the default region (poisoned: 127 -> 316 MB, the original figures). Every result shape -- str, tuple, list, variant, nullary, record, option, nested -- byte-identical across interp / C / LLVM and clean under AddressSanitizer, which is what caught the missing copy-out in the first place. codegen_identical against the previous binary: 2451 emissions over C, Wasm and RV32, zero differences -- the change is LLVM's alone. Parity 159/0, tests 2693/0.
Three unit tests that asserted "uses default region" now assert the new spelling; they are the gate for this working, and they fired.
v0.1.442 — 2026-09-06
The `exit` runtime section had no program. v0.1.434 added a gated Wasm import for exit and v0.1.439 a second one for command components; scripts/section_coverage.sh says a gated section with no program that asks for it is untested code that looks tested, and it was right -- wasm_exit_used had no row. test/parity/exit_zero.mere is that program, and it is the only parity program that calls exit at all, so the ordinary suite now also compares the status of a program that says it succeeded. Parity is 159 programs.
A nonzero status still cannot live in the parity corpus: the harness skips any program the interpreter does not exit 0 on, and its failure section is built for programs that fail with a diagnostic. scripts/exit_status_check.sh covers that half.
v0.1.441 — 2026-09-06
Durable computation, and the asymmetry that shapes it. test/durable/jobs.mere (v0.1.437) records finished jobs, so what survives a kill is a boundary between pieces -- which only helps when the work divides into independent ones. test/durable/fold.mere is one computation with a running state, checkpointed part-way through and resumed mid-stream.
Writing it turned up a documentation defect. to_json was described as serialising any value, and it does -- a Vec becomes an array, a Map an object -- but `of_json` reads back neither:
| value | to_json | round-trip |
|---|---|---|
| int / float / str / tuple / option / list / variant / record | yes | yes |
Vec | [1] | of_json: expected a variant value for Vec |
Map | {"1":"a"} | of_json: k is not a case of Map |
StrBuf | "x" | of_json: x is not a case of StrBuf |
The witness form already says why in its own words -- "a closure or a handle cannot say what to decode into" -- but nothing said the pair was one-directional for containers, and the place it bites is exactly a program that checkpoints: the state it wants to save is a Map, it saves it happily, and it cannot read it back. docs/stdlib-reference.md now carries the table and the way through (keep checkpointed state in the shapes that round-trip; rebuild containers from them on resume), which is what fold.mere does.
scripts/durable_check.sh gains the fold and one question that only a mid-computation checkpoint can be asked: the run that finishes must report starting somewhere other than zero. Without it, a resume and a restart that happened to be fast look the same. Poisoned by checkpointing the index but not the accumulator, the gate names the difference.
v0.1.440 — 2026-09-06
What it costs to bound an incremental cache with the tool the language already has. contrib/inc holds every superseded value forever, because the default region does not give memory back -- test/inc/agg.mere retains about 756 bytes per edit where the useful change is roughly 200. region R loop is the answer the language offers: deep-copy the carry into a fresh arena every N edits, release the old one whole. test/inc/agg_region.mere carries the entire engine through one.
1024 leaves, 2000 edits, same answer either way:
| cumulative allocation | peak arena | |
|---|---|---|
| the engine, no region | 2.16 MB, none of it returned | — |
region R loop, one generation | 1.67 MB | 3.1 MB |
| every 500 edits | 3.07 MB | 1.0 MB |
| every 100 edits | 10.5 MB | 1.0 MB |
| every 25 edits | 38.6 MB | 1.0 MB |
The construct is priced by the live set, and here the live set is the whole cache. Reclaiming four times as often costs nearly four times the allocation, because each generation copies everything that is still alive, and in an incremental cache everything is still alive -- the garbage from one edit is two strings. benchmarks/churn is the same construct on the opposite shape (a bounded live set, unbounded garbage) and pays a 1.44x premium for 20x less resident; this shape pays 18x at one generation per 25 edits.
That is not a bug in either. It is the trade the construct offers, stated in both directions now that both shapes have been measured.
scripts/inc_check.sh checks only that the carried engine still computes the right answer, on interp and compiled. The price above is a measurement, not a threshold to guard.
v0.1.439 — 2026-09-06
A command component ends through wasi, and the status it can carry is one bit (Q-114, second half). v0.1.434 gave the ordinary Wasm build a host import for exit; a component has no env host, so it kept trapping. A command component does have a wasi adapter under it, and wasi_snapshot_preview1.proc_exit now ends it.
What that buys is exactly exit 0, and the reason is in the component's own WIT: wasi:cli/exit is exit: func(status: result) -- success or failure, with no number in it. So exit 0 gives 0 and exit 7 gives 1. The 1 is the interface's limit, not this compiler's, and scripts/exit_status_check.sh pins both values with that reason written next to them, so a future wasi that carries a status turns the gate red instead of passing quietly.
Two shapes still trap, and both are honest. A reactor component has no adapter and no process to end. And a program whose main expression IS the exit -- ... in exit 7, whose type is 'a -- is emitted as a reactor rather than a command, because the command shape is only built for unit and int; exit has to be a statement to reach the wasi path.
Poisoned by putting the trap back: both component legs report 134.
v0.1.438 — 2026-09-06
`region R { }` does not return memory on the LLVM backend, and now there is a number for it (Q-116). The backend has no current region: every value allocation names @__lang_default_region however many region blocks are around it, and region R { } is a stack alloca that only explicitly region-typed things go into. It emits no per-type copy-out either -- 42 __mcopy sites on the C backend, none here -- because nothing it holds needed copying out.
test/regionreclaim/pertree.mere builds a tree inside a region block and lets only a scalar out. At 100 iterations and depth 16:
| backend | 40 iterations | 100 iterations | |
|---|---|---|---|
| C | 2.5 MB | 2.5 MB | reclaims (flat in the iteration count) |
| Wasm | — | completes | reclaims (unreclaimed needs ~105 MB of a fixed 64 MiB) |
| LLVM | 127 MB | 316 MB | does not reclaim |
All three print the same answer, which is why the parity suite never saw this: parity compares what a program prints, and this is a difference in what it holds. The documentation was honest by omission -- memory-model.md said "on the C backend" and its backend note named interp and Wasm -- but a limitation nobody states out loud and nobody measures is one a reader will assume away. It is now stated, with the measurement.
scripts/region_reclaim_check.sh asserts C's flatness, Wasm's completion, and LLVM's growth. The LLVM leg fails if the gap closes, printing what to update, so the numbers above cannot go stale quietly. Every comparison is a ratio with a wide band, because peak RSS is quantised.
Not fixed here. Giving the LLVM backend a current region plus a per-type copy-out is the C backend's v0.1.31 arc over again, and it wants to be its own.
v0.1.437 — 2026-09-06
A program whose progress survives being killed. The pieces existed -- contrib/store/kvlog.mere is an append-only store with an fsync per write, and its replay already stops at a record torn half-way through an append -- but nothing had put them together, so "what does a crash cost here" had no answer. test/durable/jobs.mere runs N independent jobs, appends each finished one, and on restart replays the log and skips what is already there. What is durable is the progress, not the computation: a job that was running when the process died runs again from its start.
scripts/durable_check.sh kills it with SIGKILL, repeatedly, and requires three things rather than one. The resumed answer must equal an uninterrupted reference run. At least one attempt must have been killed while it was still working -- a kill that lands after the program finished proves nothing, and without this check a machine fast enough to finish inside the delay would report success while testing nothing. And the same kill schedule, run against the same program with its log turned off, must never finish; if it does, the schedule is not interrupting anything and the first two checks were passing for the wrong reason.
Two poisons: never recording progress, and recording only every other job. Both turn the gate red.
v0.1.436 — 2026-09-06
`contrib/inc` — recompute only what changed. Every program in this repository computes its answer in one pass, which is the right shape for a compiler or a renderer and the shape the region model is built for. The other shape -- a long-lived process asked the same question again after a small edit -- had no library and no measurement. This is the smaller half of one: a cache with a declared dependency graph (inc_declare), an invalidation walk over the reverse edges (inc_touch), and a topological order over the dirty subgraph (inc_order). Dependencies are not traced; tracing needs the engine to call back into the consumer's compute function, which is a different design with a different set of failures.
Two consumers in test/inc/, because one cannot show whether an engine generalises: an aggregate (leaves, groups, root) and a build graph (headers read by many objects, archives, one binary). On 256 leaves and 20 edits the engine recomputes 333 nodes of the 5733 a from-scratch implementation would, and the count is asserted exactly -- one settle over the graph, then three nodes per edit -- rather than as a bound, which would have passed for a range of wrong engines.
The gate carries its own negative controls (scripts/inc_check.sh). Mode 2 stores a new input value and never tells its readers; the gate requires it to differ from the oracle, so the comparison is known to be able to fail. Mode 3 recomputes every node on every edit; the gate requires its output to match the oracle, which is the point -- comparing answers cannot tell an incremental engine from a cache that always misses, and only the count can.
A hole the poison found. Emitting a node before the nodes it reads passed the whole gate. Both consumers declared their graphs bottom-up, so walking the declaration list already produced a correct order and the topological walk was never load-bearing. The build graph now declares top-down -- binary, archives, objects, headers, which is how a build file reads anyway -- and the same poison is caught. The gate ran two poisons; one of them was invisible until a consumer changed.
v0.1.435 — 2026-09-06
`u8x16_first_true` — the lowest non-zero lane, or -1. The u8x16 operation set was derived from one algorithm, the UTF-8 validator, which only ever asks whether anything matched (any_true) and how many (reduce_add). It never needs to know which lane, so nothing returned a position -- and a second consumer immediately does: a literal prefilter, the loop a grep runs before trying the regex, has to know where the candidate is.
Measured on a 2 MiB buffer, finding every occurrence of one byte (best of three, seconds):
| one hit every | scalar | u8x16 + scalar rescan | u8x16 + first_true |
|---|---|---|---|
| 4096 B | 0.69 | 0.10 | 0.08 |
| 256 B | 0.79 | 0.12 | 0.10 |
| 64 B | 0.74 | 0.20 | 0.13 |
| 32 B | 0.77 | 0.31 | 0.19 |
| 16 B | 0.86 | 0.52 | 0.68 |
Without a lane index a matching block costs up to sixteen scalar loads, so the fallback degrades toward the scalar loop exactly as hits get denser -- and at one hit per 16 bytes, where every block matches, the vector work stops paying at all and first_true is the slowest of the three. That last row is the boundary, and it is in the table rather than left out.
The semantics is non-zero, not "high bit set", which is what a raw movemask answers: a lane holding 1 is true here. u8x16_eq produces 0xFF lanes so the common first_true (u8x16_eq a b) is unaffected, but every other producer would have been. C uses two 64-bit words and a trailing-zero count on little-endian targets and a lane loop elsewhere; LLVM compares against zero and counts trailing zeros of the resulting mask; Wasm builds the mask from "lane == 0" and inverts it, because i8x16.bitmask is a high-bit instruction.
Not on the RISC-V backends. The lowering is two RVV instructions (vmsne.vi then vfirst.m), both outside the subset the emulator implements; the backend refuses the name with the message it already gives for a builtin it cannot lower.
test/parity/simd_first_true.mere covers the empty vector, a lane that is non-zero with its high bit clear, lane 0, lane 15, and both outcomes of the eq shape. Parity is 158 programs.
v0.1.434 — 2026-09-06
`exit n` ends the program with n on the Wasm backend too (Q-114). The emission was "evaluate the code, drop it, unreachable" -- there is no process inside a module, so the status had nowhere to go. Every host reports that trap as a failure, which made exit 0 -- a program saying it succeeded -- come back as 1, and turned exit 3 into 1 as well. The status now leaves through a host import (env.exit_proc), emitted only for programs that call exit, so a page whose program never does needs no new import. scripts/run_wasm.js provides it. Component mode keeps the old trap: its env imports are dropped, and routing exit through wasi proc_exit is the second half of Q-114.
What hid it. The parity suite compares stdout, and its failure section compares the status of programs that fail, where the diagnostic is the last line of stdout on Wasm and the first line of stderr elsewhere. A program that exits with a status and no diagnostic fits neither shape. And there was nothing to fit: of 157 parity programs, none called `exit` -- nor does any example or contrib module reachable from a Wasm build, which is why no browser page is affected by the new import. The builtin was simply outside everything the differential suite ran.
scripts/exit_status_check.sh compares the number and nothing else, across interp / C / LLVM / Wasm, for exit 0, exit 7 and a program with no exit at all (so a gate that reported 7 for everything would fail). It counts its own backend runs and fails when a missing toolchain would have left it green while measuring less than it claims. Run against the previous binary it names both regressions.
v0.1.433 — 2026-09-06
A function defined inside a `region` body could not allocate a container on the C backend. region R { ... } binds __region_R as a local of the enclosing C function, and a container constructor in the body names it. An inner fn is lifted to its own top-level C function, where that local does not exist -- and the emitter spliced the name anyway, producing a translation unit clang rejects with use of undeclared identifier '__region_R'. Nine emission sites did it (vec_new, map_new, strbuf_new, bytebuf_new, lb_new, vec_of_bytes, read_file_bytes, file_pread, and the vec-returning byte reader), because the rule was written out at each of them.
It is now written once. region_var_of answers the C local when the region has one in scope and the RUNTIME current region when it does not -- the block makes itself current for its body, so inside the block the two are the same pointer, and when the helper escapes and runs somewhere else the current region is the one actually live at the call, which is the answer a lexical name cannot give.
The trigger was narrow enough to hide: an inner fn that allocates a string was fine (strings already follow the current region), one that allocates nothing was fine, and a top-level fn called from the body was fine (it receives the region at run time). Only a container constructor inside a function defined in the body reached it -- so region R { check (build d) }, the shape the footprint work used throughout, never did. The interpreter and Wasm ran these programs correctly the whole time and LLVM refuses them by name, which made the C backend the only one that answered with a build failure.
&R v now refuses an out-of-scope R instead of splicing it, with the same message LLVM gives: that construct names its arena explicitly, so redirecting it to the current region would answer a question the program did not ask.
test/parity/region_inner_fn_container.mere covers three constructors and a two-level nesting; parity is 157 programs.
v0.1.432 — 2026-09-06
Three more Linux-only CI failures, visible once v0.1.431 let the C gates compile again.
Vectors through heap memory state their alignment (Q-113). A u8x16 captured by a closure or stored in a record is written with store <16 x i8>; with no alignment LLVM assumes the type's natural 16, x86 emits movaps, and the bump allocator's 8-aligned box faults. arm64 does not fault and neither does Rosetta, so the range-version corpus program utf8_simd_small passed on every development machine and on an x86-64 container, and failed only on the CI runner. The LLVM backend now appends align 8 to every vector load and store that does not state one, in one place at the end of emission; a unit test asserts no bare vector memory operation is left.
The loop-safety fixpoint is linear. rv_loop_safe_toplevels looked every top-level name up in a list for every top-level name, so a program of 16000 bindings spent 2.5 s in it and the infer_scaling gate (which bounds type inference at 20x for 8x the bindings) read 35-42x. Hashtables; the gate reads 4x again.
SECTIONS rows for two gated runtime sections. lb_used (the list builder, v0.1.416) and llvm:simd_runtime (v0.1.423) had no row in test/parity/SECTIONS, so section_coverage failed.
`range_version_check` says why a leg failed: exit status, stderr, the first differing lines, and the compiler's complaint.
v0.1.431 — 2026-09-06
The emitted C includes `<stdint.h>`. Since v0.1.419 a boxed variant's tag rides in its pointer, and the X__tag / X__node / X__mk macros cast through uintptr_t. The prelude never included <stdint.h> itself: the SIMD section's <arm_neon.h> (arm64) and <immintrin.h> (x86 with SSSE3) bring it in, so every development machine and an arm64 Linux compiled the output, while the scalar fallback taken on a baseline x86-64 -- the CI runner -- left the typedef undeclared and every C-backend program failed to compile there, taking every CI gate that compiles C with it (parity, ctest, proto_parity, the memu emulator for os_check, ...). CI had been red since v0.1.419; the build matrix stayed green because it compiles no emitted C. Reproduced with clang -U__aarch64__ on the emitted C, and a unit test now asserts the header is in the prelude before the first uintptr_t.
v0.1.430 — 2026-09-05
RISC-V keeps SIMD values in vector registers (Q-112). A u8x16 expression tree is now evaluated in v1..v7 and boxed once, at its root, where before every builtin loaded its operands from their boxes and stored its result into a fresh 16-byte block. Operands that are not vector builtins (boxed variables, calls, scalar arguments) are evaluated first, in source order, onto the stack, so no call runs while a vector value is live in a register. A let-bound u8x16 that is used only as an operand of vector builtins lives in v8..v15 when no call can run between its binding and its last use (a tail call after the last use is fine; a call before it keeps the value boxed). The builtins with a scalar result (u8x16_extract, u8x16_any_true, u8x16_reduce_add) read their operand tree from the registers and never box. The UTF-8 validator's step, which bound seven u8x16 values per 16 bytes, no longer allocates for them: its RV32 listing goes from 43 box stores to 16 (the values carried into the next iteration are call arguments, so they are still boxed), and the 64 KiB validation runs about 10% faster on the Mere-written RV32 and RV64 cores.
The SIMD builtin name tables and the operand-only test moved from the Wasm backend into Ast so both backends share one definition.
v0.1.429 — 2026-09-05
_On the Wasm backend a SIMD value now travels unboxed: a v128 on the stack between operations, a v128 local for a let-bound value that is only ever an operand, and a 16-byte box only when the value becomes a Mere value -- a function argument, a result, a field. Every operation used to box its result._
emit_simd_v leaves a v128 on the stack for any SIMD-valued expression: the lane operations inline (v128.and, i8x16.sub_sat_u, f64x2.mul, ...), the loads and the shifts through _v runtime variants that take and return v128, and anything else -- a parameter, a call's result -- through the box it already is. A let whose value is SIMD-typed and whose name appears only as a direct operand of SIMD builtins (never under a lambda, never anywhere else) gets a v128 local instead of a box. show, to_json, ==-refusal and the interpreter's answers are untouched; the seven SIMD programs in the parity and range-version corpora print the same bytes as before.
What it buys, measured by where the bump arena runs out (nothing is freed on this backend): axpy_simd went from failing at n = 20,000 (a hundred passes) to passing at 50,000; the in-program UTF-8 validator from 64 KiB to 256 KiB (twenty passes). What it does not buy: the per-iteration boxes that remain -- a vector passed to a function or returned from one -- so 1 MiB still runs out. Wrapping the step in region R { ... } was tried and does not rescue it either: the region frees the step's scratch, but the values carried to the next iteration are copied out of it, and those copies are the per-iteration allocation. A v128 that crosses a function boundary without a box would change this backend's value model (every value an i64 slot), and is not this slice.
v0.1.428 — 2026-09-05
_The two bytes builtins the UTF-8 benchmark programs still needed on the RISC-V backends: hex_of_bytes (a new runtime helper) and read_bytes (a file read as bytes is a file read as a str here, so the RV prelude defines it over read_file). benchmarks/utf8valid_simd/bench.mere itself now runs on memu's RV32 core, on a file, and prints the C backend's answer._
Measured on a 64 KiB slice of utf8.txt, twenty passes, --ram 256: both utf8valid and utf8valid_simd print the C backend's valid codepoints 44112 on the emulated CPU (2.75 s and 1.85 s of emulation). The 4 MiB file does not fit: neither backend reclaims here -- the scalar validator's curried cont helper allocates closures per byte and the lane version boxes every vector result -- so a run is bounded by RAM, which is the next thing to fix on this target (keep vector values in registers across an expression, as the Wasm backend now does).
v0.1.427 — 2026-09-05
_A soundness hole in range-check versioning, closed: the loop's EXIT branch runs with the index already outside the guarded range, and its accesses were being made unchecked along with the step's. if i == n then vec_get v i else ... read v[n] unchecked in the compiled fast copy and printed 46 where the interpreter failed. Only the step is rewritten now._
Found writing the arc's retrospective, not by a gate: no case in test/range_version/ had an access in the exit branch, so the on/off comparison that exists exactly for this never saw it. fail/exit_branch_access holds it now, and a unit test reads the fast copy's C to see one checked and one unchecked access. The lesson is the one the gate already knew: the checked loop is the oracle only over the cases someone wrote.
Also: the exit test may carry the stride -- i + 16 > n, i + c >= n, and the mirrored n < i + c forms -- with the offset moved onto the bound, and i > n / i <= n / n >= i are normalised to the >= / < shapes the guard was written for. A sixteen-byte block loop over u8x16_load is planned and dispatched (test/range_version/block_loop_exit); the UTF-8 validator's block loop is planned too (its call site passes computed vectors, so it runs the checked loop, and says so). range_version_check 20/20.
v0.1.426 — 2026-09-05
_The RISC-V backends run u8x16 through the RISC-V Vector extension, and have bytes. The UTF-8 validator written with sixteen-byte lanes now runs on the Mere-written CPU -- test/range_version/utf8_simd_small.mere prints the same answers on memu's RV32 and RV64 cores as the interpreter._
RVV 1.0, the subset memu implements and holds against QEMU (-cpu rv32,v=true,vlen=128): VLEN 128, LMUL 1, SEW e8, e16 only to read a widening reduction. A u8x16 value on these targets is a pointer to a 16-byte box on the bump heap; an operation sets vl and vtype itself, loads its operands into v1 / v2 with vle8.v, computes into v3 and stores it into a fresh box with vse8.v, so no vector register is live across two operations and nothing in the ABI knows about vector state. The mapping: u8x16_and / or / xor are vand / vor / vxor.vv, sub_sat is vssubu.vv, shr is vsrl.vx, eq is vmseq.vv into v0 then vmerge.vim between a zero vector and -1, swizzle is vrgather.vv (an index of 16 or more gives 0, as on Wasm and NEON), shift_in prev cur k is vslidedown.vx by 16-k over prev then vslideup.vx by k over cur, any_true is vredor.vs and a non-zero test, reduce_add is vwredsumu.vs read back under e16, splat is vmv.v.x, extract is vslidedown.vx then vmv.x.s. _start sets mstatus.VS, because a real machine traps every vector instruction while the extension is Off (memu has no such state; on a core without V the bits are WARL zero). f64x2 stays refused by name: these targets have no floating-point unit.
bytes on these targets is the str block -- [len word][bytes], word-padded -- so bytes_of_str and str_of_bytes are the identity and bytes_concat is __str_concat; bytes_len, bytes_get (checked, unsigned compare), bytes_slice and bytes_of_hex are new runtime helpers. The RV prelude's "bytes_of_str needs a host" stub is gone with the reason for it. test/parity/simd_u8x16.mere and the UTF-8 validator case run on both cores under scripts/rv_exec_check.sh (MEMU set) and match the C backend's output. The RV disassembler (mere -rvs, the debugger's listing) names the subset too; it read every vector word as .word, which is where a reader stops.
v0.1.425 — 2026-09-05
_show and to_json of a SIMD value on the C, LLVM and Wasm backends, in the interpreter's spelling -- f64x2(1.5, 1.5), u8x16[abab...], [1.0, 2.5], "0101..." -- replacing the refusal v0.1.422 put at the call site. And a pre-existing divergence found on the way: to_json 1.5 printed null on C and Wasm, 1.5 on interp and LLVM._
The two lanes of an f64x2 go through __lang_str_of_float, the formatter all four backends already share for floats, so the digits agree by construction; a u8x16 is 32 hex digits. C: asprintf / snprintf into a Mere str. LLVM: the lanes extracted, the formats minted like the tuple ones. Wasm: the box loaded, f64x2.extract_lane into the float formatter import, and a 16-step hex writer into a scratch buffer for the bytes. test/parity/simd_show.mere pins all of it, plus to_json 1.5 and show 2.5, on interp / C / LLVM / Wasm.
The float to_json case had no arm in C's and Wasm's per-type generators and fell to the null the unit type prints. Both now call the shared formatter.
v0.1.424 — 2026-09-05
_The u8x16 lane operations a UTF-8 validator needs -- load / from_bytes, and / or / xor, sub_sat, eq, swizzle, shr, shift_in, any_true, reduce_add -- on every backend, and the validator itself: benchmarks/utf8valid_simd (Keiser-Lemire, sixteen bytes per step) against benchmarks/utf8valid (the scalar state machine) and C._
This is the row explicit lanes were for. No vectorizer builds a UTF-8 validator from a byte state machine; three nibble-indexed table lookups, a saturating subtract two and three bytes back, and a compare of the two answers do it sixteen bytes at a time. The operation set was derived from that algorithm and from axpy, not listed by hand, and it is what the stdlib reference now has.
The one operation with no portable spelling is the byte table lookup (u8x16_swizzle): NEON tbl, SSSE3 pshufb, Wasm i8x16.swizzle. C emits all three behind #if, with a scalar loop after them, and normalises pshufb's "index 16..127 wraps" to Wasm's and NEON's "gives 0" by setting bit 7 on such indices first. LLVM IR has no generic swizzle either; there it is a 16-step extract / select / insert the backend pattern-matches where it can. u8x16_shift_in (the previous block's last k bytes in front of the current one) is a constant shufflevector per k on C, and a 32-byte scratch read at offset 16-k on LLVM and Wasm, so k may be a runtime value.
The interpreter fixes the corner rules and the compiled backends match it: sub_sat saturates at 0, eq gives 0xFF per equal lane, swizzle gives 0 outside 0..15, reduce_add sums as integers (4080 for sixteen 255s), a 16-byte load past the end fails like an index. test/parity/simd_u8x16.mere exercises each lane by lane; test/range_version/utf8_simd_small.mere runs the validator on valid text, a truncated sequence, an overlong, a surrogate and a stray byte, on interp / C / LLVM / Wasm. gen_data.py now writes utf8.txt (4 MiB of mixed 1- to 4-byte text, cut at a codepoint boundary) for the two rows.
The number, from benchmarks/run.py (median of three, twenty passes over the 4 MiB of utf8.txt): utf8valid (the scalar Mere state machine) 72 ms, C's scalar machine 54 ms, utf8valid_simd 29 ms -- two and a half times the scalar Mere and nearly twice C (the loop alone, without process startup, is about 3x and 2.5x), with the same answer on every input including the corrupted ones. That is what the two types were added for, and it is the opposite of axpy_simd's result: lanes pay where the vectorizer cannot follow, not where it already does.
v0.1.423 — 2026-09-05
_The f64x2 lane operations -- make, add / sub / mul / div, reduce_add, and load / store of two consecutive lanes of a Vec[R, float] -- on every backend, and range-check versioning extended to a stride and a two-lane access, so an explicit-SIMD axpy has one exit per iteration too._
f64x2_load v i reads lanes [i, i+2) and fails past the end as vec_get does; f64x2_store writes them. C: a memcpy of 16 bytes the compiler turns into one vector load or store (the buffer is only 8-aligned). LLVM: load / store <2 x double>, align 8. Wasm: a float Vec's slots hold POINTERS to 8-byte float boxes, so a lane load gathers two boxes and a lane store boxes two doubles -- correct, and the reason Wasm's number for this row will not be C's until the value model changes. Arithmetic is lane-wise; f64x2_reduce_add is lane 0 + lane 1, in that order, on every backend.
Range-check versioning (v0.1.420) now accepts a self call that advances the index by any one positive literal, and an access of width two: the guard checks [e, e + 2) at both endpoints. A stride above one with an EQUALITY exit needs one more conjunct -- that the loop lands on the bound exactly -- because the original would otherwise run past it and fail on the first access beyond, and the fast copy must not read what the original would have refused. test/range_version/simd_stride2.mere and poison_stride_miss.mere pin both sides; test/parity/simd_f64x2.mere runs the operations on interp / C / LLVM / Wasm.
benchmarks/axpy_simd is the axpy row with explicit lanes, and it is SLOWER: 0.12 s against the auto-vectorized axpy row's 0.08 s and C's 0.07 s. Its loop body is the minimal one (two vector loads, fmul, fadd, one store, the counter) -- the first version reloaded v->data every iteration because the lane accessors went through memcpy, which may alias anything; lane-by-lane double* accesses fixed that and did not move the time. What is left is width: clang vectorizes the scalar loop at VF 2 with an interleave of 4, eight doubles per iteration, and two per step is a quarter of the loads in flight on a memory-bound loop. Unrolling the Mere source by hand to four per step gives 0.10 s. The MANIFEST says so: explicit 128-bit lanes are for the kernels the auto-vectorizer cannot build (the byte-lane ones, next slice), not for beating it where it already works.
Known limit, pre-existing in kind: on Wasm every SIMD value and every lane is a box in the bump arena, so a long SIMD loop exhausts memory where the scalar one (fewer, smaller boxes) still fits -- axpy_simd at n = 20000 x 100 runs out where axpy does not. A register-resident v128 in that backend is the fix, and is not this slice.
v0.1.422 — 2026-09-05
_Two 128-bit SIMD types, f64x2 and u8x16, reach every backend with the four operations that build a value and read a lane back. The lane operations come with the kernels that need them; this slice is the type itself._
Why 128 bits: it is what NEON, SSE2 and Wasm's v128 all have, so a program written against these two types is portable exactly. Why now: range-check versioning (v0.1.420) gets clang to vectorize the loops clang can vectorize; what it cannot -- a UTF-8 validator, a JSON structural scan, anything that is a byte-lane trick -- needs the lanes spelled out, and on Wasm, whose text this compiler assembles with no optimizer in between, there was never another way.
Ast.TySimd (F64x2 | U8x16). C: the compiler's vector_size(16) extension, by value. LLVM: <2 x double> / <16 x i8>, first-class. Wasm: v128, held as a 16-byte box because this backend keeps every value in an i64 slot (as it does floats). Interpreter: V_f64x2 / V_u8x16, the oracle for lane order, for u8x16_splat keeping the low 8 bits, for the lane-range failure and for show (f64x2(1.5, 1.5), u8x16[00..]). RV32IM/RV64IM refuse the type by name (the V extension is Q-110). == and < on a SIMD value are refused by the typer: lane-wise equality is a vector, not a bool, and a whole-vector answer would have to pick one meaning -- so every backend answers the same. f64x2_splat / f64x2_extract / u8x16_splat / u8x16_extract; test/parity/simd_seed.mere runs them on interp / C / LLVM / Wasm. show of a SIMD value on the compiled backends is a clean refusal for now.
v0.1.421 — 2026-09-05
_Range-check versioning reaches matmul's inner loop: an access whose index is affine in the loop parameter -- a[i * n + k], b[k * n + j], v[n - 1 - i] -- qualifies, and the guard checks both endpoints of each such index._
v0.1.420 only versioned an access at exactly the loop parameter. A monotonic function of k over k in [i0, N-1] takes its values between its two endpoints, so for an index that is +, -, * and unary minus over the parameter (once) and invariants, checking 0 <= e(i0) < len and 0 <= e(N-1) < len checks every iteration. % is refused because it is not monotonic; / because the guard would evaluate it before the loop, moving a division by zero ahead of the loop's own effects. test/range_version/affine_index.mere (the matmul shape), affine_reverse.mere and poison_modulo_index.mere pin the three answers; range_version_check 14/14. matmul's dot is now planned and dispatched (its fast copy is inner-lifted three levels deep) and its time does not move: 144 ms against C's 132, the same as before, because the inner loop is a floating-point reduction clang will not reorder and two removed checks vanish under the multiply-adds. The passes of benchmarks/axpy tie C (80 extra passes: 50 ms on both rows); the 10 ms the row carries above C is building the two Vecs by vec_push, so the row now runs a hundred passes and says so.
v0.1.420 — 2026-09-05
_A loop over a Vec checks its whole index range once, before the loop, and runs an unchecked copy when the check passes -- so the loop has one exit and clang vectorizes it. axpy over two Vec[R, float]s went from 2.8x hand-written C to a tie; the checked loop is kept and runs whenever the range check fails, so no program prints or fails differently._
Measured first: the gap between Mere's C and hand-written C on c[i] = c[i] + alpha * a[i] was 2.8x, and SIMD was only 1.3-1.4x of it. The rest was the shape of mere_vec_*_get/set: a bounds check per element is an early exit per access, three of them in that loop, and clang neither vectorizes a loop with several early exits nor hoists the data / len loads that sit behind them. Removing the checks by hand from the emitted C made clang vectorize it and tie C.
Range-check versioning (Ast.range_version_program, Q-108) does that removal soundly. It runs where par_map is lowered, before typing, on every backend at once. A let rec f = fn (i: int) -> ... if i == n then base else step whose step tail-calls f (i + 1) and whose body reads vec_get v i / vec_set v i x / bytes_get b i on loop-invariant containers gets a sibling f__rvfast with those accesses replaced by __vec_get_unchecked / __vec_set_unchecked / __bytes_get_unchecked; each saturated call site with atomic arguments becomes if i >= 0 && i <= n && n <= vec_len v && ... then f__rvfast i .. else f i ... The guard true means no removed check could ever have fired; the guard false runs the original code. The pass gives up -- and the case stays exactly as written -- when the exit bound is not pure arithmetic over invariants, when the body has a lambda or calls anything that is not a builtin, a loop-safe top-level function or a loop-safe local helper (a builtin that takes a function counts as unsafe, derived from the typer's environment), when anything could change a Vec's length, or when a name it relies on is rebound. A call site whose arguments are not atoms, or where the loop's name or the guard's operands are shadowed, is left to the checked loop. MERE_NO_RANGE_VERSION=1 turns the pass off; MERE_RANGE_VERSION_LOG=1 names every planned loop and every dispatched call site on stderr.
The three unchecked builtins exist on every backend; the interpreter keeps the check under them, so it stays the oracle. Two gates: scripts/range_version_check.sh runs test/range_version/ (eleven cases, poison and a mid-loop failure among them) through interp / C / LLVM / Wasm with the pass on and off and requires the pass to have planned exactly what each case's header names; scripts/vectorize_check.sh counts vector arithmetic in the functions reachable from main of the emitted C -- at least one with the pass on, none with it off. "Reachable from main" because the first measurement counted in the function named like the loop and read 0 while the copy clang had inlined was vectorized elsewhere.
Two side findings. exit 0 did not build on the LLVM backend (Q-111): its type is the bottom 'a, the generic closure-call path asked ty_tag for it, and every benchmarks/*/bench.mere ends that way. It has a lowering now. And a shell gate that toggled the pass with VAR=1 shell_function had the assignment persist -- POSIX mode -- and reported "planned: none" everywhere while staying green; the gate now checks what was planned against what the case expects, and toggles through env on the command.
crc32's byte loop is planned too (its helper crc_byte has a local loop of its own, which a loop-safe check that looked only at top-level functions had called unsafe), but the per-byte work dominates and its time does not move. matmul's inner loops index a[i * n + k], not the induction variable, and are not touched: affine indices are the next slice. parity 152/152, range_version_check 11/11, vectorize_check on: 8 / off: 0.
v0.1.419 — 2026-09-05
_A boxed variant node loses its tag word on the C backend: the tag rides in the pointer's low bits. A cons cell is 16 bytes instead of 24, a one-word node 8 instead of 16, and the three benchmarks that are mostly nodes drop by the arithmetic -- binarytrees 175.6 -> 117.1 MB allocated, json 82.2 -> 66.8, a str list 24 -> 16 bytes a word._
Every region allocation is 8-aligned, so a pointer carries three free bits. A recursive variant with at most eight constructors now keeps its tag there: the node is the payload union alone, the value is node | tag, and a nullary constructor is the tag OR'd onto the type's shared static -- no allocation, never NULL. A type with more than eight constructors keeps the old { int tag; union payload; } node. Generated code never spells the layout: three macros emitted beside each boxed type's typedef -- T__tag(v), T__node(v), T__mk(t, p) -- answer the tag, the node, and the value, and the choice is made once, where the type is declared. Construction, pattern tests, show / to_json / == / cmp, the region copiers, len on a list, vec_to_list, the list builder, of_json and the hand-written str list runtime (str_split, str_join, args, read_lines, list_dir, the UTF-8 splitter) all go through them. The copiers also stop re-allocating a nullary node: the static outlives every arena, so it is handed back as it is.
This was the last of the three causes the footprint arc named (the default region never frees; buffers left behind on growth; representation), and the one held open as a design question: whether a 33% denser node was worth touching every site that knew the layout. The sites are now ten, and they all say the same three words. Measured against the other languages' rows: binarytrees at 118.7 MiB resident naive (was 169.1) and json under Ruby's 73.9 MiB once resident numbers are re-recorded. The interpreter, LLVM, Wasm and RV32I backends keep their own layouts -- the parity suite compares what programs print, not how they are laid out, and every one of the 152 programs still prints the same bytes on all four.
Bands: binarytrees [96, 128] MiB, json [48, 96] MiB, wordfreq's floor at 12 MiB -- each floor caught the drop as a breach first, which is what a floor is for. Downstream: mere-ruby, forty thousand lines with many variants past the eight-constructor line, compiles to 19.6 MB of C under the new layout, that C compiles, and its 180-program corpus matches the reference Ruby to the byte -- the old and the packed layouts side by side in one program.
v0.1.418 — 2026-09-05
_A Wasm region block can hand out a list, a tuple, a record, a variant, a float or bytes. Until now only a scalar had ever left one: the copiers for everything else were written for the 4-byte value model and wat2wasm rejected the module._
$__mcopy_<T> is what carries a block's result into the enclosing arena before the bump pointer rolls back. Its tuple, record and variant arms still loaded and stored 4-byte slots at 4-byte offsets, its float arm loaded a double through an i64 and returned an i32, and bytes had no arm at all -- none of it had run since values widened to i64 (v0.1.153), because no test returned a boxed value from a Wasm region block, and the first program that did (the ListBuf parity test, v0.1.416) was rewritten around it and the hole filed as a question. This is the answer: every arm now reads the layouts the emitters actually produce -- 8-byte slots for tuples and records, a 16-byte { tag, payload } node (or a bare 8-byte tag) for variants with the payload handed to its own type's copier, an aligned 8-byte box for floats, and a header-plus-bytes copy through the bytes allocator.
test/parity/region_result_boxed.mere returns each kind from a block that also allocated scratch, reads it afterwards, and matches on c / llvm / wasm against the interpreter. The question's reproducing command (region R { Cons ("a" ++ "b", Nil) }) assembles and prints 1.
v0.1.417 — 2026-09-05
_The README's parity count was two behind, and the gate that exists to notice that noticed it -- in CI, after v0.1.415 had shipped with the number wrong._
first_run_check reads the parity program count out of the README and compares it with test/parity; v0.1.414 and v0.1.416 each added a program and neither release updated the sentence, so the v0.1.415 and v0.1.416 CI runs went red on that step alone. 149 -> 151, and the test count beside it 2617 -> 2620. Nothing else changes. The lesson is the one the gate already states: a number in a README rots silently, and the release that adds a program is the release that has to move it.
v0.1.416 — 2026-09-05
_ListBuf: a list built in order, for one cell per element. The JSON parser stops allocating the 23 MB it used to reverse and throw away._
An immutable list grows at the front, so code that produces one first-to-last has two idioms: recurse without a tail call, which overflows the stack in the tens of thousands, or cons in reverse and reverse at the end. In a bump arena the second pays twice -- every cell of the reversed accumulator is garbage the moment rev returns -- and the attribution table for the JSON benchmark charged 30 MB of a 105 MB run to exactly that. Wrapping the parser's loops in a region was tried and dropped: the copy-out multiplies with nesting depth and doubles the peak on a document that is one container deep.
lb_new / lb_push / lb_to_list (typer, interpreter, C, LLVM, Wasm; the RV32I backend refuses at emit time) append by writing the previous cell's tail. Nothing can observe that, because no cell is visible until lb_to_list hands the head out -- and at that moment the builder freezes. The cells reference the pushed values without copying them, which two runtime refusals make sound: a push after lb_to_list, and a push while a region other than the builder's is current (the cell would point at a value that dies with that region). Both are the same sentence on every backend and catchable with try_or; the interpreter, which has no arenas, keeps a region id so it refuses the same programs. The type carries the region marker like Vec, so a builder cannot leave its region or be stored across one -- two new rows in test/escape/ROUTES, both GATED.
contrib/json builds its arrays and objects with it: 105.2 -> 82.2 MB on the benchmark document, 101.9 -> 80.0 MiB resident, the same bytes out; the row now sits under Go's 90.8 MiB. Parity: test/parity/list_builder.mere (order, empty, strings and tuples, a list of lists, a push from inside a function, the frozen refusal caught, a builder consumed inside a region) and two failing programs under test/parity/fail/ (four backends, one message each). Documented in the stdlib reference, the memory model and patterns 8.4; the host matrix gains the three rows. The self-hosted Wasm codegen (contrib/codegen) learned the three builtins too, since it compiles contrib/json in the test suite -- its Nil and Cons tags are allocated lazily, so they travel to the helper as arguments.
Found on the way, recorded rather than fixed: the Wasm backend's copy-out of a boxed value from a region block ($__mcopy_<T> for tuples, records and variants) still has the 4-byte layout from before the i64 value model, and wat2wasm rejects the module; only scalar results ever worked. The parity test for the builder answers a scalar from its region section for that reason, and the hole is an open question with a reproducing command.
v0.1.415 — 2026-09-05
_mere --suggest-regions says where a region R { } would bound the footprint. It infers nothing and changes nothing: it is the compiler naming the expression, so that the decision lands in the source, where the language wants it visible._
Mere does not reclaim by default, and the design keeps it that way -- a reclamation point the source does not show is the implicit memory management the principles exclude. What was missing was not the construct but knowing where to put it: the binarytrees row went from 169 MiB to 5.5 MiB by wrapping one expression, and finding that expression took reading the allocator. The report walks the typed program every backend starts from and names two shapes, both safe to wrap by construction because what crosses the boundary is a scalar and copies out for nothing: a call, comparison or match whose value is scalar and one of whose operands is a freshly built heap value (check (build d), str_eq (substring line 0 4) "path", match str_split line " " with), and a non-recursive function whose result is scalar and whose body allocates. Recursive bodies are not suggested wholesale -- a region around a tail call nests one open arena per iteration -- so the report points at the expression inside instead, and it stays quiet inside existing regions, in the prelude, and in straight-line top-level code.
Measured on the repository: bench.mere gets two lines, the check (build d) that mattered and a print of a built string; bench_pertree.mere gets only the print. Across 135 contrib files that compile on their own, 822 lines: print of a built string 113, strbuf_push of a built string 99, int_of_str (substring ...) 87, a match on a built value 66, str_eq against a substring 41, map_set with a built key 27. Every one spot-checked was a true statement about garbage; whether the garbage matters is the program's business, which is why this is a report and not a warning. scripts/suggest_check.sh (in CI) holds it to the finding it exists for, and to silence where the region is already written.
v0.1.414 — 2026-09-05
_Containers grow in place. The one-shot benchmarks that were paying twice for every buffer stop paying twice, and the suite gains the two rows that show what a region costs when it is drawn around the right expression._
The cross-language record said Mere's footprint was five to fifty times a Rust program's on allocation-heavy rows while its time was never behind, and attributing the bytes (scripts/alloc_sites.sh, new: every allocation charged to its return address, three frames deep) split that into three causes. The first is the model working as documented -- the default region never frees -- and the answer is to write region where the garbage is: binarytrees with a block around each TREE rather than each batch peaks at 5.5 MiB, the Rust row's footprint, at a third of the C row's wall clock (bench_pertree.mere, one line different). The second was the allocator: vec_push, strbuf_push and bytebuf_push doubled into a fresh buffer and left the old one in the arena, so a container built by pushing cost twice its final size -- matmul allocated 12,583,160 B for three 2 MiB matrices, which is 3 x (2 + 2) MiB to within 248 bytes. __lang_region_grow now extends the buffer where it sits when it is the region's most recent allocation, realloc's it when it lives in a dedicated block of its own (v0.1.307's policy for large allocations), and copies only when it must. matmul 12.6 -> 7.9 MB and 13.5 -> 9.0 MiB resident; the read_file_bytes row 65 -> 36 MB; a 200,000-int Vec 4 -> 2 MiB, which region_slack_check.sh now holds in both directions. The C backend does all three paths, LLVM the in-place one, Wasm the in-place one only while no region block is open -- its blocks are a rollback of the one bump pointer, and extending an outer container's buffer into one would be undone with it. The third cause is representation (a boxed variant node is 24 bytes to Rust's 16) and is a design question left open, with the numbers next to it.
Two library findings from the same tables. contrib/json copied every keyword out of the input to compare it and every number out to parse it -- 7.3 MB of strings on the benchmark document that existed to be read once; it compares and accumulates in place now (112.9 -> 105.2 MB). And the streaming wordfreq row was twice as slow as the one-shot program not because of its 25,000 regions but because file_read_line read a byte at a time through fgetc and malloc'd per line; it is one getline into a per-thread buffer now. The row's ranking is vec_sort rather than list_sort_by inside a region: an immutable list's merge sort allocates about 3n log n cells, and vec_sort is the same stable sort in place with malloc/free scratch -- the named arena's peak went from 15.7 MB to the 1 MiB seed. What was NOT done is also written down: the plan to wrap the JSON parser's object loop in a region was dropped when the copy-out turned out to multiply with nesting depth and double the peak on a document that is one container deep.
Gates: test/parity/vec_grow_inplace.mere (both growth paths, alternating Vecs, a StrBuf push larger than its capacity, string elements; c / llvm / wasm match interp), the vec-growth probe in region_slack_check.sh, the matmul band tightened to 9 MiB, peak_max on the wordfreq row at 2 MiB. Docs: the memory model's growth paragraph, patterns section 8.4 on where to draw a region and what not to materialise, the stdlib reference on the two sorts' allocation shapes.
v0.1.413 — 2026-09-03
_The downstream gate had been green while checking nothing, and when it finally checked something it called a timeout a compile failure. Both are fixed; the gate now refuses to pass on zero._
downstream_check compiles thirteen programs written in Mere -- mere-ruby among them -- with mere -c from inside each checkout, to ask whether a language change broke them. In CI the checkouts come from a clone step that had no if: !cancelled(), so the moment any earlier gate failed, that step was skipped; downstream_check (which does carry the if:) then ran against an empty directory, reported "13 absent, 0 checked" -- and exited 0. That is how the v0.1.410 row was green while mere-ruby was not being compiled by anyone, during exactly the hours other gates were red. The script's own header promised to count absences loudly; it counted them and then said yes. With MERE_DOWNSTREAM set and nothing present, it now fails, in words.
When the clones did arrive (v0.1.412), mere-ruby's 40,000-line main.mere hit the 180-second wall clock on the runner and the gate said "no longer compiles against this compiler" -- bounded.sh's exit 201 was folded into the same sentence as a real refusal. The two are now different sentences, and the bound is 420: the file takes ~68 s to emit C on an M-series laptop and the runner is two to three times slower, so 180 sat on the edge.
Measured while here: v0.1.411's monomorphization fix made that particular compile ~16% slower (58.6 s -> 67.8 s, typer unchanged at 2.1 s, emitted C the same size to within 0.2%, a mid-size program unchanged at 0.07 s) -- more work in the instantiation fixpoint on the one program with hundreds of polymorphic helpers, not more output. Recorded as an observation, not a regression in what is emitted.
v0.1.412 — 2026-09-03
_The three operating-system examples had not compiled since v0.1.367, and no gate knew. They compile, they run, and os_check.sh runs them from now on._
examples/riscv_bare_shell.mere (a two-task kernel with a shell), riscv_bare_user.mere with riscv_user_prog.mere (a kernel that loads a separately compiled, ordinary Mere program as a user process and answers its syscalls) and riscv_bare_selfhost.mere are the programs behind the sentence "an OS on a self-made CPU". Each wrote its timer-interrupt cause as the literal 2147483655 -- 0x80000007, the register's top bit plus seven. When v0.1.367 gave the 32-bit backend a literal check, that number stopped fitting, all three stopped compiling, and nothing said so, because nothing ran them: the QEMU gate runs hello, timer and sched, and rv_exec runs the parity corpus. The literal check was right -- 2147483655 is not a signed 32-bit number -- and the examples were saying the wrong thing: it is a bit pattern. They now build it from its parts, bit_or (bit_shl 1 31) 7, which needs no exemption and says what it means.
scripts/os_check.sh builds the three, runs the shell (banner, the command list, a survived access fault, a clean halt) and the kernel with its user process (five lines and "exited after 9 syscalls") on memu's RV32 core, and compares every line that does not depend on how fast stdin arrives. It sits in CI beside qemu_virt on the same memu checkout. The fix took one line each; the gate is what the fix was missing.
The README also gains pointers to the larger programs written in the language -- mere-ruby, memu, mbrowse -- next to the benchmark table, since a reader who arrives from a talk abstract should find them without knowing the names.
v0.1.411 — 2026-09-03
_A match-pattern that binds the same name as a top-level function was two bugs at once: the monomorphizer attributed the local's type to the function, and the LLVM backend direct-called the function where the local belonged. CI's three red gates are green._
contrib/http/mount's entries binds h in two match arms at two function types -- str -> str from MExact, str list -> str -> str from MPattern -- and the test program also defines a top-level h : str -> str. Three things went wrong with that one name, in order:
The monomorphization scans handled shadowing by a fn parameter and not by a match pattern. `find_concrete_arrow` and the recovery scan walked every arm's body and read the pattern-bound `h`'s arrow as an instantiation of the top-level skeleton `h` -- which is monomorphic -- so the C, LLVM and Wasm backends refused the program ("cannot instantiate `h` at ..."). The scans now skip an arm whose pattern binds the name: unlike a `let`, an arm can never be the polymorphic fn's own definition, so skipping it hides no use site. (A third walker with the same shape, `specialize_single_use_local_fns`, is a plain traversal with no name to shadow and is left alone.) With the refusal gone, the LLVM backend segfaulted. Its direct-call arm was guarded by "the name is a top-level fn" alone, while its Var arm already said "if a local shadows a top-level fn, prefer it." So the MPattern arm's h, a two-argument closure loaded from the payload and then ignored, was emitted as call @mu_h(<the list>) -- the one-argument top-level fn -- and the str it returned was read as a closure struct and jumped through. The guard now also requires the name not be a local, and the call falls through to the closure path. The Wasm backend had the identical guard and the identical bug in a different costume: the one-argument $h read the list as a str and the heap ran away ("out of memory"). Same one-line fix, mirroring its own Var arm's List.assoc_opt name !locals. Interp and the C backend were right all along. Why CI only noticed now. `mount_check.sh` passes when at least two backends run. The C backend had refused this program since the gate was written (2026-08-29); interp and Wasm made two. v0.1.407-409 gave Wasm the shared monomorphizer -- and with it the same refusal -- so the count fell to one and the gate finally said so. Not a regression in the shared pass: the gap was the C backend's, and moving the pass made it universal, and visible.
test/parity/match_pattern_shadows_toplevel.mere pins the shape on all four backends. Also from this CI pass: the six __rv_* internals behind hosted file I/O (v0.1.404) are allow-listed for doc coverage with their portable spellings named, and the README's parity count is 149, as the check that exists because a number in prose rots was right to insist.
v0.1.410 — 2026-09-03
_Pages went red the moment docs/changelog.md passed 1,000,000 bytes: the site generator escapes each document one character per recursion, on the interpreter, whose call-depth cap is exactly 1,000,000._
escape_json_str in contrib/site/build.mere is a textbook tail-recursive loop over a string. The interpreter has no tail-call elimination -- every App is an OCaml recursion behind a depth counter -- and the counter's default of one million is deliberate: a program that deep has already died on every compiled backend, so the interpreter stops early and says so. The changelog crossed that line at 1,014,387 bytes during v0.1.401-409, and eval error: stack overflow (recursion too deep) at loop (i + 1) was the SSG hitting the guard, not a regression -- the same loop overflows identically on v0.1.400.
The guard has a knob for exactly this case, "a program that genuinely wants the host's own ceiling": build_full.sh now runs the SSG with MERE_MAX_DEPTH=10000000. OCaml 5 allocates fiber stacks on the heap, so the depth is real on the CI runner too; the full site builds locally with it.
Two honest notes. The interpreter's lack of TCO is a real limitation and this is a workaround at the one call site that hit it, not a fix of the interpreter. And the 32-bit CPU's count_lets-class scaling pass from v0.1.394 had a fourth casualty found this week by the other session -- try_or's catch record put its s1..s10 save area at word 20..29 of a 15-word record, because (20 + i * 4) was a byte offset whose 20 meant five words -- which is what made mere-ruby boot on the 64-bit machine (v0.1.40x). A value and a size wearing the same literal, one more time.
v0.1.409 — 2026-09-03
_All four compiled backends share one monomorphization pass. Porting the last two found a real wrong answer in Wasm and a real miscompile in LLVM — each one a piece the copy was missing rather than a mistake in it._
The C backend's fixpoint moved to lib/monomorph.ml in v0.1.407 and the RISC-V backend started using it in v0.1.408. LLVM and Wasm still had their own copies, _llvm- and _wasm-suffixed, and reading them side by side is what this slice started from:
| pristine clones | promotion (single → multi) | resolved bodies in the re-scan | recovery pass | |
|---|---|---|---|---|
monomorph.ml (C) | yes | yes | yes | yes |
codegen_llvm.ml | yes | yes | yes | no |
codegen_wasm.ml | no | no | no | yes |
Wasm answered a polymorphic `==` wrongly. Those first three absences are one bug from a user's side: a helper resolved at int and then reached at str through a polymorphic caller that resolves later stayed at int, and the str call got the int body. test/parity/poly_promote_second_type.mere is that program — written from the table above rather than from a symptom, and pinned before anything was ported, because a gate only catches the failure someone thought of. It answered str via poly: MISSED on Wasm where the interpreter, C and LLVM all answered found.
LLVM emitted a call to a function it never defined. No recovery pass meant a referenced-but-never-concretized fn was dropped instead of emitted with its tyvars erased, so examples/plugin/plugins/loops_forever.mere produced IR that clang rejected outright: use of undefined value '@mu_go'. It had presumably never been compiled through LLVM.
Neither copy was wrong; each was missing something the C one had grown. That is what near-identical copies do — they stop being able to drift together.
The instance namer travels with the table. The three ty_tags genuinely differ: C's erases a residual tyvar to int and defaults a container's unresolved region slot to __heap, LLVM's raises on both, Wasm's spells Map differently and had no case for bytes at all. Any of those is a fine symbol name inside one backend; what is not fine is the pass naming a specialization one way while a call site names it another. So Monomorph.inst_table carries the mangle that produced it, and instance_of reads it from there — the disagreement is now unrepresentable rather than merely unlikely. lift_fn_skels takes a ?subset string, because that difference ("... in C subset" / "... in Wasm subset") is a message and not a rule.
One failure got moved rather than fixed, and running it said so. Making Wasm's namer total with deep_erase_tyvars got three softfloat programs emitting again — and produced invalid WAT: undefined function variable $_dc_mul2k__Vec_int_int__int__Vec_int_int. Erasing a region slot to int names one type twice, Vec_int_int where the region never resolved and Vec___heap_int where it did. That is the shape the C backend already had a name for (mpng P5) and a function for: ty_as_tagged, which is now shared and which both spellings go through.
Wasm's ty_tag also gained bytes — simply missing, so it refused to name a type the rest of that backend supports. Two programs whose refusal said "I cannot name this type" now say what actually stops them (channel_* has no Wasm lowering). And that refusal now names the type it is refusing, which is how the missing case was found at all.
All 17 emission differences were classified by name, not counted: 2 real bug fixes, 1 behaviourally identical but more precisely named (LLVM store_watch), 9 measured equal to the interpreter, 2 measured equal to the previous Wasm output (host-dependent, so the interpreter cannot arbitrate), 3 refusals that now name the real limit, and test/mount/cases.mere, which Wasm now refuses — consistent with C and LLVM, which have both refused it all along. Wasm accepted it through blindness, not capability: its own make_spec carried the same refusal and never reached it.
Gates: 2620 unit tests, parity 148/148 (+15 failing-programs = 163/163). C, RV32 and RV64 emission byte-identical to v0.1.408 across all 765 files (2295 emissions, 0 differences), so rv_exec stands at 82/0 at 64 bits and 76/0 at 32.
v0.1.408 — 2026-09-03
_The RISC-V backend monomorphizes. A == inside a polymorphic function compares values there now instead of machine words, graphql_stack_portable leaves the known-different list at both widths, and the pin that asserted the wrong answer is replaced by a test that asserts the right one. Q-102 closed._
A value on this backend is an untagged machine word, and the emitter reads .ty only to choose an instruction sequence. So a == whose operand type was still a type variable at emit time compiled to a word comparison: exact for ints, and a comparison of two HEAP POINTERS for strings and compound values. Two equal strings in different blocks answered false, and the stdlib's list_member missed names that were there.
There is no fix at emit time — there is nothing at run time to ask. The fix is to hand the emitter monomorphic functions, so lib/monomorph.ml (extracted in v0.1.407) grew an AST→AST entry point: specialize_toplevel binds one function per concrete instantiation under its mangled name and rewrites every reference whose use-site type names one. The original polymorphic binding stays, because a reference whose arrow still holds a type variable has to name something, and leaving it means such a program compiles exactly as it did before instead of failing to find a label; an original nothing references is dropped by the backend's own reachability.
The rewrite honours shadowing, which as a tree rewrite is not free: the top-level let chain IS a chain of Lets, so counting those as binders would mark every top-level function as shadowed by its own definition and rewrite nothing — green, and doing nothing. The chain is walked separately from everything under it, where a binder really is one.
The specialization landed on the first try and the pin still printed "pinned". compile_cmp asked "what type is the left operand?" three different ways, and the string branch was the odd spelling: l.Ast.ty = Some Ast.TyStr, a structural comparison. A type variable LINKED to str is Some (TyVar {link = Some TyStr}), which is not equal to Some TyStr. While nothing on this backend had ever been specialized, nothing could produce that shape, so the wrong spelling could not show — it waited for the day it mattered. One lty, asked once, at the top.
A type pass was living inside the layout loop. mere-ruby's RV64 compile went from 4s to 208s, so the cost was instrumented rather than guessed at: the pass was running three times. build_items_sized re-runs the emitter until the jump width and the globals' placement stop moving each other, and the pass sat inside it. Three times the cost is the smaller half. The pass is not idempotent — it unifies type variables in place, so the second run reads a tree the first one made concrete and answers differently: 112 multi-instantiated functions on the first pass over mere-ruby, 500 on the second. And its pristine skeleton clones, the copies that keep every instantiation independently possible, are taken at the start of a run — so on a second run they are clones of an already-fixed skeleton, which is exactly what make_spec refuses when it can see it. It runs once now, in prepare_main, outside the loop. 208s → 63s, which is what the C backend has always paid for the same program (60s); this backend was fast because it was not doing the work.
Nothing in the test suite could see that one. The pre- and post-hoist emissions are byte-identical across all 764 .mere files at both RV widths (954 identical, 574 refused by both, 0 differences) — mere-ruby itself was the only program in reach whose output changed, because nothing smaller is polymorphic enough for a second pass to find more instantiations in.
test/rv/poly_eq_word.mere is gone, along with its bespoke block in the gate. It was a pin whose expected output WAS the wrong answer, so it broke on being fixed, as designed. Its replacement, test/parity/poly_eq_mono.mere, asserts the right answer at str / int / tuple / chained-through-poly instantiations, and lives in test/parity so all six implementations answer it instead of one gate having a special case. A pin says "this is still wrong"; a test says "this is still right", and only the second survives being fixed.
graphql_stack_portable left KNOWN_DIFF and KNOWN_DIFF64 — again not by my judgment but on the freshness check's demand, which failed at both widths with "is in KNOWN_DIFF but now agrees — remove it".
Gates: 2620 unit tests, parity 162/162 (161 + the new test), rv_exec 82/0 at 64 bits and 76/0 at 32 (both +1: the new test, and graphql_stack_portable moving from known-different to passing), known-different down to 2 at 64 and 3 at 32. C, LLVM and Wasm emission byte-identical to v0.1.406 across all 764 files. mere-ruby's corpus on RV64 is unchanged at 142/162 — none of the remaining 20 was this bug, which is what the RV64 arc's own classification of them said.
v0.1.407 — 2026-09-03
_Monomorphization moves out of the C backend into lib/monomorph.ml, unchanged. A refactor is only a refactor if something checks: scripts/codegen_identical.sh compares two mere binaries byte for byte over every .mere file in the tree._
The pass that turns a polymorphic function used at several concrete types into one specialized function per type — the ~265-line fixpoint with its pristine skeleton clones, its promotion of a single-resolved fn when a second type shows up later, and its per-pass re-scan — lived inside codegen_c.ml and was the C backend's alone. The RISC-V backend has none of it, which is why a == inside a polymorphic function compiles to a word comparison there: exact for ints, a POINTER comparison for strings and compound values (Q-102, pinned in test/rv/poly_eq_word.mere). Sharing the pass is the prerequisite for fixing that, and this slice is only the sharing.
Two things had to change for a second backend to be able to call it, and both were the reason "just call codegen_c's" was not available:
- It published into a module global.
resolve_fn_typesreset and filled
Codegen_c.multi_inst_fns, so a RISC-V compilation would have written into the C backend's table — invisible in a one-backend process, and real cross-contamination in one that runs several (the LSP). The table is returned now; each backend holds its own.
- Its refusals were worded as C's.
MonomorphraisesUnsupported/
Error without a backend's name, and the caller supplies the wording — so the C backend's messages are unchanged to the byte while the RISC-V backend can say something true about itself.
mangled_inst_name (and ty_tag, which it is built from) moved with the pass rather than being copied, and the instance-picking rule at call sites — "if this name is multi-instantiated and this use site's type walks to a concrete arrow, name that instance" — is now Monomorph.instance_of, called from all three of the C backend's dispatch sites instead of written out three times.
`scripts/codegen_identical.sh` is the check that this changed nothing. "Behaviour is unchanged" is a claim; emitting every backend's output for every .mere file in the tree with two binaries and comparing bytes is a check. It is stronger than running the behavioural suites and for a different reason than "more files": parity compares what a few dozen programs PRINT, so a change in emitted code that no test's output depends on is invisible to it. 3820 emissions (764 files × 5 backends): 2895 byte-identical, 925 refused by both with the same message, 0 differences. It also caught the one real mistake in this slice — wrapping only the fixpoint left lift_fn_skels's refusal escaping as an uncaught OCaml exception instead of a located diagnostic, on examples/typed.mere. And it was poisoned to prove it can see: perturbing the mangled suffix produced 128 CODE-DIFFs, which also says the corpus exercises the moved code rather than stepping around it.
Gates: 2620 unit tests, parity 161/161, rv_exec 80/0 at 64 bits and 75/0 at 32.
v0.1.406 — 2026-09-03
_StrBuf is a real byte buffer, and the prelude's string builders and splitter stop being quadratic in total allocation. str_edges agrees with the C backend at 64 bits and leaves the known-different list._
The StrBuf a program gets on the RISC-V backend was a one-word cell whose push REPLACED the held string with its concatenation — simple, correct, and O(n²) in total allocation. On a bump allocator that never frees, that is not a slowdown but a death: test/parity/str_edges builds a 200KB string with str_repeat, which allocated ~2GB of dead intermediates and walked the heap into the stack with 128MB of RAM. That was str_edges's ACTUAL 64-bit difference all along — the list's stated reason ("64-bit values; this backend's int is 32 bits") was the 32-bit column's reason, wearing the other column's name.
StrBuf is [len][cap][dataptr] now, bytes, doubling on growth, amortized O(1) push; to_str materializes a fresh block and the buffer stays usable. The prelude's builders — str_repeat, str_rev, to_upper, to_lower — go through it.
str_split was the same quadratic in a different dress: it took substring s start (str_len s) PER PIECE — the whole remaining tail — so splitting that 200KB string into 20001 pieces copied ~2GB of tails. It finds from an offset now, without materializing anything but the pieces. str_replace had BOTH shapes at once (tail copies and acc ++ growth) and goes through the buffer and the offset finder together.
str_edges at 64 bits is now byte-identical to the C backend and is out of KNOWN_DIFF64. And out of the 32-bit list too — not by my judgment but by the freshness check's: the run after the fix failed with "str_edges is in KNOWN_DIFF but now agrees — remove it", meaning the 32-bit entry's stated reason (big-integer lines) had gone stale underneath it at some point. That check exists for exactly this, and this is the third time today it has caught a list outliving its reasons.
Gates: 2620 unit tests, parity 161/161, rv_exec 80/0 at 64 bits and 75/0 at 32 (str_edges counts as passing in both columns now), rv_float 2880 identical, rv_prelude 127 names, softfloat 102 names, qemu_virt 7/7, rvd_oracle 7/0.
v0.1.405 — 2026-09-03
_The float library is computed on the RISC-V backend: sqrt correctly rounded through softfloat's integer digit-by-digit root, the transcendentals in the prelude with measured, documented accuracy, and the exact ones exact. Thirteen functions stop stopping._
sqrt lives in contrib/softfloat (sqrt.mere) because it can hold that library's bar: the root of a non-square is irrational, so the true value is never exactly halfway between two doubles, and 56 computed bits plus a nonzero-remainder sticky decide every rounding. Restoring digit-by-digit like div's long division; the radicand (a 53-bit significand shifted up 58) never exists as a value -- the loop consumes two implicit bits per round and its live values stay under 2^58, inside L6's 90 bits. An odd exponent donates a doubling to the significand first, parity taken WITHOUT trusting % on a negative int. test/float/softfloat_sqrt.mere holds it to the hardware bit for bit -- exact squares, both subnormal edges, the extremes, sqrt(-0) = -0, and golden-ratio sweeps -- and the narrow probe now reaches fsqrt so the RV32 arm actually compiles it.
exp, log, f_pow, sin, cos, tan and atan2 are prelude Mere over the four softfloat operators, with accuracy MEASURED against libm on 10k-point sweeps and written where the code is: exp <= 1 ulp, log <= 2, atan2 <= 3, sin/cos <= ~10 and tan <= ~14 for |x| <= 1.6e6 (the three-piece pi/2 reduction's exact range -- the reduced angle is a HEAD+TAIL pair, so a near-zero crossing of a large argument keeps its accuracy instead of dying at the double's edge). f_pow routes y = +/-0.5 through the correctly rounded sqrt (9 0.5 must print 3.0, and exp(0.5 log 9) is two ulps shy), takes small integer exponents by binary exponentiation (exact where the result is exact), and otherwise runs exp(y log x) with the product done exactly -- Dekker split -- over a RENORMALIZED two-piece log, after the first version treated the log series' tail (a real 3.4e-3 of the VALUE) as a rounding crumb and answered 0.9^-700 one percent wrong.
floor, ceil, round, f_min and f_max are exact: bit surgery and compares, written width-safe (no literal at or above 2^31, sign bits made by shifting, low bits cleared by shift-down-shift-up, which is fill-agnostic). round is half-away-from-zero decided on the true fraction, not on x + 0.5, which can round up in float and push a value below the half over it. f_min/f_max are the C backend's exact ternary, NOT libm's fmin -- those disagree about NaN. pi and e are prelude constants now too.
The prelude gate's job reversed here: it used to assert the float shims' stop message survived, and now asserts sqrt does NOT carry one -- the stub coming back would be the regression. host_matrix flips 13 float cells and 4 file cells from stub to yes, and the C backend's refusal for the raw __rv_ names now attributes every one of them, so the matrix's probe can name its subject instead of reporting `unattributed`.
Measured end to end: mere-ruby's corpus as files on the 64-bit machine is 142 of 162 identical, up from 137 -- Math.sqrt, x 0.5, the numeric tower's pi and Float classification all went green. rv_float_check runs the whole 2880-line softfloat comparison ON the emulator against the machine's own doubles: identical.
Gates: 2620 unit tests, parity 161/161, rv_exec 79/0 at 64 bits and 74/0 at 32, read_file 3/0, write_file/read_stdin 3/0, rv_float 2880 lines identical, qemu_virt 7/7, rvd_oracle 7/0, rv_prelude 127 names, softfloat 102 names.
v0.1.404 — 2026-09-03
_write_file and read_stdin on the hosted RISC-V target — the write half of what v0.1.402's read_file started — and substring's range refusal now names the range and the length on this backend too._
write_file is openat(O_WRONLY|O_CREAT|O_TRUNC, 0644), a short-write-safe loop, and close, all Linux-numbered; the failure is a catchable fail whose negative return IS the errno. read_stdin is the same slurp read_file uses, pointed at fd 0 — the emulator serves read(0) from its own stdin, once, and a second read_stdin answers "" exactly as a drained process stdin would.
The emulator buffers a write fd and hands the bytes to the host at close. Two of its choices were forced by failures on the way in. The open VALIDATES the path by creating the empty file right there — which is what O_CREAT|O_TRUNC means anyway — because a bad path discovered at close has nowhere to send its bytes, and the uncaught host-side fail took the EMULATOR down: the guest's bad path reported as our crash, and the guest's try_or never saw anything. And the close swallows late host errors through try_or for the same reason, which is also the oracle's own shape: C's write_file checks the open and discards fwrite's count and fclose's answer.
--bare refuses both by name, and the refusal reason is part of the gate, not just the refusal (host_read_file_bare taught that: a missing top-level main refuses everything and proves nothing). read_stdin's --bare message no longer blames "the filesystem" — there is no host to read from, and stdin is not a filesystem.
substring's out-of-range message matches the C backend's now: substring: range [20005, 3) invalid for str of length 20008 where this backend used to say "range out of bounds" — which names neither the argument that was nonsense nor the length it exceeded. The check moved into the prelude (three compares and a concat are Mere's job; the assembly helper keeps its unsigned check as a backstop), and the raw slice is reachable only through the private name the wrapper calls.
The host-services comment in the prelude was rewritten around what remains, because its list had gone stale — it still said there was no host to read a file from, one screen above read_file doing exactly that. run stays refused PERMANENTLY, with its reason stated: there is no shell on the other side of an ecall, and an emulator that faked one by running commands itself, on the host, with its own privileges, would make every guest program a host program.
test/rv/host_write_file.mere gates the pair against the C backend at both widths, in the harness's own cwd, with the SAME stdin piped to both sides — feeding them differently would report a backend difference that is really a harness difference. Covered: write-then-read-back through the independent read path, a shorter second write (O_TRUNC, where append-by-accident shows up), NUL bytes, the empty write, a write past one 4096-byte chunk, a catchable miss on a bad directory, and the drained second read_stdin.
Gates: 2620 unit tests, parity 161/161, rv_exec 79/0 at 64 bits and 74/0 at 32, read_file 3/0, write_file/read_stdin 3/0, qemu_virt 7/7, rvd_oracle 7/0.
v0.1.403 — 2026-09-03
_map_recycle on the RISC-V backend kept the entries it was contracted to clear, so every pooled call frame in mere-ruby arrived still holding the previous call's locals — a bare identifier in a Ruby method could resolve to another method's variable. One line, and the last silently-wrong value on the 64-bit machine goes away._
map_recycle m is map_clear plus give-the-arena-back (v0.1.300). The RV map is an assoc list with no arena behind it, and the no-op it had was reasoned from that half — nothing to wind back — to skipping the half it CAN do. The reasoning even sat next to a correct map_clear in the same file. What it cost surfaced a long way off: mere-ruby pools its call frames and cleans a dead one with a single map_recycle, so on this backend every frame from the pool still held the previous call's bindings. def f(v); q = v * 10; q; end left q visible to the next method; a Comparable's n <=> o.n read both sides from the same leaked slot and answered 0 where every other backend answers -1. Nothing crashed. The values were plausible; they were merely someone else's.
The diagnosis order is worth recording. The symptom was <=> answering 0, and <=> was innocent: method lookup agreed with the C backend on everything (respond_to?, instance_methods, method(:x).arity), send reproduced it, [n, o.n] came back [2, 2], and only BARE identifiers were wrong — @n and self.n were right. That named "a leaked method-local shadowing the reader", and a two-method reproducer confirmed locals surviving between invocations. map_new was checked and is genuinely fresh per call, which left the pool, and the pool cleans with exactly one call: map_recycle.
test/parity/map_recycle_clears.mere pins the clear half on every backend, including reuse-after-recycle — a pool does not just clear a map, it hands it to a DIFFERENT call that trusts it to be empty. Compiled with the no-op, the pin answers LEAKED q at both widths; nothing else in the corpus ever recycled a map and then asked it anything, which is how a no-op survived under a suite this size.
Also here: rv_exec_check.sh states its denominator. Programs the backend refuses at compile time were dropped with a bare continue, so "73 passed" read as coverage of the corpus when it was really coverage of the roughly half this backend accepts — 70 of 144 at 32 bits and 62 at 64 were skipped, invisibly. The summary now counts them and names them. A gate that does not state its denominator claims the whole corpus every time it goes green.
mere-ruby's corpus, run as files on the 64-bit machine: 137 of 162 identical, up from 135 — corpus 148, the silently-wrong comparison, is one of the two that went green. The remaining 25 differences are host services this target still refuses, float transcendentals, and emulated speed — every one a refusal or a documented gap, none a wrong value.
Gates: 2620 unit tests, parity 161/161, rv_exec 79/0 at 64 bits and 74/0 at 32 (read_file 3/0), qemu_virt and rvd_oracle unchanged from v0.1.402.
v0.1.402 — 2026-09-03
_read_file and file_exists work on the hosted RISC-V target, so the interpreter this backend carries can finally be handed a script instead of only -e. Getting there found a worse bug: a one-argument top-level let rec compiled to real recursion, and only ran because clang chose to flatten it._
read_file, file_exists, write_file, run and read_stdin were all answered with a stop, on the reasoning that --bare hands the program the machine and there is no host behind it. That is true of --bare and was never true of the hosted target, which already asks the host to write and to exit through the same Linux-numbered ecalls. The cost was specific: mere-ruby could only ever be given a program with -e, because reading a .rb off disk went through read_file — so its own 162-program corpus could not be run on the machine at all, and the only end-to-end evidence was hand-written one-liners.
Two of them are answered now, through openat(56), read(63), close(57) and faccessat(48) — Linux's own numbers, so a binary built this way is servable by anything that speaks the ABI and not only by the emulator in this tree. The failure is a fail and therefore catchable, which is what an interpreter needs in order to raise Errno::ENOENT of its own; the negative return IS the errno, so the message says which. --bare still refuses all of them BY NAME, and test/rv/host_read_file_bare.mere exists to check that it refuses for that reason: the first version of that check ran the hosted program under --bare and passed on "the program needs a top-level main", reading a refusal it had caused itself and calling it evidence.
The bug underneath
Adding the syscall service to memu's emulator turned the whole RV corpus red at both widths, with the emulator's own "stack overflow (recursion too deep)". That reads exactly like a regression in the compiler under test, and it was not one: the same prog.bin ran on the previous emulator binary.
The emulator's instruction loop is let rec drive = fn n -> ... drive (n + 1), so its recursion depth is the emulated instruction COUNT. It had never been a loop. The tail-call transform (Q-029) is reachable only through the direct N-ary form — the plain call path never consults self_tail_goto — and that form was gated on params >= 2, because one parameter has nothing to uncurry. The two facts composed: every one-argument top-level let rec in every program compiled to real recursion, one C frame per iteration. They ran anyway, because clang's sibling-call optimisation flattened them, and adding code to the same translation unit was enough to change its mind.
So the gate is >= 1 now. The transform is a promise the language makes, and the optimiser was the thing deciding whether it was kept. This is a codegen_c table, so the blast radius is the C backend: an RV binary is byte-for-byte the same size as before.
It is pinned on the EMITTED CODE — three unit tests asserting goto __mere_tail for one- and two-argument shapes — and not only by running something, because at -O2 clang flattens the recursion anyway: a program that loops past any stack still passes with the transform missing. test/parity/tail_loop_one_arg.mere loops five million times as the runtime half, and poisoning it (MERE_NO_TAIL_LOOP=1, -O0) kills it, which is how the pin is known to be a pin.
Five unit-test expectations moved, all of them incidental to this change: four call sites now name mu_f__direct, and a -g line-directive count went from 1 to 2 because a one-argument function's body is emitted twice. Each was rewritten to assert what it means rather than what it happened to spell — "the user's run is called, not the shell-exec builtin" is now mu_run__direct(__da0).
What the corpus says now
Running mere-ruby's corpus through -e against the C backend: 124 of 162 identical. The 38 are host services this target still refuses (Dir, ENV, require, sockets, Marshal, zlib, OpenSSL, YAML, Date), float transcendentals, map and ivar iteration order, and RAM sizing — plus one that is none of those and matters more than the rest: a user-defined <=> answers 0 where every other backend answers -1. Traced, and it is not <=>: METHOD-LOCAL VARIABLES LEAK BETWEEN INVOCATIONS on this backend. def f(v); q = v * 10; q; end followed by a method that never binds q finds q still holding 30. map_new is genuinely fresh per call, so the leak is further in. Recorded as the next investigation rather than guessed at.
And one of the 38 turned out to be small
A map iterates each distinct key once, in the order the keys were FIRST inserted, with that key's most recent value -- on every backend except this one. The RV map is an assoc list that map_set prepends to, and its iterator walked from the head, so it visited keys in REVERSE insertion order. Every value was correct, which is why it survived: map_get, map_has and map_len are all order-blind, and only a program that prints a WHOLE map can see it. mere-ruby is one, because ruby's Hash is insertion-ordered and its inspect shows it, and the symptom was {"MRB_B"=>"2", "MRB_A"=>"1"} for a hash built the other way round.
Reversing the walk alone would have been the trade for the worse -- the right order with each key paired with its OLDEST value -- so the walk goes over the reversed list while the value still comes from the head end. test/parity/map_iter_order.mere pins all three shapes: plain order, a re-set key keeping its place while taking the new value, and a deleted-then-reinserted key going to the end, which is what separates insertion order from sorted order and from whatever slot a key happened to fall in.
Gates: 2620 unit tests, parity 160/160, rv_exec 78/0 at 64 bits and 73/0 at 32 (plus 3/0 for read_file across both widths and the --bare refusal), qemu_virt 7/7, rvd_oracle 7/0.
v0.1.401 — 2026-09-02
_A try_or's catch record saved the catcher's ten callee-saved registers 15 words PAST the end of the record -- onto bump space the failing thunk then allocated for itself -- so a caught failure handed the catcher its named bindings back as the thunk's garbage. mere-ruby now boots on the 64-bit machine._
The record is [prev][sp][fp][catch][default][s1..s10] and is allocated with alloc_words 15: five header words plus one per saved register. v0.1.388 introduced the save as enc_s (20 + i * 4) -- byte 20, words 5..14, exactly filling those 15 words. The width parameterization in v0.1.394 rewrote it as (20 + i) * wsz (), and the byte offset 20 became the word index 20. The whole save area moved to words 20..29 of a 15-word record.
What that costs is a matter of ordering. The record is allocated; the registers are written 5..14 words above the new heap top; then default is evaluated and the thunk is called, and every allocation either makes walks over the save area. On the unwind, the catcher's names are restored from whatever the thunk last put there. env in mere-ruby's run_src lives in s10 -- the deepest slot -- so it took roughly fifteen words of allocation to reach, and came back as 1.
Two earlier investigations had ruled out the frame (v0.1.398's slot invariant reports zero overflow) and the ABI (a shadow-stack check over every JAL/JALR found no callee failing to restore s1-s10). Both negatives were correct and neither could see this: the save is heap-relative, not frame-relative, so no fp-slave store ever carries the value; and an unwind is a longjmp, not a return, so a function writing its OWN s10 from its OWN code is behaving perfectly legally. The thing that answered in one line was the disassembler: lv_set env "self" begins mv a0, s10, which settles "register or memory" without tracing a single address.
The layout is now declared once (tor_prev .. tor_sreg), the allocation is DERIVED from it (tor_words = 5 + Array.length sregs), and tor_off refuses to emit an offset outside the record -- so the save side and the restore side cannot spell it differently, and the size cannot drift from the layout because it IS the layout. The check was added BEFORE the fix and confirmed to fire on both widths; a check that has never fired is not known to be a check.
This was broken at 32 bits too, from v0.1.394 on. Every strided offset expression that release rewrote was re-derived from the diff: i * 4, (i+1)*4, (sreg_base+k)*4, (total+2)*4, (ra_slot+1)*4, fsz+(i-8)*4 and 32*4 all had word-index addends and are all correct. 20 + i * 4 was the only one whose addend was a byte offset. The defect class has one member and it is closed -- enumerated, not guessed.
test/parity/try_or_saves_bindings.mere pins "the unwind restores the registers" and stayed green through all of this: its thunk allocates a few words, the save area was fifteen away, and the clobber never reached the registers it used. Restoring from the wrong place and restoring from a place nothing has overwritten yet print the same thing. test/parity/try_or_save_area.mere closes that: the catcher holds a live binding in every one of the ten registers -- strings, so a clobbered pointer cannot read as a plausible number, and integers at distinct powers of two, so the sum names which register was lost -- and the thunk allocates past all ten slots before failing. A third case gives default an allocation of its own, since it is evaluated after the save. Compiled with the old offsets, it traps at both widths.
A side effect worth naming: the writes are now inside the allocation, so alloc_words's own out-of-memory check covers them. Before, near the heap top, the save area could be written past the checked limit.
Also in this slice, mere -rvd learned to read RV64, and got a gate. It printed ? for f3=3 -- LD and SD, which are most of a 64-bit binary, since every slot, cell and field access is one of them -- and did not decode opcodes 0x1B or 0x3B at all, so every addiw/addw/mulw/divw fell through to .word. Shift amounts are read as 6 bits now and srli/srai is told apart by bit 30, which is right at both widths (srli x, y, 32 has f7 = 1 on RV64 and used to print as srai -- worse than a ?, because it is confidently wrong). Unreadable lines in a mere-ruby listing: 41.4% before, 11.6% after, and the remainder is the rodata blob rather than instructions. The listing being unreadable is why this bug had to be found by hand-decoding the same words the disassembler exists to print.
That it had no gate is why it drifted. The backend has three -- parity compares outputs, rv_exec runs the code, qemu_virt cross-checks the bytes against an emulator nobody here wrote -- and all three are about what the compiler EMITS. Nothing checked what we PRINT of it, because a disassembler is not part of any program's behaviour. scripts/rvd_oracle_check.sh closes that with the same argument qemu_virt makes: the oracle is riscv64-elf-objdump, and a decoder we wrote agreeing with itself proves nothing about an encoding. It compares every instruction in the code region of four real binaries at both widths -- 1.6k to 21k instructions each, mnemonics over everything and full operands over the load/store family, since a mnemonic-only check passes a listing that prints the right instruction with the wrong offset, and an offset is what this slice was about. The code region's end is read from the debug map's first string block rather than estimated: a guess reached into rodata and the oracle dutifully named fmadd.s for the one-character literals "C", "G" and "K", whose ASCII codes are also the F-extension opcodes. Run against the pre-fix disassembler, all four 64-bit rows fail and name the LD/SD gap; all three 32-bit rows pass, as they must, so the summary line says which widths and which cases actually ran.
mere-ruby runs Ruby on the Mere-written 64-bit CPU. Blocks, hashes, begin/rescue with live locals across the rescue, recursion, bignums and 64-bit integer results all work; Math.sqrt does not, and that is the RV backend's pre-existing compile-time refusal (softfloat covers the four operators, the comparisons and int conversion, at both widths), delivered to the program as a catchable failure exactly as designed.
Gates: 2617 unit tests, parity 158/158, rv_exec 76/0 at 64 bits and 71/0 at 32, qemu_virt 7/7, rvd_oracle 7/0.
v0.1.400 — 2026-09-02
_The package lock's content hash included .git, so the same pinned revision hashed to a different value on every clone -- the Pages build could never satisfy its own integrity check._
mere install clones a dependency, cp -Rs the checkout (.git and all) into the module dir, and hashes that tree to record in mere.lock. The hash walked .git too: pack files, FETCH_HEAD, the index, and reflogs, all carrying per-clone timestamps and per-machine packing order. So a coordinate that pins an immutable SHA still produced a6a6ddab here, 94aad445 on the next install, and 593c448a on the CI runner -- and the integrity check, whose whole job is that a pinned rev reproduces the same content, could not pass. Pages had been red on exactly this since the check was added.
walk_rel now skips .git. The hash is over the working tree the coordinate actually pins, which is reproducible: stable across three re-installs here, and now the same bytes CI checks out. mere.lock is re-pinned to the content hash (4b9a3a99); the rev is unchanged. The docs site (mere SSG + every wat2wasm target) builds with it.
A hash meant to prove reproducibility that is itself not reproducible is worse than none -- it fails the honest case. Excluding the one directory that is machine state, not content, is what makes the guarantee true.
v0.1.399 — 2026-09-02
_CI had been red since v0.1.357 -- four gates that only fail on Linux, three of them added during the red stretch and so never once green there. The macOS dev machine hid every one._
doc coverage -- this session's eight new backend builtins were in no doc. `raw_peekw` / `raw_pokew` are real bare-metal accessors and now sit beside `raw_peek8` in the stdlib reference; the six `__rv_ internals (what args, time, random_int lower to on -rv, plus the two diagnostics) go in docs/UNDOCUMENTED_ALLOW with the reason a program must never name them. * **first_run_check** -- the README's "139 parity programs" was three behind the 142 on disk. The check exists precisely because a number in prose rots; it was right, the prose was stale. * **run status** -- a CPU-limit kill is SIGXCPU (152) on macOS but SIGKILL (137) on the CI container, where the CPU cap makes the soft limit the hard one. The gate is about interp and C AGREEING on how a child died, and they do -- both 137 there. The exact signal is the platform's, so the two cpu-limit rows now accept any signal death (>128, equal on both backends); the self-kill -9 rows stay exact at 137. * **softfloat** -- the sign and payload of a NaN that an OPERATION produces are unspecified by IEEE 754 and differ by platform: 0.0/0.0 is +nan on ARM/macOS and -nan on x86/Linux. The add/mul/div gates compared those bit-for-bit, so they tested the platform, not the arithmetic. When both hardware and the library return a NaN, that is now agreement. The round-trip gate stays exact -- there no operation makes the NaN, a bit pattern goes in and must come back. And str_of_float's NaN string ("-nan" on glibc, "nan" on macOS, "nan" from our exact printer) is compared sign-insensitively in the decimal gate, for the same reason.
Nothing in the compiler changed; these are the gates learning what is portable and what is the host's. The macOS-only-green trap is one this repo has a memory about, now paid down for the float and process-status gates.
v0.1.398 — 2026-09-02
_count_lets undercounted region bodies, and now a hard invariant proves it can never again: a frame's reserved slots must equal the slots its body hands out._
count_lets predicts how many binding slots a function needs; the frame is sized from that; the body then hands out slots as it compiles. The two are separate walks of the same tree, and count_lets was missing Region_block and Region_loop -- it fell through to _ -> 0 for exactly the nodes vars_in and free_vars_of both recurse into. Three walkers, one hole.
A function with bindings inside a region therefore sized its frame too small. Its highest slots landed BELOW its own stack pointer, in the region a called function's frame writes into. At 32 bits it happened to miss anything live; at 64 bits mere-ruby's gc_collect reserved 2 slots and used 17, and the overflow landed on a live binding.
This was not deduced -- it was measured. A temporary check compared !slot_ctr against total at the end of every function: it named gc_collect (over by 15), run_while_top (by 4) and run_dowhile (by 5), and after the fix it reported zero across the whole parity corpus, the rv corpus, and mere-ruby, at both widths. That check is now permanent and a HARD error: count_lets and the slot walk must agree at every compile, because a frame sized too small is silent corruption a long way from its cause, and the two walks drifting apart is exactly how it happens.
This is a real 64-bit correctness fix -- any program that calls a region-containing function with live bindings in the caller was exposed -- and the invariant is the confirmation, not a comment claiming one. The 32-bit suite, both parity columns and the softfloat/prelude gates all pass with the hard check in place.
_(mere-ruby's own boot on the 64-bit machine is still not through: its env binding reads as 1 at first use, and that is NOT this bug -- with the fix and zero slot-overflow, env is neither stored to a frame slot nor clobbered in a register, which points at a save/restore path and is the next investigation. Separating it cleanly from the count_lets undercount, with the invariant now guarding that whole class, is this version's result.)_
v0.1.397 — 2026-09-02
_75 of 79 runnable corpus programs agree at 64 bits, and rv_exec_check runs both widths now. Five more bugs, and the most instructive pair had been unreachable at 32 for the same reason they were bugs at all._
`chr`, `vec_get`, `vec_set`, `char_at` and `substring` never checked their bounds on this backend. The test that exists to see exactly that (index_edges, and prop_int's randomized `chr`) uses literals too wide to compile at 32 bits -- so the first machine that could run the test was the 64-bit one, where `chr` of a sixty-bit number quietly stored the low byte and `vec_get v (-1)` read the word below the buffer. All five refuse now, with unsigned comparisons so a negative index is one huge index and the same refusal. (And the first version of the checks branched with the operands swapped -- `bltu len, i` -- which refused everything IN bounds; index_edges said so on the next run.) `__strbuf_new` and `__strbuf_push` had frames spelled `8` -- two words at the old width. At 64 the push frame moved sp by 8 and stored ra at sp+8: one word ABOVE its own frame, in the caller's stack. Every strbuf call corrupted its caller. This is the third member of the value-vs-size family, wearing a different literal. Same family, one more: `emit_variant_eq`'s three-word frame was `-12`.
The 32-bit column is untouched (2617 unit tests, parity 142, rv_exec 70/5 all green), and the four differences left at 64 all have standing causes, named in the script: the word-comparison ==, the arena that does not exist, the RAM the sweep does not grant, and the prelude's concat-quadratic string builders -- that last one now recorded as real work (strbuf wants to be an actual byte buffer; today it is str_concat in a loop on every push).
Still open, with evidence attached: mere-ruby's boot on the 64-bit machine. lv_owner receives env = 1 -- a boolean where an environment belongs -- on its very first call, so the corruption is upstream of everything probed so far, and it is not any of: the strbuf frames (fixed), the map runtime (75 corpus programs exercise it), or the tuple-destructure shape (reproduced clean in isolation). That is the opening question of the next session, instruments in hand.
v0.1.396 — 2026-09-02
_Hosted RV64: the parity corpus runs on a 64-bit CPU written in Mere, 67 programs agreeing with the C backend -- including the 64-bit-value tests the 32-bit machine could never hold._
memu grew rv64i_run.mere: the same decode structure, device map, syscall numbers and halt reasons as the 32-bit core, at a width where the HOST int and the register width coincide -- headroom zero. Unsigned comparison is (a xor minint) < (b xor minint); logical right shift must clear what the host's arithmetic shift dragged in; and the file is compiled-only, because a register with bit 63 meaningful does not fit the interpreter's 63-bit int -- the reverse of the interp-only trap, stated in the header.
Three compiler bugs surfaced, each caught by a differential:
`li`'s 32-bit arm was wrong for values just under 2^31. The guard was on v; the truth is on HI: lui's 20-bit immediate is sign-extended, and a v in [2^31-2048, 2^31) rounds hi up to 0x80000 -- out of signed-20 range. `li 2147483645` materialised -2147483651, and `2147483647 / 5` had a negative numerator before the divide ran. Found because print_int of a literal disagreed with print_int of the same value computed -- and it had made `random_int`'s rejection loop spin forever, every draw rejected against a negative limit. The shift builtins' saturation points were 32's. bit_shl x 40 compiled to li 0, on the machine where it is an ordinary shift. The bounds are the width's now, constant and dynamic counts both. `__rv_clock`'s tuple meant rv32's timespec layout. Linux's timespec64 is four 32-bit cells there and two native words on rv64, so the prelude's `time` branches once on the new `__rv_xlen` -- a compile-time constant, and the one piece of width the prelude ever needs to see.
Also learned twice over: a Mere program that ends in a bare expression has its final value echoed by the C runtime, and an emulator must exit 0 explicitly or every guest's output grows a trailing () -- the 32-bit core's closing comment says exactly this, and the 64-bit core now ends the same way.
Left named for next time: mere-ruby's own boot traps inside rvmap_has on the 64-bit machine (s-registers holding what looks like string BYTES loaded as words), and five smaller corpus differences (index_edges, prop_int, strbuf, url_percent, capture_after_call, str_edges) beyond the three with standing causes. The 64-bit sweep is manual until those settle; then it joins rv_exec_check like the 32-bit one did.
v0.1.395 — 2026-09-02
_The trap machinery and the context switcher run at 64 bits: qemu_virt is seven images -- hello, timer and sched at BOTH widths. Three bugs, all one family: a value and a size wearing the same literal._
The trap save layout collided with itself. 32 saved registers at 8 bytes end exactly at +0x100 -- where the handler slot lives, now 8 bytes wide, with the old depth word at +0x104 inside its upper half. `depth = 1` set bit 32 of the handler closure's address and the first timer interrupt jumped into nothing. The depth slot moved to +0x110, at both widths. Vec's initial capacity was the constant 4, not the word size. The offset- scaling pass turned li t2, 4 into li t2, wsz because it looked like every other addi _, _, 4. On RV64 that made cap = 8 above a 4-cell buffer: the first growth never fired and the fifth push wrote through the stale capacity into the cell's own length field. Diagnosing it took one wrong turn that the pinned lesson about boundary values names exactly: a probe pushing 1..5 read v[4] = 5 and looked CORRECT, because the corrupted read returns the length, and the length was also 5. `exit(4)`'s status code scaled too -- exit(8) on one width, exit(4) on the other, the kind of difference nothing notices until a script branches on `$?`.
And the scheduler example was itself width-32: it walked the register save area with raw_peek32 save (i * 4), which on RV64 reads half of every register. The missing vocabulary was CELL-indexed, machine-word-wide access -- raw_peekw / raw_pokew -- and with it the same source runs at both widths.
The lesson the three bugs share, written down where the scaling happened: a size scales with the width and a value does not, and a textual pass cannot tell them apart -- while the existing width's whole suite stays green, because there the substitution is the identity. Only running the NEW width sees it. The audit that closed the slice enumerated every li from x0 whose immediate had been scaled (zero remain) -- and the reverse ledger, operations whose NAME carries a width (raw_peek32, exit statuses, protocol constants), stays unscaled by contract.
v0.1.394 — 2026-09-02
_-rv64: the same backend, twice as wide. hello runs on qemu-system-riscv64, byte-identical output to the RV32 run._
One backend, two widths -- an xlen knob rather than a second file, because two files drift. What actually changes:
LW/SW become LD/SD (149 sites, every one now spelled through `ldf3`/`stf3`) and every cell offset, frame slot, record field and length header scales through `wsz ()` -- at 32 all of it is the identity, and the whole RV32 suite ran unchanged after every pass of the surgery. `li` learned the 64-bit truth about lui: it SIGN-EXTENDS. li rd 0x80000000 the 32-bit way materialises 0xFFFFFFFF80000000, which as an address points at nothing. Outside signed 32 bits the value is built in 12-bit chunks; on RV32 the masked lui+addi remains, because there every value is a 32-bit pattern and 0x807E0000 is a legal unsigned address -- the first version gated on signed range and the unit suite caught it within seconds. `LoadAddr` is pc-relative (auipc) on RV64 for the same reason, and the same 8 bytes on both widths, so the layout logic does not care. `raw_peek32`/`raw_poke32` are 32 bits BY NAME, on either width: a device register is as wide as the device says, not as wide as the CPU. The blanket f3 parameterisation had quietly turned the finisher poke into an 8-byte store, which QEMU's test device ignores -- the guest printed everything and the machine never powered off, and a pipe through head hid it by killing qemu with SIGPIPE. The peek is LWU on RV64, so a device value with bit 31 set does not come back negative.
riscv_virt_hello -- heap, string building, recursion, device windows -- compiles with -rv64 --bare --load-base 0x80000000 and prints the same four lines on qemu-system-riscv64 that the RV32 build prints on qemu-system-riscv32 and on the Mere-written emulator. That is now a row in qemu_virt.sh.
Named and not yet done, in the script where the gap lives: timer and sched (the trap machinery) do not run at 64 yet, and args writes its block in 4-byte fields by hand, so it is width-32 by construction. Those are the next slice of the arc, along with the RV64 emulator in memu and the hosted path.
v0.1.393 — 2026-09-02
_The last unexplained RV32I difference has a name: == inside a polymorphic function is a word comparison, because this backend does not monomorphize._
graphql_stack_portable returned {"__type":null} and "resolved to a type \"Post\" that does not exist inside the schema" while every other backend printed the right answer. The chase went through the schema parse (identical, byte for byte), a probe inside the executor (the type-name list CONTAINS "Post", and list_member still said no), and ended in the stdlib:
let rec list_member = fn xs -> fn v ->
... if h == v then true else list_member t v;
The C backend compiles that function twice -- mu_member__list_str__ and mu_member__list_int__ -- and the string one compares content. This backend compiles it once; at codegen the operand type is still a type variable, and a type it cannot see is compared as the 32-bit word. Exact for ints. A POINTER for strings and compound values.
Worse than assumed, and measured: it misses even for two IDENTICAL literals, because string literals are not interned -- the "Query" in a list and the "Query" at a call site are two rodata blocks.
The fix is monomorphization, and that is designed work, not a patch -- recorded as such. What this version does is make the boundary impossible to forget: test/rv/poly_eq_word.mere asserts the CURRENT wrong answers, so the day the backend learns to specialize, the pin fails and points at itself and at the known-different entry to delete. Symmetrical with the KNOWN_DIFF freshness check, and for the same reason: a limitation nobody re-measures outlives its own fix.
With this, all five remaining known-different programs have named causes: two are 64-bit values on a 32-bit int, one wants more RAM than the sweep gives it, one measures an arena that does not exist here, and this one.
v0.1.392 — 2026-09-02
_time and random_int are real on the hosted RV32I path: the emulator answers Linux's clock_gettime64 and getrandom, and Time.now.year says 2026 on the Mere-written CPU._
The hosted -rv target already spoke two Linux rv32 syscalls (write 64, exit 93). It now speaks two more, for the same reason those numbers were chosen: a guest written against this machine is not learning a private convention.
`__rv_clock` -> clock_gettime64 (403), a fresh 4-word block returned as the tuple `(sec_lo, sec_hi, nsec_lo, nsec_hi)` -- a tuple on this backend IS a plain block of fields, so the block the kernel filled is the value. __rv_urandom32 -> getrandom (278) of one word, through the print scratch buffer, which is dead between print calls.
Both intrinsics REFUSE under --bare, at compile time: a machine has devices, not syscalls, and inventing a clock or a seed would let a program measure nothing and believe it.
On top of them the prelude builds the portable spellings. time assembles the two unsigned halves of sec in FLOAT space -- softfloat eating its own dogfood -- which also keeps the answer right past 2038, where the low half goes negative on a 32-bit int. random_int does rejection sampling over the 31-bit pool, because r % n favours the low residues whenever n does not divide 2^31, and the bound contract (positive, message included) is the interpreter's.
The gate is properties, not a diff -- a clock and an entropy source are nondeterministic by contract, so test/rv/host_services.mere checks that the epoch is sane, time does not go backwards, random_int 1 is 0, forty draws stay in range and are not all identical, and random_int 0 fails with the exact message. That last-but-one property is not hypothetical: getrandom's buf is a0, NOT write()'s a1 one branch up in the emulator, and the first transcription filled zero bytes at address four -- every draw read the same stale word, and "forty identical draws" is precisely what the gate said.
The emulator also answers -ENOSYS for an unknown syscall now, as Linux does. It used to fall through silently: a0 kept whatever it held, and a guest calling a syscall this machine had not learned yet read its own argument back as the result.
With a real clock, time_clock left the known-different list (the freshness check flagged it), and mere-ruby's Time surfaced a 37-bit ordinal packing that had been wrapping silently on 32 bits -- fixed on its side by making the ordinal epoch seconds, with the corpus at 162/162.
Matrix: time and random_int go stub -> yes; 20 stubs remain.
v0.1.391 — 2026-09-02
_float_of_str / str_of_float computed exactly, in integers: Ruby's 0.1 now parses, adds and prints on the Mere-written CPU._
$ rvrun 64 -- -e 'x = 0.1 + 0.2; puts x'
0.30000000000000004
$ rvrun 64 -- -e 'puts "%05.2f" % 3.14159'
03.14
contrib/softfloat/dec.mere: a double's decimal expansion is EXACT -- m 2^e, multiplying by two never lengthens the fraction and dividing by two adds at most one digit -- so the working value is an array of base-10 digits and every rounding decision reads real digits. No tables, no estimates.
The contract is the one the interpreter and the C runtime already share, spelled out at the top of the file: trim, drop _, decimal or hex or inf/infinity/nan, whole string or the exact error message; printing is %.{p}g for p = 12..17 stopping at the first p that parses back to the same bits, .0 appended when the output has none of . e E n i. The parse side hands pack a 56-bit significand and a sticky bit and lets the existing rounding do its one job; overflow is inf, underflow is a SIGNED zero, and 1e1000000 is decided from the leading digit's decimal exponent without building a million-digit array.
Checked against the machine, not against a table. The new arm of softfloat_check runs a fixed battery (the %.12g round-trip artifact on subnormals, 1e23, ties-to-even at 2^53+1, hex, the largest finite double) and then two thousand xorshift bit patterns, comparing every print against printf and every parse against strtod -- zero disagreements, on the interpreter and the C backend both. rv_float_check runs the same shape on the emulated CPU (twelve patterns, not two thousand: a full-range conversion builds ~700-digit arrays, a region reclaims nothing on this backend, and the count is a measured RAM budget).
Two portability notes earned along the way, both already in this repo's memory and both stepped in anyway:
The battery's xorshift diverged between backends: `>> 17` on a value with bit 31 set is an ARITHMETIC shift on RV32I, where that value is negative. The 15 real bits are masked now, and the comment says why. rv_exec_check flaked at a 60-second alarm, still flaked at 240, and the timeout was never the reason: nul_in_str's output contains a NUL, and grep without -a was swallowing lines behind a binary-file heuristic that looks at buffer boundaries. Two programs sat in the known-different list on the strength of that filter; the list's own freshness check is what got them out.
The no-float rule in softfloat_check now strips string literals before it greps: dec.mere's parse error says "is not a valid float" because that is the message the other backends produce, and a string cannot be a type.
Matrix: float_of_str and str_of_float go stub -> yes; 22 stubs remain (transcendentals, f_pow, f_min/f_max, and the host services).
v0.1.390 — 2026-09-02
_ruby -e 'puts 1 + 1' prints 2 on a RISC-V CPU written in Mere. The last bug was a region rollback reclaiming a Map's live cons cells._
$ mere -rv --ram 64 main.mere > prog.bin # mere-ruby, 39,719 lines
$ rvrun 64 -- -e 'puts 1 + 1' # on the Mere-written emulator
2
Blocks, hashes, classes, methods, ranges, string formatting -- a real slice of Ruby -- run end to end on a machine whose every layer is self-written. What still stops is the decimal float conversions (0.1 in Ruby source needs float_of_str), which were already on the list.
The bug, and why every earlier layer hid it
region R { body } used to park gp and roll it back at the closing brace -- sound only if nothing that OUTLIVES the region allocates from the bump heap inside it. On this backend that premise is false: Map is prelude-lowered Mere code, and map_set m k v on a map that lives outside the region allocates its new cons node INSIDE. The rollback declared the node reusable; the next closure allocation overwrote it; the map still pointed at it.
The typer cannot see this. The map is Map[__heap, ..] and mutating it is not an escape, because on every other backend map internals live in the map's own arena -- which is exactly why interp, C, LLVM and Wasm all print the right answer for the same program. mere-ruby runs every top-level statement inside region STMT, so one map_set per statement corrupted its constant table, and the crash arrived four million instructions later inside __str_concat.
So a region on RV32I now reclaims nothing: compile the body, keep the heap. Correct and hungrier. The rollback comes back the day prelude structures can allocate from a second, persistent bump area -- recorded as design work, and the two unit tests that used to pin the rollback now pin its ABSENCE, so that day fails them and gets pointed here.
test/parity/region_map_escape.mere holds the five-line shape, with the allocation churn after the region that a dangling node needs before it can testify -- unreused memory still holds its bytes, and reading it back looks like a pass. Poisoned by restoring the rollback: map_get returns garbage and map_len returns a memory address, the exact signature mere-ruby showed.
What it took to find
The chase is a story about instruments, and all of them are now permanent:
Every encoder refuses a value that does not fit its field instead of masking it. The first thing that surfaced: frames beyond ~500 locals had their fp offsets silently wrapped into OTHER SLOTS (`-3484 does not fit 12 bits`). Real bug, fixed with two-instruction wide access through s11 -- and not the bug being chased. `__rv_word`: the machine word any value is, on the target only. "Which address is this map cell" has no other spelling, and the corruption was finally pinned by printing a tuple's address before and after a statement boundary: 7178284 before, a stack address after. memu grew a pc ring, a register dump on abnormal halts, a heap/stack invariant watch, and an exact-address store watch (`rvrun 64 watch <addr>`). The store watch produced the verdict: the same three heap words written once by `rvmap_set` and once, two statements later, by an unrelated closure allocation. Same address, two owners -- the allocator had been rewound.
Hypotheses measured and discarded on the way, in order: the new try_or's s-reg saves (disabled them -- byte-identical corruption), the assoc-list Map (104 real keys, zero mismatches), map_len itself (correct to 500 entries), deep recursion (correct to depth 512), heap/stack collision (the new invariant watch, first verified able to fire, saw none), frame width (real, fixed, symptom unchanged).
v0.1.389 — 2026-09-02
_The parity suite has 142 programs and RV32I had never been run against any of them. Sweeping it found two silent wrong answers._
scripts/parity.sh exists because four backends should agree. RV32I is not one of the four, because running its output needs a machine. v0.1.388 added rv_exec_check.sh with five programs; this points it at the whole corpus, with the C backend as the reference -- both are compiled programs, so neither is doing something the other cannot.
`<` on a nullary variant compared heap pointers. derive (Eq, Ord) color orders by declaration order, which is the tag. The tag compare existed for == and stopped there, so ordering fell through to the integer path and compared the two one-word blocks' addresses. That is not a random wrong answer: operands are evaluated left then right, so the left one is always the older allocation and a < b came out true for every pair, whichever way round it was written. lt Red Blue and lt Blue Red were both true.
Ordering on a tuple, record or payload-carrying constructor is now refused by name. == on those has a generated structural helper; < has none, and a pointer comparison is not an approximation of one.
The Map compared every key with `str_eq`. The rv prelude's Map is an assoc list, and its key comparison reads the first word as a length and the rest as characters. For a nullary constructor -- a one-word [tag] block -- that walks off the end, and map_get c Green answered key not found for a key that was there. It did not fail; it read out of bounds and said no.
A key whose type is known and is not str is now refused, naming the type. Only a known non-str type: an unresolved type variable is let through, because a polymorphic helper called with string keys is a real program. That gap is deliberate and written down.
Both refusals are gated in rv_prelude_check.sh and not left to the sweep -- the sweep skips a program that does not compile, so a refusal quietly deleted would become a wrong answer again with nothing going red.
What the sweep says now
64 programs agree, 8 differ for reasons recorded by name in the script, and 63 of the 64 agree except for one line: the reference prints the program's own final value and an RV32I binary does not. That is a real parity difference and it is spelled out as its own accepted shape rather than filtered away, so a difference in the last line is still a difference.
The known-different list is checked in both directions. A program in it that starts agreeing fails the gate, which is not hypothetical -- nul_in_str was briefly in the list on the strength of a comparison that had merged stderr into the output, and the check said so on the first run.
v0.1.388 — 2026-09-02
_A try_or that fired handed the catcher back the failed callee's registers._
Five lines of Mere:
let g = fn (n: int) -> let a = n + 1 in let b = a + 2 in fail ("b" ++ str_of_int (a + b));
let f = fn (u: unit) -> let p = 5 in let v = try_or (fn (z: unit) -> g 1) 7 in p + v;
let _ = print_int (f ());
Every backend prints 12. RV32I printed 9 -- p came back holding 2, which is g's a.
A named binding on this backend lives in a callee-saved register. The unwind restored sp and fp by hand, which is enough to land back in the right frame and not enough to find the caller's values in it: the failing thunk had already written its own bindings over them. The record grows from 5 words to 15 and carries s1..s10, which is what setjmp saves and for the same reason. s11 is not among them -- it is the far-jump scratch and is dead between jumps.
The suite had a try_or test. Its thunk had no bindings of its own, so there was nothing to overwrite and all four backends agreed.
An unimplemented service is now a failure a program can catch
emit_abort -- what an extern fn call lowers to on a target with no C library -- wrote its message and called exit. So a try_or around such a call did nothing: the abort left the process from inside the handler's reach. It now goes through the same unwind fail uses, which is where that sequence already lived; the copy in emit_abort was the write-and-exit half only.
That matters because coping is the program's job. mere-ruby sets $$ from getpid at startup, and wrapping that line was ineffective until this changed.
The gate that was missing
scripts/parity.sh compares four backends by output and RV32I is not one of them, because running its output needs a machine. So everything about that backend that is a matter of runtime behaviour rather than emitted instructions had no gate: qemu_virt.sh covers the bare-metal side and nothing covered the hosted side. scripts/rv_exec_check.sh does now -- each program run on the interpreter and on the Mere-written CPU, outputs required to match, so no expected values are written down. The one exception is labelled: an extern fn that nothing implements has no behaviour on another backend to compare against, so that case alone carries its expectation.
Both fixes were poisoned and both were caught.
mere-ruby on the self-made CPU
$ ./rvrun 64
usage: mere-ruby [-v] [-h] [-e CODE] [-I DIR] [--] <script.rb | -> [args]
That is a 39,719-line Ruby interpreter, built by this compiler, initialising on a RISC-V CPU written in Mere, and printing its own usage message -- byte for byte what the native build prints. It gets there by coping with a machine that has no shell, no pid and no arena: the environment reads as empty, $$ is left unset rather than set to a made-up number, and the arena byte counters are not kept. Each of those is a decision the program makes, which is why the catchable abort had to come first.
-e still stops, in mere-ruby's own code rather than at a host wall. The message that reports it used to be mere-ruby: (built-in exception raised) -- naming neither what was raised nor why, while raise_exc was holding both arguments. It now says TypeError: Dir is not a class/module, which is where the next slice starts.
v0.1.387 — 2026-09-02
_Float arithmetic on a backend with no float: 2,880 results, bit-identical to the machine's own doubles._
contrib/softfloat has held an IEEE 754 double in 15-bit limbs for a while, and the RV32I backend has refused + on two floats for just as long, with a message saying the library "is not yet injected into the -rv prelude". It is now.
The prelude carries the library. It cannot import it -- an installed compiler has no contrib/ on disk, and the -rv path has no file system to resolve a path against -- so scripts/gen_rv_softfloat.py bakes the source into lib/rv_softfloat.ml, renamed into a private namespace: __sf_ on the lowercase names, __rv after the type names. The library keeps its plain names for anyone who imports it directly. One source, two namespaces, and sh scripts/gen_rv_softfloat.sh --check diffs the result on every run.
The rename is not paranoia. The prelude is PREPENDED to the user's program, so a top-level add in the program shadows the prelude's -- and would become what float + calls. mere-ruby has 806 top-level names and collides with none of these 76, which is luck, not a design.
The operators are calls, not lowerings. Bin and Cmp on floats rewrite to App (App (Var "__fadd", l), r) and go through the ordinary application path, the way print_bool already did: arity, argument order and tail position are handled once. A float operator is also the one case where the callee's name appears nowhere in the source, so vars_in had to learn to mark it -- without that, codegen emits a call to a function it never laid down, which is exactly what a poison run reported (undefined label u___feq).
`-x` on a float was computing a wrong number, quietly. Unary minus had one arm, and it negated the word -- which for a float is the pointer to its two halves. It is a sign-bit flip, defined on NaN too, so it goes through the library like everything else.
The gate: our arithmetic against the machine's
scripts/rv_float_check.sh runs test/float/rv_float_ops.mere twice -- once where a float is a C double, once on RV32I -- and requires identical output. Every one of 17 values against every other, four operators and six comparisons plus unary minus: 2,880 lines, and they match. There is no table of expectations anywhere in it; the reference is whatever the hardware does.
The values are bit patterns, not decimals: the largest subnormal, the smallest normal, a tie that round-to-nearest-even has to decide, the largest finite double, a signalling NaN, a negative NaN. The output is bits too, printed as four 16-bit chunks -- float_bits_hi cannot be printed directly, because it is unsigned 32-bit and the same bits come out negative on a signed 32-bit int.
It found two real bugs in the library:
`a - NaN` came out negated. `sub` was `add a (neg b)`, and negating a NaN flips its sign bit; every machine with a float unit returns the NaN it was given. The add/mul/div gates had never subtracted a NaN. `add` quieted one operand and not the other. Only a signalling NaN can see that, and the test had only a quiet one until this found the first bug and the fix went looking for company.
What it deliberately does not compare: with BOTH operands NaN, which payload propagates is unspecified, and this machine does not agree with itself -- its add returns the first operand's and its sub returns the second's. Those nine pairs print a skip line naming the reason, so the exclusion is in the output rather than in somebody's memory. The comparisons still run on them; NaN comparisons are specified.
The RV32I half needs a 32-bit machine, which is the hole scripts/softfloat_check.sh names in its own header. MEMU=<checkout> supplies one; without it the gate runs the hardware half and says the other did not run, rather than printing ok.
Also
`f_add` … `f_ge`, `f_abs`, `f_neg`, `float_of_int`, `int_of_float`: twelve matrix cells from `stub` to `yes`, sharing the operators' implementation rather than getting a second one that could disagree with `+`. `f_min`/`f_max` stay stubs on purpose: `if a < b then a else b` is a definition, but the host's is fmin/fmax with their own NaN rules, and guessing puts a wrong answer where an honest refusal is. The shim message no longer says softfloat "is not yet injected into the -rv prelude". It was true when written and stopped being true the day it was, and would have gone on telling users to do something already done. `%` and `++` on floats: the typer rejects those before codegen sees them, so the float abort had become a branch no input can reach. Deleted rather than left next to a comment claiming it handles a case. contrib/softfloat/bits.mere -- the hardware bridge, moved in from test/, because the prelude needs it and the compiler cannot depend on a test file. It is the one file in the library allowed to mention float, and the gate now checks the rule against a GLOB minus that one name: a seventh file added to the library is checked the day it appears, which the six hand-listed paths would not have done. The exemption is checked too -- a bits.mere that stopped using float would mean the bridge had moved. That gate's headline count was inflated: its path list named `div.mere` and `conv.mere` twice, so "75 exported names" counted repeats. Derived and deduped, it is 76. test/test_basic.ml asserted the OPPOSITE of the new behaviour -- that a float operator reached the abort -- and checked it by looking for li a7, 93, the exit syscall that _start, __oom and __pat_fail all end with. It held for every program this backend ever compiled. The needles moved to scripts/rv_prelude_check.sh, which drives the real compiler: the unit harness types with Typer.initial_env and never prepends the prelude, so the callee is not a known binding there and no direct call is emitted at all.
With this, mere -rv --ram 64 on that interpreter and rvrun 64 -- -e 'puts 1 + 1' gets past float arithmetic and stops at random_int. Which call site reaches it during a puts 1 + 1 is not yet known -- every one of them is inside a function, and none is a global initializer.
v0.1.386 — 2026-09-01
_args on a target with no host: the loader leaves the command line in RAM._
args was the last thing standing between a 39,719-line Ruby subset interpreter and running on this backend. It is a host service everywhere else -- ask the process what it was started with -- and -rv has no host to ask, so it was a shim that failed with a message. A program that cannot be told what to do can only do one thing.
What replaces it is what a kernel does for a Unix process: the loader leaves the arguments in memory and the program reads them from there. The block sits in the reserved top of RAM, derived from the RAM size like the framebuffer, so it is right at every --ram provided the loader was told the same one -- a requirement -rv already puts on it.
+0 magic "ARGV"
+4 count
+8 + 4*i pointer to argument i, already a [len][bytes] string block
... the blocks, padded so each starts word-aligned
Three things this gets right on purpose:
The magic word. RAM no loader wrote is not zero in general, and a count read out of garbage hands the program pointers into nothing. With it, an untouched block reads as no arguments, which is the true answer. No argv[0]. The machine did not load the program by name and has no name to offer. Inventing one would be a lie a program could branch on. No copying. What the loader leaves already has this backend's string layout, so a pointer out of the block is the string.
args itself is four lines of prelude over two loads. The list is assembled where it is typed rather than in codegen, where a wrong tag would be silent.
The two loads are declared to the typechecker, not just to codegen: a name codegen knows and the typer does not is not a builtin, it is an unbound variable. Every other backend refuses them and points at args. They get their own refusal rather than sharing the bare-metal one, which had filed them in the matrix under "bare" -- the wrong reason. What a hosted process lacks is the block, not the privilege.
`examples/riscv_virt_args.mere` is its own loader: it writes the block through a window and reads it back with args, so what it checks is the reader the compiler emits, with no loader and no emulator in the picture. QEMU decodes the instructions. It runs identically on QEMU and on the Mere-written emulator.
The lengths are 1, 4 and 5 for a reason. The first version of this used one argument and passed while the index was not scaled at all: the shift amount had gone into the immediate field where the value belongs, assembling to slli a0, a0, 0. Index 0 lands on the first pointer slot either way. Index 1 loads from a misaligned address and traps.
before: 0 is printed once a count and three good pointers are already in RAM and only the magic is missing. Asking before writing anything would not check the guard at all -- QEMU hands the program zeroed RAM, so a reader that ignored the magic entirely would read the count as 0 and look correct.
Also fixed: qemu_virt.sh compared the emulator's own rvrun: lines -- its halt reason, its pc trace -- against QEMU output that has no equivalent. All four images had been failing that half since the emulator learned to name its halt reason, which is a differential test reporting on its own diagnostics.
With this, mere -rv --ram 32 on that interpreter and rvrun 32 -- -e 'puts 1 + 1' gets as far as float arithmetic, which this backend does not lower yet. The argument now arrives; the next wall is a real one.
v0.1.385 — 2026-09-01
_The wide jump landed on top of __str_concat's copy pointer, because grepping for t6 found nothing and the runtime helpers use it under its number._
v0.1.384 gave the wide jump x31 as scratch on the grounds that nothing else in this backend used it. That was checked by searching for the identifier t6. The hand-written runtime helpers -- __str_concat, __str_eq, __print_int -- are emitted as raw enc_* calls with numeric register operands, so the search saw none of them. __str_concat's copy loop keeps its destination in x31, and the wide jump at the bottom of that loop overwrote it every iteration.
A guest built that way sat at pc=296 -- inside .sc_l1 -- for four hundred million instructions, printing nothing, which is exactly what a program stuck before its first write looks like.
Auditing the numbers rather than the names says there is no free register: t3..t6 are the helpers', a3..a5 are arguments four through six of any call, and s1..s11 are the named bindings. So one is reserved: s11 leaves the named-binding pool and carries the jump address. A function with more than ten live names now spills one more to memory, which is the whole cost.
With that, a 39,719-line Ruby subset interpreter boots on a CPU written in Mere and prints from its own code:
RV32I: args needs a host, and --bare hands the program the machine instead
That is the prelude's own shim reporting that --bare has no command line — which means the interpreter initialised, ran, and got as far as looking for its arguments. What it does after that failure is caught by a try_or is the next question; it keeps running and the instruction trace does not tick, which those two facts cannot both be true of, so something there is still wrong.
Also in memu: rvrun <ram> trace reports the pc every 100M instructions on stderr, which is what found pc=296. Off by default; off costs one comparison per instruction.
v0.1.384 — 2026-09-01
_Wide jumps, for the programs that need them._ auipc + jalr reaches ±2GB, and costs four bytes per jump — so only the programs bigger than J-type's ±1MB pay. The pass that measures the code decides, which is the two-pass layout from v0.1.382 earning its keep a second time.
t6 carries the address. Nothing else in this backend uses x31, and the trap entry saves x1..x31, so a trap landing between the auipc and the jalr cannot lose it.
Two knobs feed each other — wide jumps grow the code, and a bigger code region moves the globals — so the layout is a bounded fixed-point loop that says so if it does not settle. It settles in one extra round either way.
Four places answered "how many bytes is this item". The assembler's address pass, the encoder, the listing and the debug map, and three of them still said a Jal was four bytes after the wide form arrived — which would have put every address in a listing and every line in a debug map off by however many jumps preceded it. There is one item_size now, and the other three call it.
A 3.97 MB interpreter assembles to 4,536,425 bytes and runs: it executed for ten minutes of wall clock without halting, where before the change it took a trap at pc=4195244 in 0.13 seconds. Whether it is making progress is a different question, and the emulator cannot answer it yet: it buffers the guest's output until the guest halts, so a run that is killed shows nothing at all. That is the same shape as the silent halt fixed in memu earlier today, and it is next.
Small programs are byte-identical, checked by stashing the change and comparing.
v0.1.383 — 2026-09-01
_The comment said a bare B-type "silently truncates", and used J-type to avoid it. J-type does the same thing one megabyte out._ enc_j masked its immediate to 21 bits and nothing checked the range, so a call farther than 1 MB was encoded as a call to somewhere else.
That is what a 3.97 MB interpreter did: it ran off the end of its own code and took a trap at pc=4195244, past a 4,163,521-byte binary. It is a loud failure now — a jump of 3,168,864 bytes says so and names the reach it does not fit.
The fix that makes such a program work is auipc + jalr, two instructions with ±2GB of reach, chosen when the code is big enough to need it. The two-pass layout from v0.1.382 already measures the size, so the second pass can know whether to widen every jump — which keeps a small program's bytes identical, because it is only the big ones that pay. That is the next change; this one stops the wrong answer.
The driver catches Failure now, for the reason it already catches Out_of_memory and Stack_overflow: OCaml's default is "Fatal error: exception Failure(...)" and exit 2, which reads as a crash when the honest answer is that this backend cannot assemble that program.
And the emulator was hiding the evidence. memu's RV32 runner halted on an exit syscall, on a trap with no handler, on a write to the test finisher, and on MULH — all four with vec_set st 1 1 and not a word. A guest that stopped on its first unimplemented instruction looked exactly like one that ran to completion and printed nothing, which is what a 4 MB guest looked like. It names the reason and the pc now, on stderr so a gate comparing stdout is unaffected. That is what turned "no output, exit 0" into "halted at pc=4195244 -- a trap with no handler installed", and the range check followed from there.
v0.1.382 — 2026-09-01
_The globals were at a fixed offset, and a program's code grew past it._ globals_base was load_base + 0x200000 whatever the program was. A 39,719-line interpreter emits 3.97 MB, so its first global store would have landed on its own code.
It follows the code now, rounded up to the megabyte with a megabyte of room after it, and never below the old value — so every program that fit before assembles to the same bytes as before, which is what the suite checks.
Sizing it needs the code size and the code needs the address, so the items are emitted twice. That terminates because no item's size depends on an address: LoadAddr is eight bytes whatever it loads, and a li of any global address is two instructions at every base above 4 KB. The second pass asserts the size did not move rather than trusting that argument — and the assertion is a guard for a future change, not something this commit demonstrated: shifting the rounding boundary does not make an emission base-dependent, so there is no cheap poison for it.
code_size sits next to assemble, which computes the same sizes in its first pass. A size computed by one rule and encoded by another is a silent mismatch, so they are adjacent.
The interpreter's globals now land at 0x500000, above its code. It still does not run: with the globals at 5 MB and the stack at 8.25 MB there is about 3 MB of heap, and that program wants gigabytes. What this removes is the self-corruption, not the memory requirement.
v0.1.381 — 2026-09-01
_A 39,719-line Ruby subset interpreter now compiles for RV32I._ Four megabytes of flat RV32IM, from a program that a week ago stopped on the first of its float lines. What was left after or-patterns was one language feature and one clock.
A top-level function used as a value gets a one-word closure whose pointer is an adapter, not the function. A closure here is called with the closure in a0 and the argument in a1; a one-argument top-level function wants the argument in a0. The adapter is that move and a tail jump, emitted once per function however many times it is used.
A curried one is refused, and the message says how many arguments it takes and why: a partial application has to allocate a closure per argument, and this backend has no currying layer. to_s, which is what the interpreter passes around, takes one.
time is the last host-service shim. A fixed number would make a program that measures elapsed time report 0 rather than say it cannot measure.
It compiles; it does not yet run. The binary is 3.97 MB and the code region is globals_base - load_base, which is 2 MB — so the code overlaps the globals and the heap by almost two megabytes and would corrupt itself on the first store. globals_base is a fixed offset from the load base, and making it follow the emitted code is the next thing. The binary itself is right: it opens with the stack setup and the store that clears the try_or handler word.
v0.1.380 — 2026-09-01
_Or-patterns on RV32I, refused by a rule where they bind._ Try the first alternative; jump past the second once it matched; on both mismatches, fail. Both the scrutinee and the stack pointer are kept, because the first alternative may be a container pattern that parked its pointer and jumped out from inside — the same thing v0.1.379 taught compile_match to undo at an arm's fail label.
An or-pattern that binds is refused, and the message says which rule: both alternatives would have to bind into the same slot, and slots are handed out while the pattern is walked, so they would name different ones and the arm body would read whichever the compiler saw last. A narrow refusal that states its rule beats a wide one that does not.
Two of three poisons walked through the first test set, which had the first alternative failing and the second either failing or matching:
- The stack-pointer restore matters only when the first alternative is a
container that failed part-way, AND the second matches, AND the whole match sits in a call argument — then the leftover word is popped as one of that call's other arguments. Any two of the three and it hides.
- The jump past the second alternative matters only when the first matches
and the second would not: without it, control falls into the second attempt and the arm is rejected for a value it had accepted.
Both cases are in test/parity/or_pattern.mere now, and the comment says why each one is there rather than leaving the next reader to guess.
v0.1.379 — 2026-09-01
_Four bytes that were never the problem; what they displaced was._ A container pattern parks its pointer on the stack and unparks it on the way out the bottom, so a sub-pattern mismatch left it there. Read as a leak, four bytes per failed arm sounds like something to fix eventually.
It is not a leak. The pushes an enclosing call has already made for its other arguments sit on the same stack, so the extra word was popped as one of them. f (n - 1) (a + (match (1,2) with (1,3) -> 0 | _ -> 1)) called f with a heap address where n belonged, and the loop ran until the heap met the stack — ten iterations was enough to report out of memory, which is what made the leak reading impossible: forty bytes cannot close a six-megabyte gap.
compile_match now records the stack pointer as it stood before the arm's pattern and restores it at the arm's fail label, which handles arbitrary nesting because every push a pattern makes is below that mark.
The bisection is in test/parity/match_arm_stack.mere, and each case is there because removing it hides the bug: the mismatch is needed (a matching pattern unparks), the container is needed (an int pattern parks nothing), the enclosing call is needed (nothing else has pushes to displace), and the tail position is not — the non-tail case fails identically, which is what ruled out the frame teardown.
That file found a second bug on its first run, in a different backend. An int literal pattern inside a container -- (1, y), R { p = 1 }, A 1 -- made the LLVM backend emit icmp eq i32 against an i64, so such a program did not compile there at all. int widened to i64 in v0.1.96 and that one line was left behind. The top-level match 1 with 1 takes a different path and was always right, and no file in a 148-program parity suite had a literal inside a container, so nothing had ever asked. bool and str in the same position were already correct; only int was wrong.
Same shape as v0.1.367, twice over: a widening that did not reach every site, and a check that was deleted for the backends that outgrew it.
v0.1.378 — 2026-09-01
_The FFI boundary, the host services behind it, and a matrix that lied twice about both._
A call to an extern fn now aborts at runtime naming the symbol, instead of being refused at compile time. fb_set / key / present still lower: those are MMIO on the machine itself, not calls into a library. Everything else is a C symbol there is nothing here to link against, and refusing refuses the whole program for a call it may never make — the Ruby subset interpreter this is being carried for declares twelve, seven of them TCP, on paths a script never reaches. An extern used as a value is still refused, because higher-order is unsupported here and there would be nothing to abort inside.
The host services go the same way, as prelude shims rather than codegen cases so they are visible in the prelude's own source: run, read_file, write_file, file_exists, read_stdin, args, and bytes_of_str / print_bytes — the last two not a host service but the bytes type, which has no representation here at all. random_int is a shim rather than an implementation because a deterministic sequence returned from something named random is the kind of wrong that stays quiet.
The reclamation API is not shims. map_compact and map_recycle exist to return bytes an arena is holding; this backend's Map is the prelude's assoc list of ordinary values, so there is nothing behind it to return and a no-op is the honest answer. map_clear does have work. map_bytes and vec_bytes ask HOW MANY bytes the arena holds, and answering 0 would read as "this uses no memory" rather than "the question does not apply", so they stop and say which it is.
`host_matrix.sh` caught the same lie twice. In v0.1.375 the float shims turned 27 cells to yes, and stub was added for them — keyed on that family's exact phrase. The host shims arrived with a different sentence and nine more cells flipped to `yes`, which only the matrix's own diff noticed. It keys on the RV32I: prefix every abort message shares now, and rv_prelude_check.sh asserts that both families produce it. A prefix survives the next family; a phrase does not.
RV32I reads 60 yes, 40 refused, 36 stub, 12 unattributed, 2 bare.
rv_prelude_check.sh also had to be updated rather than believed: it asserted the compile-time refusal for externs, which this change removed, and it said so.
What is left for that interpreter is a language feature, not a service: or-patterns, which this backend does not compile. Every host builtin it reaches is handled.
v0.1.377 — 2026-09-01
_try_or on RV32I, and five poisons the first test set could not see._ There is no longjmp here, so unwinding is a jump to a recorded address: try_or builds a five-word record — previous handler, sp, fp, catch address, default — installs it in a word reserved at globals_base, and fail reads that word before deciding whether to exit.
The record is on the heap, not the stack, and that is deliberate: it stays readable after sp has been moved back, where a record below the restored sp would not be.
Nesting works because the record keeps the PREVIOUS handler and both paths put it back. A single global slot holding "the current handler" would be overwritten by an inner try_or and never restored, and the inner one would then catch the outer one's failures forever after.
The reserved word is at globals_base, ahead of the top-level value bindings, rather than in the print scratch region — that region's whole description is "the buffer the print helpers build digits in", and putting unrelated state there would make the description untrue.
Five of six poisons got past the first test set, which had looked convincing: catch a failure, catch it from twenty frames down, nest two of them, and run a 2,000-deep recursion afterwards to show the stack was intact. All four passed with the unwinding half-broken.
failnot restoring the previous handler is invisible with one level of
nesting, because the normal path restores it too. It needs an inner try_or that CATCHES and is then followed by another failure.
- The normal path not restoring it is invisible in the value: the next
fail restores the handler, the second attempt reaches the right one, and only the code between the inner try_or and the failure runs twice. That needed a counter, not a value.
spnot being restored does not corrupt anything once. It drifts. One
try_or and a 2,000-deep recursion afterwards prove nothing; 20,000 of them meet the heap.
test/parity/try_or_unwind.mere is the eight cases, MATCH on C, LLVM and Wasm, and verified on the Mere-written RV32IM emulator. All six poisons now fail it.
One timeout in the poison harness was not detected as missing on this platform, and its shell error read as "the poison was caught" for four runs. A poison that does not land looks exactly like a gate that does not catch — twice in two days.
The Ruby subset interpreter's remaining wall is now only the FFI boundary: twelve extern fn declarations, seven of them TCP.
v0.1.376 — 2026-09-01
_Three walls down, and the one left is not a builtin._ Enumerating what still stops the Ruby subset interpreter on RV32I -- by shadowing each refused name with a diverging function, which types as 'a -> 'b and so needs no signature -- gave a list of three: map_len, print_err, try_or.
map_len is a redirect, not a lowering. Maps here are the prelude's assoc list reached by rewriting the map_* names, so rvmap_len counts DISTINCT keys: rvmap_set prepends and a key set twice is in the list twice, with the newer one shadowing, so counting nodes would count the shadowed ones. Both halves of the redirect table had to learn the name — the reachability seed and the call site — because missing the first emits a call to a function that was never emitted.
print_err sets the descriptor. The emulator's write syscall ignores it; QEMU's does not, and a diagnostic on stdout is a diagnostic in the wrong stream.
try_or is left. It is the catchable-failure mechanism, and fail here writes and exits with nothing to unwind to.
And the wall after those is not a builtin at all. An extern fn is a name the program declared, and this backend answered unbound variable for it — blaming the user for a target that has no C library to link against. That is the shape Q-070 closed for builtins, one declaration further out. It names the limit now, and a name nobody declared still reports as unbound, which the gate checks in both directions.
The Ruby subset interpreter declares twelve of them: seven for TCP, four for libc memory, and getpid. Those are the boundary now — one builtin and an FFI surface, from a program that a week ago stopped on the first of 39,719 lines of float.
Two of the three gates for this could not live in test_basic.ml. It calls the backend directly, so it cannot see the prelude, and its own typer answers unbound variable for an extern before codegen is reached. They are in scripts/rv_prelude_check.sh.
v0.1.375 — 2026-09-01
_A cell that compiles and then aborts is not yes._ Merging the RV32I float work made host_matrix.sh fail, which is the gate working: the -rv prelude now defines the float builtins, so a probe for sqrt or atan2 compiles and 27 cells were about to flip to yes on a backend that cannot do the arithmetic.
stub is the third value for that, next to the bare this table already had. yes would be the flattering lie and refused the other one. The int builtins the same change added — abs, max, min, clamp, even, odd, gcd — are yes, because they are implementations; so are float_bits_hi / float_bits_lo / float_of_bits, which are the representation.
RV32I now reads 59 yes, 50 refused, 27 stub, 12 unattributed, 2 bare.
The comment about how this is held was false when it was written. Detecting a stub means looking for the shim's message in the emitted binary, so the cell depends on that wording, and the comment said rv_prelude_check.sh asserts the phrase. It asserted only the word softfloat. A poison run reworded the rest, 27 cells turned to yes, and that script stayed green. It checks the phrase now, and the same poison fails both.
stub is in the summary line too, for the reason the others are: a number nobody prints is a number nobody notices growing.
v0.1.374 — 2026-09-01
_A float on RV32I is two words, the arithmetic is not lowered, and the program says so at runtime instead of quietly adding two pointers._
compile_bin does not look at types. Letting a float literal through without a guard would have sent 1.5 + 2.5 to the integer add and returned a number, so the guard came in the same change as the literal. A float value is a two-word block — the halves of the IEEE 754 pattern — which is the shape float_bits_hi / float_bits_lo already ask for; those two and float_of_bits are real here now, not scaffolding.
The arithmetic is a scaffold and says so: an operator aborts at runtime with a message naming contrib/softfloat. A compile-time refusal was the other option and is worse for the program this exists for — a Ruby subset interpreter carries float code on paths a script never reaches, and refusing at compile time refuses the whole program.
The shims needed their types written out. Left to inference they are 'a -> 'b, which unifies with anything and moves the failure elsewhere: the typer stopped on float_of_str n / float_of_str d with expected float, got int, because a fresh type variable let / resolve as integer division.
abs / max / min / clamp / even / odd / gcd are also new to this backend and are not scaffolding — they are int functions that were on the refused list only because nobody had written them.
New gate, scripts/rv_prelude_check.sh, because the prelude is injected by the driver and test_basic.ml calls the backend directly — it could not see any of this. The probe list is derived from the prelude source, so a name added without one fails by name. It found 25 existing prelude names that nothing had ever compiled, and pins a bug it turned up on the way: rvmap_set and rvmap_get used on the same map fail to type-check, and the message is unbound variable: rvmap_new — for a name that is bound. Either alone is fine, and it reproduces on the interpreter, so it is not a backend hole. The pin expects the failure and reports when it stops happening.
What this unblocked: mere -rv on that Ruby subset interpreter now compiles past every float in its 39,719 lines. The next wall is not float and not this compiler — it is 0xFFFFFFFF in a gzip library it depends on, which the literal check added in v0.1.367 refuses on a 32-bit target.
v0.1.373 — 2026-09-01
_The magnitude of a negative integer is not 0 - n._ contrib/softfloat/conv converts between a double and the backend's own int. Both directions have a rule that was measured and not assumed: int to double rounds to nearest with ties to even (2^53+1 becomes 2^53, 2^53+3 becomes 2^53+4), and double to int truncates toward zero (-1.9 becomes -1).
The most negative integer has no positive counterpart, and negating it wraps back to itself. The gate caught -4611686018427387904 — a value that cannot be written as a literal, since the lexer caps at 2^62-1, and had to be reached by arithmetic to be tested at all. n / 2 always has a counterpart, so the magnitude is built from half of it, doubled in the limbs, with the odd bit put back. Poisoning that back to 0 - n fails exactly one check, which is the right number: there is exactly one such value.
str_of_float is not here. Mere's prints shortest-round-trip decimal (0.1 is 0.1, not 0.1000000000000000055511151231257827), which is a different piece of work from arithmetic — Ryu's, not IEEE's — and it gets its own slice rather than an approximation hidden inside this one.
Also: l6_of_int / l6_to_int in contrib/softfloat/limbs, five limbs wide, which is more than any backend's int.
v0.1.372 — 2026-09-01
_Division, and an invariant that only holds if the loop is entered correctly._ contrib/softfloat/div is restoring long division: 56 quotient bits, one per round, and the remainder left at the end is the sticky.
The loop keeps 2^i * a = q * b + rem with rem < b, and one subtraction per round can only hold that if the remainder is already below b when the loop starts. Both significands are normalised to 53 bits first, so a / b is in (0.5, 2) — and a can be the larger. The first version began with rem = a, and every quotient where that happened came back at half its true size: 286 of them, all with the same shape, which is what said it was one condition and not a scatter of arithmetic slips.
Normalising the operands before dividing is the other half. A subnormal divisor makes the true quotient far wider than 56 bits, and a loop that produces exactly 56 would drop the top and answer with a small number where the truth is an overflow. The exponent pays for the shift.
pack grew a bound to go with it: below 1 the underflow shift is capped at 60 bits, because past that everything remaining is sticky, and a normalised subnormal over a large normal can arrive thousands of exponents below.
Seven poisons, each caught, and the counts say what each one is: removing the pre-step fails 286, one round fewer fails 648, the exponent off by one fails 652, and 0/0 answering zero instead of NaN fails 4 — the whole table has only four such pairs.
Verified on the Mere-written RV32IM emulator as well: 1/3, an overflow, an underflow to the smallest subnormal, x/0 and 0/0 all print the limbs the interpreter prints.
v0.1.371 — 2026-09-01
_Multiplication, which is what the 15-bit limb was for._ Two 53-bit significands make a product of up to 106 bits, and the schoolbook loop builds it from limb-by-limb products of at most 30 bits — small enough for RV32I's signed 32-bit word, on one condition: the carry is propagated at every step. Accumulating four columns before carrying reaches 2^32, which is right on a 64-bit int and wrong on the target.
The bug the gate found is the one a tolerance would have accepted. The first version shifted the product down by a fixed 49, which is correct for two normals — their product is 105 or 106 bits — and wrong the moment either operand is subnormal, because then the bits below 49 are significand and not sticky. 0,1 * 2146435071,4294967295 came back with the high half right and the low half zero. The shift is now computed from the product's own width.
Subnormal results needed a step in pack that addition never reaches: a product's exponent can land below 1, and the answer there is not zero but a subnormal, obtained by shifting right until the exponent is 1 and rounding there.
What the gate cannot see, said out loud. The limb width exists so a product of two limbs fits a signed 32-bit int, and both arms of scripts/softfloat_check.sh run where int is 63 or 64 bits wide. A loop that accumulated before carrying would pass both and overflow on the target. Compiling for RV32I proves the code is expressible there, not that the widths hold. The check that does hold it is running the library on a 32-bit machine: done here on the Mere-written RV32IM emulator, where the fullest cases — every mantissa bit set, on both sides — print the same limbs as the interpreter. That is not in CI, because the emulator lives in another repository; the script now says so rather than implying the coverage.
v0.1.370 — 2026-09-01
_IEEE 754 addition in integers, and two NaN rules that had to be measured._ contrib/softfloat/add adds two doubles with no floating point anywhere: significand shifted left by three so the guard, round and sticky bits have somewhere to live, aligned by the exponent difference, added or subtracted by magnitude, normalised, and rounded to nearest with ties to even.
Subnormals need no separate path. The exponent carried around is the effective one — a value is significand * 2^(E - 1075), with E = 1 for subnormals — and the normalisation step simply cannot lower E below 1. What falls out is a significand that stays small, which is what a subnormal is.
The gate is test/float/softfloat_add.mere, and it compares bits, not values: 484 pairs from the format's landmarks, 9,000 swept pairs, and 1,800 more in the top binade. Bit equality is what puts -0.0, the NaN payloads, the subnormal boundary and tie-to-even inside the check instead of on a list of things to remember.
Two rules came from the gate rather than from the specification. The first: a signalling NaN operand comes back quieted, which cost 45 failures. The second: a signalling NaN outranks a quiet one whichever side it is on, so quiet + signalling is the signalling one quieted, not the first operand. That cost one failure, and it is not a rule anybody would guess.
Poisoning found one thin spot worth naming. Moving the overflow threshold from 2047 to 2048 failed exactly two checks, because only max + max and min + min reached it. The interesting overflow is not that pair but a sum that is finite until the round bit is applied, so the sweep now spends 600 rounds in the top binade — and the same poison fails 1,283.
Also in this slice: l6_bit, and test/float/sf_bridge.mere, which holds the conversion between hardware doubles and the limbs so that both gates share one copy and neither library file has to name float.
v0.1.369 — 2026-09-01
_Six 15-bit limbs, and the two tuple accessors the library needed to return a sticky bit._ contrib/softfloat/limbs is the working width for adding two doubles: 90 bits, wide enough for a 53-bit significand plus the guard, round and sticky below it and the carry above, in limbs narrow enough that a product of two fits RV32I's signed 32-bit word.
Its gate is test/float/softfloat_limbs.mere, and the oracle is the backend's own integers — a real oracle here rather than a tautology, because the limb code never adds two whole values, only one limb at a time with an explicit carry.
Three holes in that gate, found by poisoning it:
A hand-written table caught the carry it was written for and little else: breaking one link of the carry chain failed 2 checks and one borrow failed 1. A deterministic sweep took those to 91 and 194.
The sweep could not reach the top. Its values stay under 2^58 so the native oracle is exact, and l5 (bits 75..89) is outside a 63-bit int entirely — dropping the borrow into l4 passed everything. Up there the oracle is an identity instead of a value: (a + b) - b is a, and a shift up and back down is a. Neither needs a number the backend can hold.
l6_bitlen's base for l4 was checked by nothing at all. Moving it from 60 to 61 passed every check in the suite, because no value in the table or the sweep ever had its top bit up there. A single bit at position i has length i + 1, which needs no oracle, so the check now walks 58..89.
One poison that appeared to reveal a fourth hole turned out not to have applied: the sed looked for else if a.l5 where the source says if a.l5. A poison that does not land looks exactly like a gate that does not catch.
Also: fst and snd on RV32I. A tuple there is already a block of words with element i at i4 — what `Ast.Tuple` builds and what a `P_tuple` pattern reads — so their absence was not a missing mechanism but a missing pair of cases. Destructuring a pair in a `let` compiled; calling `snd` on the same pair did not. The library that wanted to return a value and a sticky bit together found it.
v0.1.368 — 2026-09-01
_A double's bits, held as integers narrow enough for the backend that has no float._ contrib/softfloat is the representation half: sign, exponent, and the 52-bit fraction in four 15-bit limbs, little-endian. 15 because a product of two limbs has to fit the narrowest int any backend has, which is RV32I's signed 32-bit word.
The obvious representation is the one float_bits_hi / float_bits_lo already give — two 32-bit halves — and it does not work on the target that needs it. Those halves are unsigned (float_bits_hi (-1.0) is 3220176896) and RV32I's int is signed and 32 wide, so a half does not fit. The same shape of answer as v0.1.281's, one target further down: one accessor would not fit an interpreter's 63-bit int, and two do not fit a 32-bit one.
The file never names `float`. That is the design claim, and the gate checks it as one.
The gate is scripts/softfloat_check.sh: every double in test/float/softfloat_roundtrip.mere survives the trip into limbs and back as the same 64 bits — including -0.0 and NaN payloads, which a value comparison lets through — the interpreter and the C backend print the same bytes, and every exported name compiles for RV32I. The inputs are built from bit patterns rather than written as decimals, because the boundaries of the format (the largest subnormal, the smallest normal, the limb split that straddles the two halves) are exactly the ones nobody writes down.
The first version of the RV32I arm was green and checking almost nothing. It compiled a probe that called three functions, and passed with a float-using function sitting in the library: nothing referenced it, so codegen never saw it. The probe is now test/float/softfloat_narrow.mere, written by hand, and the script asserts that every let in the library appears in it — so adding a function without exercising it fails the gate by name. A generated probe was tried and rejected: it has to guess each function's shape, and a wrong guess fails on the probe's own type error instead of on the defect.
v0.1.367 — 2026-09-01
_The check was deleted for the backends that outgrew it, and the backend that needed it arrived after the deletion._ v0.1.41 refused an integer literal that did not fit the target's int on LLVM and Wasm. Both widened to 64 bits later (v0.1.96, v0.1.127) and the rejection went with them. RV32I is 32-bit and was added after that, so it never had one, and li truncated in silence: 4294967295 printed as -1 where the interpreter printed 4294967295, and 3220176896 — the high half of -1.0's bit pattern — came back as -1074790400.
It surfaced while sizing a soft-float representation, where the question "what fits in a word here" has to have an answer the compiler agrees with.
The boundary is the interesting half. -2147483648 is in range and 2147483648 is not, and the parser hands the first over as a negation of the second — so a check that reads only the literal refuses a value that is perfectly representable. The fold happens before the check, and both directions have a test: one past the top, one past the bottom, and the boundary itself.
v0.1.366 — 2026-09-01
_exit was the one host builtin between a 39,719-line program and the wall behind it._ The RV32I backend refused exit, and fail — three lines above it in the same match — already ended with the ecall that does the work: a7 = 93, with the status hardcoded to 1. The lowering is that tail with a0 taken from the argument.
What it unblocked is the more interesting half. Compiling mere-ruby, a Ruby subset interpreter written in Mere, with -rv stopped at line 39,719 on exit. It now stops at line 7,716 on x * 0.5. That wall is not a missing builtin: the backend's value model is one 32-bit integer per register, and a Mere float is an IEEE 754 double. There is nowhere to put it.
--bare refuses exit for the reason it refuses print — the exit syscall is a courtesy of the host, and a machine handed to the program does not answer it. A user process under a kernel is not --bare, so it gets the lowering.
The backend needed no shadow check to go with it. compile_app tests locals, globals and top-level functions before any builtin name, so a program that defines its own exit calls its own; codegen_c and codegen_wasm each need an explicit user_shadows guard to reach the same answer.
v0.1.365 — 2026-08-31
_The point-in-glyph test was half-open in the wrong variable._ It counts ray crossings and dropped t = 1 while keeping t = 0, which counts a vertex two segments share exactly once — right, until the shared vertex is a maximum or minimum in y. There the curve touches the ray and turns back and the net has to be zero, but the arriving segment was dropped and the departing one kept. The leftover crossing was never cancelled, and since the ray runs rightward it made every pixel to the left of a flat-topped stem come out inside the glyph.
At 30ppem the letter d grew a bar across the top of its ascender, ed grew one across both letters, and de did not — the asymmetry is the fault, not a coincidence.
The rule is now half-open in y. Each curve is cut at its own y extremum, which leaves pieces monotonic in y, and each piece is asked whether py lies in its half-open span. A maximum is the upper end of both its pieces and belongs to neither; a minimum is the lower end of both and the two directions cancel. The three special cases the old form needed — a horizontal segment, a negative discriminant, a double root — are all just spans that fail that one test, so they are gone.
The separation matters more than the arithmetic: the count now comes from comparing endpoint heights, and a root is solved only to decide which side of px the crossing is on. Rounding can still move where a crossing is. It can no longer change how many there are.
Every existing gate is byte-identical across the change — glyf_check in all four modes, font_check, and the unit checks that reach the rasteriser through layout and paint. That is the shape a fix should have: nothing moves except what was wrong.
mbrowse gains `scripts/winding_check.sh`, and it needs no oracle. Reflecting a glyph in x and reflecting the query point with it turns the rightward ray into a leftward one over the same code, so the two answers can be compared directly — inside is inside whichever way the ray goes. It reports 87 disagreeing pixels against the old rule and none against the new one, across letters and sizes no case list had asked about. The browser-oracle raster comparison holds the renderer to real numbers, but only at the sizes it names; this one is size-blind by construction.
v0.1.364 — 2026-08-31
_A file that imports nothing was reachable by exactly one program, because of where it sat._ src/font.mere in the mbrowse repository reads a TrueType file and answers how wide a string is and where a glyph's curves go. It imports nothing, it opens no window and it knows nothing about the web — but living in a browser's src/ made that browser the only thing that could use it. It is contrib/font now, beside contrib/raster and contrib/window, which are the two halves of the same job.
Its gates do not move. The oracle for a metrics reader is a browser's measureText over the same font file, and mbrowse is the repository that has a browser to compare against: font_widths 29 of 29 and glyf_points 13 of 13, unchanged across the move, along with the unit gates that reach it through layout and paint. A metrics reader that agrees with a tool nobody uses has agreed with nothing.
Three of five backends take this file. interp, C and RV32I compile it; Wasm and LLVM do not, both inside the contour walker in _points and with different symptoms — Wasm cannot see an inner-lifted capture (os), LLVM an inner binding (ends). Neither is new and neither is caused by the move; they are the two backends' MVP limits on inner-lifted functions meeting the most nested code in the tree. So there is no bootstrap_wat_ok entry for this module, and its absence is a fact about the backends rather than an oversight.
contrib/font/README.md writes down what the reader does not do, because those are the parts that fail quietly: a CFF font has no outlines here and draws nothing, and a variable font satisfies every stated condition and still comes out at whatever its fvar default is — frequently the lightest master rather than the regular one, with no error to say so.
v0.1.363 — 2026-08-31
_Four places in this tree decided what a heading is, and only one of them had heard of a fenced code block._ contrib/markdown answered it three ways — to_html wanted # with a space and stopped at three levels, to_text needed no space and had no limit, toc counted hashes — and contrib/site answered it a fourth time in build_toc_html, then a fifth in inject_heading_ids. Only to_html tracked fences, so ## x inside an example was a <pre> in the body and a section in the table of contents at the same time.
The cause was that there was no parse. Three outputs cannot be three views of a document that does not exist, so each one re-read the lines. contrib/markdown is now the mere-markdown package: one Md.parse, and HTML, plain text and the table of contents as views of it. Its own gate runs the frozen pre-split implementation as the oracle and compares 72 outputs byte for byte, which is 72 more comparisons than the three compile-and-size checks this suite had.
This is the first external dependency in this repository. mere.toml names it, mere.lock pins the revision and hash, .mere_modules/ is ignored, and mere install reproduces it — the same shape mere-blog and mpng already use, so the package manager now has a consumer that is not a demo of itself.
contrib/site parses once and both the body and the table of contents are views of that. The generated documentation changes on 16 of 19 pages, and all of it is escaping: text and attributes were escaped inside code blocks and nowhere else, so --ram <MB> in bare-metal.md has been reaching every reader's browser as a tag and being swallowed. One page also gains an <h4>, which had rendered as a paragraph while appearing in the contents as a depth-4 heading.
The fence defect turned out not to reach this site: docs/ contains no ## or ### inside a fence, and the table of contents only shows those two levels. It was real in the other two outputs: the table of contents for http-demos.md listed open http://localhost:8080/ in a browser as a section, tutorial.md listed → 252, and tutorial-rest-api.md listed → {"deleted":true} — all of them # comments and output lines from inside shell examples.
Connecting the first real consumer found what using it alone could not: the package's Text constructor collided with contrib/html's, and the type error pointed at the consumer's file rather than the package's. Type declarations are top level, so a module does not keep its constructors. Renamed to Run / Strong / Emph / Strike / Mono / Href / Pic.
v0.1.362 — 2026-08-31
_host_matrix asks which backend has each builtin; nothing asked whether the documentation had heard of it._ Twenty-five of 221 had not been written about by hand anywhere: every name in the bytes family — a first-class type with an arc behind it — along with map_clear / compact / recycle / bytes, vec_concat / reverse / compact / bytes / of_bytes, owned_vec_get, channel_new, file_open, file_read_line, list_dir, mkdir_p, read_stdin, hex_of_bytes and str_of_bytes. vec_reverse appeared in no file under docs/ at all, the changelog included.
scripts/doc_coverage_check.sh compares mere --dump-builtins against the hand-written docs. The changelog is excluded because a name that appears only there has been announced and never documented, which is the case being looked for, and docs/host-matrix.md is excluded because it is generated from the same builtin list — counting it would let the gate pass by citing itself. An allowlist takes <name> <reason> lines for deliberate omissions, and an entry whose builtin has since been documented is an error, so the list cannot rot into a to-do nobody rereads.
The twenty-five were documented rather than excused: stdlib-reference.md now carries the bytes family with signatures, the rest of Vec / Map / OwnedVec beyond what language-reference and the tutorial cover, and the six that belong to no family. 221 of 221.
What the gate does not check is stated in its header, because a coverage number invites the other reading: it asks whether the name is spelled somewhere, not whether it is explained. That weakness bit immediately. The paragraph introducing the new roster named two of the builtins it existed to cover, as examples of what had been missing — so deleting a table row left the name in the prose ABOUT the row, and the first poison run stayed green. The sentence names none of them now, and the poison fails as it should.
Three measurements were wrong before one was right, all of them the instrument rather than the subject: grep here is a ugrep wrapper where an anchor inside a group silently matches nothing, an unquoted list of filenames is one argument under zsh rather than sixteen, and "the reference documentation" was defined as two files when the collection families are deliberately delegated to three others. The number went 52, then 69, then 221, then 25. Only the last was checked two independent ways.
dune test 2599/0, doc_coverage 221/221.
v0.1.361 — 2026-08-31
_A fourth correction, and the biggest: try_or exists._
v0.1.358 justified running each plugin in a child process like this: "None of that can be survived in-process. Mere has no way to recover from fail, and parse_and_eval reaches for it through 79 call sites in the evaluator and the parser." The claim was repeated in examples/plugin/host.mere, in scripts/plugin_host_check.sh, in examples/README.md, and in the commit.
Mere has recovered from `fail` since long before any of that. try_or : (unit -> 'a) -> 'a -> 'a catches Eval_error and returns the default. It is a builtin on the interpreter and lowers on C, LLVM and Wasm, and it is documented in docs/stdlib-reference.md with its signature and an example. The C backend's own __lang_fail_impl is written around it — the exit(1) that ends a failing program is the branch taken when no try_or is active. Nothing was hidden; it was not looked for.
Measured rather than argued: a host that wraps tokenize / parse / run in try_or survives the corpus in-process. calls_a_builtin_wrongly, does_not_parse and asks_for_the_filesystem are all caught, the host prints a verdict for each and exits 0.
What survives of the reason for a process is narrower, and it does hold. try_or returns a DEFAULT, not the message: a host built on it learns that a plugin failed and nothing about why, where the child's stderr carries unbound variable: write_file — the line the host reports and a user acts on. And try_or does nothing whatever about a plugin that does not stop; loops_forever and allocates_without_end need the OS either way. So the boundary is a process to keep the diagnostic and to bound the runaway, not because failure could not be survived. The corrected wording is now in all four places that carried the old one.
A related reading, checked before it was written down. The interpreter looked like it honours detach — a detached thread that fails leaves it at exit 0 and silent, where C exits 1. It does not. detach marks the thread so the leak report does not name it (v0.1.304), and the silence is the general one: an uncaught failure in a thread reaches the program only through join, which does re-raise it, on both the interpreter and C. There is no separate detach defect, and what is left is one question rather than two.
That question is unresolved and is now stated as one: what should an uncaught failure in a thread nobody joins do? The vocabulary for the distinction already exists — join asks for the result and its failures, detach disowns the thread, and neither is a leak the registry already reports. Three of the four backends answer differently and none of them consults that vocabulary. Recorded, not decided.
All four corrections in this arc are one mistake: a partial path reported as the whole. A second harness in the same file, an environment variable, two backends, and now a documented builtin.
dune test 2599/0, plugin_host_check 9 plugins, thread_fail_check 7 assertions.
v0.1.360 — 2026-08-30
_Two corrections to v0.1.358, and the real hole underneath them._
First correction. v0.1.358 said of the four-backend split in how a spawned thread's failure is reported: "scripts/parity.sh cannot see any of it: it discards the compiled backends' stderr, compares exit status only against its timeout, and skips a case outright when the interpreter exits nonzero." That is wrong. parity.sh has TWO harnesses, and the second one — "programs that are supposed to fail", over test/parity/fail/ — compares exit status, stdout AND the failure message across backends. It was built in v0.1.246 for exactly this purpose. Each clause of the sentence is true of the main loop; the conclusion drawn from them is false of the file, which was read as far as the first loop.
Second correction. The same entry said a fail inside a spawned thread "leaves the interpreter running and SILENT — no message on either stream." It is silent BY DEFAULT. builtin_spawn records T_died with the message and re-raises into a domain nobody joins, and MERE_THREAD_REPORT=1 names it: thread 1: died: fail: boom, never joined. The interpreter has not lost the failure; nothing asks it.
The hole is real and is at the seam. A program the interpreter FINISHES and a compiled backend does not fits in neither harness: the main loop compares stdout only and calls it MATCH, and the failure section refuses the file for exiting 0. Measured on the spawn case: interp exits 0, C exits 1, stdout byte-identical.
note_exit closes it. The main loop now compares exit status, pinned the way stdout differences are — <case>.<backend>.exit holds the status, and a pin whose case has stopped diverging is an error rather than a pass. The interpreter's status is 0 for every file that reaches there, so the expectation is 0.
Sizing that change is where a third correction belongs. The number reported before making it was "0 of 139 disagree", from a scan that compared the interpreter against C only — the same mistake as the two above, a partial path presented as the whole. Running all four found one: capture_after_call, which already carries an llvm.expected for the wrong output Q-052 describes, exits 1 under LLVM where the interpreter exits 0, deterministically in 15 runs of 15. The known miscompile reaches the exit status too, which the pin now says out loud.
And the case that found it still cannot go there. The defect destroys the determinism a pin requires. Measured: interp exits 0 in all 70 runs; C exits 0 in 5 of 60 and nonzero in 55; LLVM showed 2 zeroes in 20 and then none in 60. The proportions move with machine load. The first version of the gate asserted "C exits 1 every run", having watched exactly that 20 times, and failed on its next run. So scripts/thread_fail_check.sh holds it instead and asserts only what does not move: the interpreter always exits 0 and says nothing by default, the interpreter under MERE_THREAD_REPORT=1 still names the death and its message, C ends the process at least once in ten runs, Wasm always exits 0 with the failure in its output, and interp and C do not agree. LLVM is reported and not asserted, because asserting that something varies is flaky in the other direction.
Poisoned two ways: removing the fail from the spawned thread (four assertions fire), and making the interpreter report a dead thread by default (the silence assertion fires — which is how this gate finds out it was fixed).
dune test 2599/0, parity 139/139 with the exit-status comparison in (the run before the pin went in was 138 and 1), thread_fail_check 7 assertions.
v0.1.359 — 2026-08-30
_The boundary held on the path taken._ v0.1.358 said the interpreter is a capability boundary, and it is, but the refusal arrives when control REACHES the call. A plugin whose request sits in a branch it does not take was accepted and reported success — measured, not argued: if 1 > 2 then write_file ... else "harmless" returned "harmless", wrote nothing, and told the host nothing about the filesystem having been asked for. A host deciding whether to run something cannot use an answer that only arrives if it runs it.
contrib/parser/free_names.mere answers about the whole program instead: free_names : program -> str list, the names a program uses without binding, computed over the parsed AST. Binding forms are let through the whole pattern, let rec with the group in scope in its own bodies, fn's parameter, and a match arm's pattern in both the guard and the body. Constructors are not names here — the parser has rewritten them into EConstr / PConstr before this sees them, so Cons and Some never look like requests.
`extern fn` does not bind. That is not an oversight but the second evasion: the evaluator treats a TopExternFn as a no-op and leaves the name unbound, so a plugin declaring extern fn tcp_connect: str -> int -> int and never calling it runs to completion and returns 0. Nothing it does at runtime reveals what it asked the host for. Declaring an extern IS the request, and is counted as one.
The grant now comes from the host on the command line rather than from a constant inside the evaluator, and a plugin that asks for something outside it is refused whole — with every name it asked for, not the first one execution happened to reach — and exits 3, so the host can tell "asked for what it was not given" from "failed". The corpus is nine plugins and the gate counts four kinds, because a host answering "refused" nine times passes a check that only asks whether it survived.
The check that the verdict was READ and not RUN: both hidden plugins print a line before the thing that would give them away, and neither line may appear in the output. A host that evaluated its way to the same verdict would print it, and would have missed both.
Checked for false denials as well: shadowing an ungranted name, a let rec calling itself, a match arm's binder and a fn parameter are all bound, and none of them is reported.
Poisoned four ways, each caught by the assertion it belongs to: the runner not consulting the grant, the analysis not descending into an if (denied 3 of 4, and the hidden branch's line appears), extern fn treated as a definition (denied 3 of 4, and the declaring plugin's line appears), and the earlier poisons still hold.
dune test 2599/0, plugin_host_check 9 plugins (1 returned, 4 denied, 2 refused, 2 stopped).
v0.1.358 — 2026-08-30
_A plugin host is the one program that takes somebody else's code as input, and nobody had written one._ Half of what Mere claims about capabilities holds for free and the other half had never been asked. contrib/eval seeds its environment with print and show, so a plugin naming write_file does not reach it and the interpreter is the boundary — asked of the filesystem, not of a message: the corpus plugin that tries to write /tmp/mere_plugin_pwned.txt leaves nothing behind. But the denial and the host's death are the SAME event. There is no way to recover from fail, and parse_and_eval reaches for it through 20 call sites in the evaluator and 59 in the parser, so an unknown name, a type error and a syntax error all end the process that asked. [Corrected in v0.1.361: try_or recovers from fail on every backend and is documented. The design stands for other reasons; the sentence above does not.] Nor is there anything to stop a plugin that does not stop: no step budget, no cap on a region. The infinite loop and the runaway allocation ran until the operating system killed them.
spawn is not the way out either. A fail inside a spawned thread leaves the interpreter running and SILENT — no message on either stream — while C and LLVM end the process with the host's stdout lost 4 and 19 times out of 20, and Wasm prints the failure and carries on. Four backends, three answers, and scripts/parity.sh cannot see any of it: it discards the compiled backends' stderr, compares exit status only against its timeout, and skips a case outright when the interpreter exits nonzero. [Corrected in v0.1.360: parity.sh has a second harness that compares exactly those things. The hole is at its edge, not its absence — and the interpreter is silent by default, not silent.]
So the boundary is a process. examples/plugin is a host, a runner, and seven plugins — one that behaves, two that never stop, four that fail in the four available ways — and scripts/plugin_host_check.sh requires a verdict for each by KIND, since a host that answered "refused" seven times would satisfy a check that only asked whether it survived. The CPU bound comes from ulimit, which is the honest record: the mechanism is the shell's, not the language's.
Writing the host found a bug in the only failure channel it has. run returned 101 for a plugin killed at its CPU limit where the shell and the C backend both say 152, and 121 for SIGKILL against their 137. builtin_run had WSIGNALED n -> 128 + n, which is the right shape with the wrong n: OCaml's waitpid reports signals in OCaml's encoding, where Sys.sigkill is -7 and Sys.sigxcpu is -27. Both results land BELOW 128, where no caller testing for a signal will look. It was also intermittent in a way that hid it — adding exec to the host's command line, which should be invisible, turned "killed by signal 24" into an ordinary-looking status, because without exec a shell survives to report the death itself and the wrong arithmetic never runs.
Fixed without a signal table: the command now runs inside a subshell, so the shell being waited on always exits normally with 128+signal in the system's own numbering. scripts/run_status_check.sh checks six cases against a table of expected values as well as for agreement, because two backends that are both wrong agree perfectly. The WSIGNALED branch underneath is still in OCaml's encoding and is now merely hard to reach; a correct conversion needs a table that differs between Linux and macOS for eleven signals, so it is recorded rather than guessed at.
docs/host-matrix.md was corrected first, since it is the instrument used to decide what is missing. file_pread read refused on LLVM and Wasm because its probe opened the handle with file_open, and vec_new read refused on C and LLVM because the synthesized probe left the element type unresolved. Both have worked all along. A cell whose diagnostic names a different builtin now says unattributed instead of guessing, and there are 16 — lcm on RV32I is not missing lcm, it is missing abs.
Poisoned five ways: the old file_pread probe, the reverted subshell (three of six status cases fail, and the plugin host reports its two runaways as refused with no reason), write_file seeded into the plugin environment, which makes the pwned file appear, and a three-second wall-clock bound, which is how the runaways are held to something other than the good behaviour of the thing under test.
dune test 2599/0, host_matrix 150 builtins (16 unattributed), run_status_check 6 cases, plugin_host_check 7 plugins.
v0.1.357 — 2026-08-30
_G-1 assumes rendering is a function of its input; the language does not make it one._ time and random_int are unconditional builtins, and a function that reads the clock has the type str -> str like any other — measured, not assumed: mere -t on fn (x: str) -> x ++ show (time ()) answers (str -> str). Nothing seals them, so nothing stops a view from being nondeterministic, and render_agreement compares a server against a browser, which are two different environments.
scripts/render_purity_check.sh checks what the language does not prevent. The ambient set is the clock, the two PRNG builtins, and the environment ones (env_var, args, read_file, read_stdin, run) — the last group belongs because G-1's two sides share no environment, argv, filesystem or stdin.
The first version of this gate claimed something false and was corrected by measurement. It said it bounded what the render path can REACH, on the theory that dead code is eliminated. That holds inside the main file — a program whose only clock call sat in an uncalled function emitted zero occurrences of __lang_time — and does not hold across an import: an uncalled function added to contrib/html/build.mere put a call site into the driver's C. So the property is coarser and stated plainly now: the view modules contain no ambient use anywhere in them, called or not.
The prelude's own baseline is measured at gate time from a print-only program rather than written down, since getenv legitimately appears once in every program. Each module is anchored by a symbol that must be in the emitted C, so a module that stopped being pulled in fails instead of passing as the emptiest possible way to be ambient-free.
Poisoned three ways: a clock read inside an uncalled module function, an environment read in the tree, and dropping the leaf module's import.
dune test 2599/0, render_agreement 6 lines, htmlbuild 15 cases, determinism_check 153 cases.
v0.1.356 — 2026-08-30
_The derivation now says which schemas it cannot speak about._ derivable refuses a READ whose freshness no write determines. schema_unmodelled is the same admission one level up, and both holes it names were measured in Postgres before it existed:
- Trigger. An
AFTER INSERT ON poststhat writesauditturned
INSERT INTO posts into a change of a read of audit — "post touched" → "post touched,post touched". The statement never names that table, and a trigger body is arbitrary PL/pgSQL: there is no honest static answer, only a refusal.
- View.
SELECT ... FROM published, a view overposts, changed on an
insert into posts: "one" → "one,two". The read names the view, the write names the table. A view is expandable in principle — its definition is in the DDL — and until that is done, saying so beats missing it.
Rules are refused for the trigger's reason. The scan matches only after CREATE (skipping OR REPLACE, MATERIALIZED, CONSTRAINT): the bare word found RETURNS trigger AS inside a function body and reported a trigger named as, which refuses too much and hides the real finding.
The gate asks pg_trigger and pg_class whether each fixture has one, so the refusal is judged by the database rather than by the file that wrote the fixture. Both directions are poisoned: refusing nothing fails with_trigger and with_view; refusing everything fails every schema that has neither, since a schema needlessly refused costs every channel in it.
18 fixtures. dune test 2599/0, live_query 53 cases, live_soundness 42 pairs (0 unsound).
v0.1.355 — 2026-08-30
_The DDL parser's coverage is now Postgres's answer, not a list someone remembered._ cascades_of reads foreign keys out of DDL text, and twice the set of spellings it survived was enumerated from memory and twice something was missing. v0.1.353 was the first: SET NULL and SET DEFAULT skipped under a note saying they "do not remove rows". Probing the spellings found two more:
- `CREATE TABLE IF NOT EXISTS c` made the child `if`. The name was taken as
the word after TABLE. A pair was recorded between the parent and a table that does not exist, the real child was never widened, and affects_via answered false for a DELETE that empties it — unsound. ALTER TABLE ONLY c, which is what pg_dump writes, had the same shape.
- Quoted identifiers kept their quotes.
CREATE TABLE "c"derived"c"
while SELECT ... FROM c derived c, so a schema written by a tool that quotes never matched a query written by hand. Both sides go through _clean, so stripping there fixes both at once.
scripts/cascade_catalog_check.sh is the part meant to outlast this. Fifteen fixtures in test/live/cascade_fixtures.mere are each built in a real Postgres, pg_constraint is asked which foreign keys exist and what they do, and that is compared to what the parser derived from the same text. The one judgement the gate makes is which actions count: c/n/d change the child's rows, a/r refuse the parent's write instead. Everything else is the catalog's answer, so a spelling nobody here thought of is still judged.
All three defects poison the gate: reverting the IF NOT EXISTS fix fails if_not_exists and alter_only, reverting SET fails inline_setnull and inline_setdefault, reverting the quote strip fails quoted.
dune test 2599/0, live_query 53 cases, live_soundness 42 pairs (0 unsound), cascade_catalog 15 schemas.
v0.1.354 — 2026-08-30
_The broadcaster never let a subscriber go._ vec_push was the only thing that ever happened to the subscriber list. Measured on the example server: 20 subscribe-and-disconnect cycles took the process from 8 open descriptors to 28, another 20 took it to 48. Every subscription cost a descriptor forever, every publish afterwards wrote to a socket nobody was reading, and the server would have stopped accepting at the descriptor limit.
A subscriber that left is discovered by writing to it. The write fails, the fd is closed, and the slot is marked free — fd = -1 rather than a shorter vec, because Mere's Vec has push, get and set and nothing that shrinks it. The list is then bounded by the PEAK number of concurrent subscribers instead of the total ever seen.
The quiet channel is the half that needs a heartbeat. A subscriber on a channel nothing is published to is never written to, so it is never discovered to be gone. The pump waits with channel_recv_timeout instead of blocking on channel_recv, and on each timeout writes an SSE comment (: then a blank line) to everyone — which is also what keeps a proxy from dropping an idle stream.
Measured after: 8 descriptors, 8 after 20 disconnects, 8 after another 20 on a quiet channel. 50 concurrent subscribers all received a publish, and all 50 descriptors came back.
The gate counts descriptors before and after ten disconnects on a quiet channel, through /proc or lsof, and fails rather than skipping when it can count neither. Removing the reap fails it; leaving the reap but stretching the heartbeat to ten minutes fails it too.
dune test 2599/0, live_query 53 cases, live_e2e 6 checks, thread_leak 7/7.
v0.1.353 — 2026-08-30
_Three of the four row-changing referential actions were missed._ cascades_of recorded a pair only for ON DELETE CASCADE, under a note saying "RESTRICT / SET NULL do not remove rows". True, and the wrong question — what matters is whether the READ changes. Asked of Postgres, one row in the child, parent deleted:
| before | after | |
|---|---|---|
ON DELETE CASCADE | 1:1 | (empty) |
ON DELETE SET NULL | 1:1 | 1:null |
ON DELETE SET DEFAULT | 1:1 | 1:9 |
ON UPDATE CASCADE | 1:1 | 1:77 (on UPDATE of the key) |
ON DELETE RESTRICT | the write is refused; nothing changed |
Each miss is an unsoundness: a write changes a read and nothing says so. The scan now runs to the end of the column definition rather than a fixed three-word window — ON DELETE RESTRICT ON UPDATE CASCADE puts CASCADE sixth, and the old scan stopped looking before it got there.
RESTRICT and NO ACTION stay excluded, and that is the negative that separates this from returning a pair for everything: a refused write changes nothing, so recording a pair for it would wake channels that cannot have gone stale.
Eight string cases in test/live/derive_cases.mere (53 total), and — because a string case is only this file's claim about what Postgres does — a tags table with ON DELETE SET NULL in the soundness fixture, so Postgres judges it. With the fix reverted the gate names it: UNSOUND w_post_del -> r_tags.
dune test 2599/0, live_query 53 cases on 4 backends, live_soundness 42 pairs against a real database (37 exact, 5 wasteful, 0 unsound), live_e2e 6 checks.
v0.1.352 — 2026-08-30
_The leak report counted threads it had not seen start._ thread_leak_check went red on a loaded runner: a program that leaks two threads reported one. It passed 300/300 locally, including under load — the shape of a race, not a wrong program.
Threads registered themselves from inside the spawned domain, so spawn returned before the entry existed and a main that spawns and exits could reach at_exit first. An undercount is the bad direction for a leak report — fewer leaks reads as a cleaner program, so the diagnostic failed quietly in the direction that gets believed.
Two tables now, because there are two questions. Did this thread exist is the parent's to answer and is known the moment Domain.spawn returns. What was it doing is the child's and is not knowable until it runs. A domain that never reached its first instruction is reported as spawned, never ran — which is what is actually known about it, rather than a state nobody measured.
The gate splits the same way. The count is checked on the real clock, where it must always hold. The wording is checked under MERE_VIRTUAL_CLOCK, whose rule is that time advances only when every live thread is parked: a sleep_ms in the program returns at a moment when the workers have reached their blocking calls, so blocked on channel_recv is a fact rather than a snapshot. With a 50ms delay injected into every child — the loaded runner, reproduced — the gate is 7/7, and it still goes red for both the old undercount and a wrong state word.
dune test 2599/0, parity 139/0, thread_leak 7/7, virtual_clock 6/6.
v0.1.351 — 2026-08-30
_A Mere program holds its own connections open._ scripts/live_e2e_check.sh already ran subscribe → write → receive over a real socket, but on Wasm under Node — the negative it proved was proved by a JavaScript host, leaving open the question of whether a Mere program can hold connections open by itself.
Two builtins on the C backend answer it. http_hijack takes a connection out of serve_mt's hands (returns the fd, marks the response sent) so the worker neither writes nor closes; http_hijacked lets the two close sites ask. Both _Thread_local, reset per request. contrib/http/sse_native.mere is the broadcaster on top: subscribers' fds live inside the broadcaster thread, and scripts/sse_native_check.sh runs the loop against a 145K native binary.
The gate has two negatives, because two different wrong implementations pass a check that only asks "did the subscriber get it". A broadcaster ignoring the channel delivers to everyone — so a subscriber on other must receive nothing. A broadcaster keeping one subscriber instead of a list delivers to the most recent — so three subscribers on posts must all receive. Both are poisoned and both go red.
The second was not hypothetical. channel_recv_opt reads like a try-receive and is not: it is while (len == 0 && !closed) cond_wait, and None means closed, never empty. A pump polling two channels with it parks in the first, which looks like a delivery bug rather than a deadlock because the first channel's traffic keeps waking it. docs/stdlib-reference.md now has the table of which of the three receives ends its wait on what. The other seven uses in the tree were already the correct shape — a worker loop ended by channel_close.
dune test 2599/0, parity 139/0. Gates run locally: sse_native (new), live_e2e, live_query, htmlbuild, mount, extern_host, wasm_stub.
v0.1.350 — 2026-08-30
_Three builtins stop lying on Wasm._ run, env_var and file_exists answered constants — 127, None, false — under notes saying a browser has no subprocess, no environ and no filesystem. True of a browser. False of scripts/run_wasm.js, which has had system, getenv, file_size, read_file and write_file all along, and false of docs/host-matrix.md, which called all three yes.
file_exists was the worst of them: it did not fail, it answered "no", so a caller's if file_exists p then … else … took the wrong branch with nothing anywhere to say why.
All three now go through the host, the way args was wired in v0.1.159 for the same reason — the runners had supplied arg_count / arg_get for a long time and only the builtin was never connected. Checked against the C backend as the oracle: run "echo x" prints x and exits 0 on both, env_var answers the same value and the same None, and file_exists "/etc/hosts" is true on both.
The host had a hole too, and the compiler's hid it. getenv in run_wasm.js read the variable, built the buffer, and fell off the end without returning — every lookup answered undefined, which arrives as 0 and reads as "not set". Nobody noticed because the Wasm backend never called it.
test/parity/env_var.wasm.expected is deleted: parity reported that the pin no longer diverges, which is the retirement condition a DIVERGE is supposed to carry, doing its job.
v0.1.349 — 2026-08-29
_One sort, four backends._ vec_sort was an insertion sort on C / LLVM / Wasm and OCaml's Array.sort on the interpreter, and both halves of that were wrong in a way the parity suite could not see.
The asymptotics differed by backend. O(n²) compiled, O(n log n) interpreted. And because the comparator is curried, every comparison built its inner closure's environment in the region and never freed it — so the compiled backends' memory was O(n²) too. Measured, sorting a Vec[R, int]:
| n | before | after |
|---|---|---|
| 10,000 | 0.59 s, 569 MB | 4.1 MB |
| 20,000 | 1.44 s, 2.3 GB | 7.3 MB |
| 40,000 | 5.07 s, 6.7 GB | 14 MB |
| 200,000 | killed by the OOM killer at 82 s | 0.45 s |
And they disagreed about stability, which is observable from a Mere program: Array.sort is unstable and an insertion sort is stable, so seven pairs keyed by fst came out 600 400 200 300 700 500 100 interpreted and 200 400 600 100 300 500 700 compiled. No parity case had ever sorted with a comparator that returns 0 for two different elements — the only input that can tell.
All four now run the same bottom-up stable merge sort, comparing cmp(right, left) < 0 so a tie keeps the earlier run. The scratch buffer is malloc/free on C and LLVM (a temporary the program cannot observe, and a region allocation would leave n slots behind per call); Wasm takes it off the bump allocator, which has no free, the same way vec_push already abandons its old buffer.
The comparison sequence is part of the contract, not an implementation detail. A comparator is an arbitrary closure — it can count, print, or write — so four backends agreeing on sorted output while running two different algorithms is a divergence waiting for someone to put a side effect in one. test/parity/vec_sort_stable.mere therefore prints the comparator call count (8192 for its 1024 elements) next to the permutation's checksum: the number is the algorithm's fingerprint, and it fails if any backend changes algorithm alone.
A third divergence fell out of the rewrite, and it was a wrong answer. The LLVM helper called the comparator closure as i32 while Mere's int is i64 there. -1 / 0 / 1 survive that truncation, so it had never shown; a comparator written a - b on values more than 2^31 apart does not. Measured on the pre-change compiler, sorting [5000000000, 1000000000, 9000000000] ascending: interp, C and Wasm printed 1000000000 5000000000 9000000000 and LLVM printed it backwards. The parity case carries that input now.
What remains is the per-comparison allocation itself: cmp.fn(env, a) still materialises the inner closure. n log n of those instead of n² is a rate, not a fix — removing it needs the closure value to carry an uncurried entry point, the value-level twin of the __direct fns from v0.1.52 / Q-066.
_Two smaller repairs found by the same probe._
The interpreter's `vec_push` was `Array.append` — a full copy per push, so filling an n-element Vec cost O(n²) there while every compiled backend doubled. The interpreter is the parity oracle, and its asymptotics bound the input size any differential gate can afford: 200k pushes took 74 s interpreted and 0.01 s compiled. V_vec now carries capacity (vecbuf, the shape bytebuf already had). A 200k-element program that reads, sorts and prints went 77 s → 2.8 s.
An integer literal out of range escaped the compiler as an uncaught OCaml exception — Failure("int_of_string"), no location, no file name. Every literal is held in the host's 63-bit int, so 4611686018427387903 is the ceiling on all backends, including the ones that compute in 64 bits. Worse was the hex form: OCaml's parser wraps rather than failing there, so 0x7FFFFFFFFFFFFFFF lexed as -1 — a bit mask, which is the reason hex literals exist here (v0.1.46), silently becoming a different number. Both now name the limit at the literal's own location.
v0.1.348 — 2026-08-28
_What the capture analysis actually catches._ No behaviour change; a correction and three pinned facts.
spawn's capture analysis was reported here as not catching a mutable Map shared across threads. It catches it. The measurement that said otherwise used mere -t, which does not run Move_check — only mere -c, the run path, and Pipeline.process do. Asked the right way:
spawn (fn () -> map_set hits …) — the Map held directly | refused, "neither Send nor Sync" |
spawn (fn () -> h ()) where h captured it | accepted |
the same, with h arriving as a parameter | accepted |
Map is false for both is_send_v and is_sync_v; the classifier was right all along. The real hole is narrower and sharper: a captured value is classified by its type, and TyArrow _ -> true for both Send and Sync — a function type says nothing about what the function closed over. So the Map is caught one level up and invisible one level in, and a request handler is exactly one level in. That is why http_serve_mt cannot vouch for what a handler shares, and why mere-blog still had to move its sessions into the database.
All three are pinned in test_basic.ml, including the two that pass: a defect asserted as accepted makes the suite fail the day it is fixed, which is when the record needs updating.
Closing it means classifying a closure by its capture set — a type-system feature, not a patch. A syntactic approximation (recurse into a captured name's definition when it is visible) catches the second row and not the third, and errs toward refusing programs that work today: mkv, contrib/http2 and serve_mt all hand closures to spawned workers. The routes that must not be closed need enumerating before any are.
v0.1.347 — 2026-08-28
_Several processes can share a port, if you ask._ tcp_listen_shared, and MERE_HTTP_REUSEPORT=1 in both servers.
Measured first: a second process could not bind at all. SO_REUSEADDR — which tcp_listen has always set — lets a restart rebind through TIME_WAIT; it does not let two live processes share a listener. So "sessions live in the database now, which means more than one process could serve" was half a claim: the sessions could, the server could not.
SO_REUSEPORT does, and the kernel spreads connections across the processes. Off by default, and that is a decision rather than an oversight: with sharing always on, a deploy script that starts a second copy by mistake does not fail — it quietly serves half the traffic from the old binary, alternating, which is the hardest kind of wrong to notice. Failing to bind is the better default.
Verified in C before Mere, in both directions, and that order mattered: the Mere side looked broken and was — a string replacement that did not match, read as a working build because the program still compiled and ran.
v0.1.346 — 2026-08-28
_A dead database stops looking like an empty one._ contrib/db/pg.mere.
Measured with a real Postgres in a container, restarted under a running mere-blog: GET /api/posts answered `[]` with a 200. The log said terminating connection due to administrator command; the API said success. A client cannot tell that from the truth, and a cache would store it.
_collect_result printed the error and carried on, returning the rows it had — none. Severity decides now, and the distinction is load-bearing:
- ERROR is a statement that failed on a connection that is still usable. A unique
violation on signup is this, and failing it would turn "that name is taken" into a
- Those still return, exactly as before.
- FATAL / PANIC mean the server has already closed the connection. There is
nothing to return and no honest way to say so in a row list, so it fails — which contrib/http/rescue.mere turns into a 500.
The V field is read rather than S: S is localized.
Downstream, mere-blog now holds each worker's connection in a slot it can replace, so a failure drops it and the next request redials. Four workers cost exactly four failed requests and then serve again — one per worker, each discovering its own dead connection once. Not zero: a request in flight when the server went cannot be saved, and its check says so rather than pretending.
v0.1.345 — 2026-08-28
_A request body stopped at its first zero byte._ Every binary upload was truncated, silently, and every test passed.
contrib/http/multipart.mere had been written, documented and never run. Running it found a defect that is not in it: http_current_body built the body with __lang_str_of_cstr, which is strlen. A 52-byte file with forty zeroes inside arrived as its first 8 bytes. Text bodies have no zero, so JSON worked everywhere and nothing ever said otherwise. The length was known the whole time — __http_req_body_len — and thrown away.
Then the checksum agreed with the bug. With the length fixed, the file was still wrong: sha256_hex also used strlen, so it hashed the same 8 bytes and reported a digest that matched a truncated file. The instrument and the subject failed the same way, which is the only arrangement in which a wrong answer looks confirmed.
Three helpers take arbitrary bytes and had this: sha256_hex, hmac_sha256_hex_str (its message), and pbkdf2_sha256_hex (its password). All now use __lang_str_size. The hex helpers beside them keep `strlen` on purpose — their input is hex text with no zero byte, and asking __lang_str_size about a plain C string reads a header that is not there. Verified against outside references (shasum, openssl dgst, hashlib.pbkdf2_hmac): all four values unchanged for text, so nothing that already worked moved.
scripts/http_upload_check.sh compares a SHA-256 taken by the server against one taken by the shell, on a file chosen to be binary in the way that matters, and separately pins sha256_hex on a string containing a zero.
v0.1.344 — 2026-08-28
_A stuck handler stops being an outage._ contrib/http/serve_rd.mere sheds.
A handler that never returns cannot be interrupted. There is no way in this language to take a thread away from a loop it will not leave, and this release does not pretend otherwise — the stuck worker stays stuck for the life of the process. What it changes is what happens to everyone else: with every worker busy and the queue full, a new request is answered 503 immediately rather than joining a line that may never move.
Measured, two workers held by a 20-second handler with four requests already in:
| the next request | |
|---|---|
MERE_HTTP_QUEUE=2 | 503 in 0.03 s |
MERE_HTTP_QUEUE=999 (the control) | still hanging at 6 s |
The control is the check. Without it, "the fifth request got a 503" says nothing about shedding — and the same binary answers both, so the comparison is between two runs of one program.
The queue defaults to the worker count: a short burst is absorbed, a stuck server starts saying so within 2N requests instead of accumulating clients. It says so on stderr once per transition, not once per request — a log line per shed request is what turns an incident into two incidents.
v0.1.343 — 2026-08-28
_Keep-alive, and the slot it costs instead._ http_response_head_ka in the runtime, and contrib/http/serve_rd.mere reuses a connection instead of closing it.
Q-081 said keep-alive could not be added to the worker pool, because an idle connection would hold a worker and going quiet is free for an attacker. Under the readiness loop it holds a descriptor — cheap, and still finite, which is the part a deployment feels: MERE_HTTP_MAX_CONN silent clients fill the table and nobody else gets in. So the connection also carries a last-touched time, and anything silent past MERE_HTTP_IDLE_MS (15 s) has its slot taken back.
The reaper only runs when the loop runs, and the loop blocked indefinitely with nothing in flight — so the first version reaped nothing and the check failed for a reason unrelated to reaping. The wait is three cases now: 5 ms with a request in flight, 1 s with connections open and none in flight, and indefinite with neither.
Connection: close is obeyed, and so is HTTP/1.0. Pipelining is refused rather than mishandled: the response is written back into the request's slot, so a second request already in the buffer is gone — a connection that arrived with extra bytes is closed after its first answer.
A grep that did not match read as a feature that did not work. The first check counted Re-using existing connection in curl's trace and got zero; curl 8 says Re-using existing connection with host. Keep-alive had been working the whole time. The check matches re-using, case-insensitively, now.
v0.1.342 — 2026-08-28
_One bad route no longer takes the server with it._ contrib/http/rescue.mere.
And the write side goes through the loop too. The worker writes what the socket takes and hands the remainder back — fd, offset, bytes left — then takes the next request. A client that asks for something and then refuses to read it costs a map entry, the same as one that dribbles its request.
Finding that out took a control, and the first two attempts had none worth having: six impolite clients against four workers answered in 0.03 s under both the old blocking write and the new handoff, because a 28 KB response fits in the socket buffer and never blocks anything. Raising the response to 896 KB separated them, and not in the way the change was made for — the old code was wrong, not slow. The descriptor is nonblocking, and it called tcp_write once and ignored the short count: 3 271 270 bytes delivered for an 896 KB response. The new one delivers 896 000, byte for byte.
So the gate checks content at a size that cannot go out in one write, not timing. rd_max_fd and rd_slot are read from the environment now, because a property that only appears at one size cannot be checked at another — and because the slot is both the connection budget (the arena is 16 MB with no free) and the response ceiling. 128 KB by default: the first guess of 32 KB is too small for a JSON list of rows. http_send_file never enters this buffer, so static assets of any size are unaffected.
A handler that fails — read_file on a path that is not there, a refused decode, an index out of range — unwound past the server loop and ended the process. Measured on the readiness server: GET /ok answers, GET /boom is a failing route, and the next GET /ok cannot connect, because nothing is listening any more. One request path took the whole site down, and it is the path nobody tested.
Not a property of any server shape: all three behaved this way, because none of them caught anything. The language has try_or; contrib/http had simply never used it. grep -rn try_or contrib/http/ was empty — the same one-line question that was empty for io_poll yesterday.
let _ = http_serve_rd 8080 8 (with_rescue (router routes not_found));
try_or thunk default takes a value, not a second thunk, so a default of let _ = http_set_status 500 in "..." would set 500 on every successful request too. The result is wrapped in an option instead: the default stays pure (None) and the status change lives in the branch only failure reaches.
What it does not catch, said plainly: try_or catches a Mere fail. A stack overflow, a segfault in an extern, or a worker killed by a signal are not that. What this buys is that a bug in one handler stops being an outage. The failure message goes to stderr — it can carry a path, a query or a row — and the client gets five words.
The gate runs the same binary with the middleware off first, and the control has to show the server dying, or "the rescued server survived" is not evidence about the middleware.
v0.1.341 — 2026-08-28
_Readiness for the I/O, workers for the handlers._ contrib/http/serve_rd.mere.
The capability was not missing. The adoption was. io_poll_new/add/mod/del/wait/get and io_set_nonblocking landed in v0.1.313, under a commit message that read "every server so far either blocked or spawned", with a dogfood — mpoll, 193 lines, "one thread, many connections" — that serves HTTP over a readiness loop and holds each connection's progress as a value in a Map, because a connection with no thread has no stack to remember where it was. contrib/http, the module every web program here actually imports, then went on offering exactly the two shapes mpoll's header called insufficient. v0.1.340 added the second one again, from scratch.
Neither half is enough, and this is the measurement rather than the argument:
| slow client (8 stalled connections) | slow handler (8 × 400 ms) | |
|---|---|---|
http_serve_mt (8 workers) | no answer in 5 s | 0.47 s |
| readiness loop alone | fine | would serialise |
http_serve_rd | 0.04 s | 0.46 s |
A client that connects, sends half a request and goes quiet holds a worker under the pool — and opening a socket and saying nothing is free for an attacker, where making a handler slow is not. Under the loop it holds one map entry. Q-081, which recorded that keep-alive could not be added to the worker pool because an idle connection would hold a worker, stops describing anything: under this loop an idle connection holds nothing.
The gate runs both servers against the same eight stalled clients, and the control is the check: the worker pool must FAIL to answer, or "the readiness server answered" says nothing about readiness.
`contrib/db/pg_pool.mere`'s premise expired, and the note says so rather than being quietly edited: it opens by explaining that "every http_serve handler runs to completion before the next request is dispatched", which was true of http_serve and is not true of the other two. Two handlers sharing one fd interleave on a Postgres socket and corrupt the protocol, so on a concurrent server that module is not a pool. What to use instead is per-worker state, which is not a bag of connections at all.
Still blocking on the write side: a client that sends a request and refuses to read the answer holds a worker. Fixing it means handing the fd back to the loop with a Writing state — which is what mpoll does and what this should grow. The read side is the half that matters first, because dribbling a request is the cheap attack.
v0.1.340 — 2026-08-28
_The server answers more than one request at a time._ contrib/http/serve_mt.mere is an accept loop with a fixed worker pool, and the measurement that prompted it:
| 8 requests, 400 ms handler | |
|---|---|
http_serve (the C loop) | 3.27 s |
http_serve_mt, 8 workers | 0.45 s |
http_serve is accept → read → handle → write → close, sequential. Nothing said so and nothing measured it, so "Mere serves HTTP" was true while "Mere serves a web application" was not — a handler doing a 50 ms query capped the process at 20 requests a second.
The pool is written in Mere, not added to the runtime. spawn, channel_new, channel_recv_opt — the same shape contrib/http2/server.mere arrived at, including its reason for a fixed pool rather than a thread per connection: mem_alloc is a bump allocator with no free, so a buffer per connection leaks one per connection. The per-request state the http_set_* externs write is now _Thread_local, and four new externs (http_begin_request, http_response_head, http_response_sent, http_end_request) let Mere drive a request through it. The handler type is unchanged, so router and every middleware written against it drop in.
A claim in the first draft of that module was false, and the correction is the finding. It said the compiler would refuse a handler sharing a mutable Map across requests, because spawn's capture analysis sees what a spawned closure captures. Measured in three steps: spawn capturing a Map directly compiles. So does a closure capturing one, so does that closure arriving as a parameter. Map is not !Send — which mkv wrote down when it chose to route every access through one owner thread, and then stopped being a live question because mkv had worked around it. The module comment now says what is true: moving a handler to the pool turns state that was safe because the loop was sequential into state that races, and nothing will tell you. Recorded as Q-080.
http_serve_mt_ctx adds per-worker state: make_ctx worker_index is called once per worker and its result is passed to every request that worker serves. That is what a connection pool is here — not a shared bag of connections with checkout and return, but state that never crosses a thread boundary, so a handler that fails cannot fail to give one back, and the number of connections equals the number of workers by construction.
The context is an int, and that is a real limit rather than a simplification: a polymorphic context does not compile, because the capture analysis refuses to move a value of unknown type across a thread boundary and Mere has no way to write a Send bound on a type parameter (sync type declares a concrete type, not a constraint). Worth noting that this is the analysis working — it refuses what it cannot prove, and separately fails to classify Map, which is Q-080.
scripts/http_concurrency_check.sh runs the same binary twice — eight workers and one — because a machine fast enough to make the sleep irrelevant would pass the concurrency check for the wrong reason. The control is the check.
(Its first draft hung: a bare wait waits for the server too, and the server does not exit. A timing gate that can hang is the one thing it must not be.)
v0.1.339 — 2026-08-28
_A fix for something v0.1.338 broke and shipped._ http_serve_tls moves out of contrib/http/http.mere into contrib/http/tls.mere.
Putting it in http.mere made every contrib/http program require OpenSSL, plaintext ones included. The C backend links OpenSSL into a program that declares a TLS extern, and an import declares everything the imported module declares — so import "http/http.mere" was enough. Before v0.1.338 a plaintext contrib/http server built with no OpenSSL at all; v0.1.338 took that away, and it was released before anyone noticed.
Two reasons it got through, both worth naming. The app that would have shown it — mere-blog — already links OpenSSL for Postgres, so its build line did not change. And the property had no check: nothing in the tree asserted that a plaintext server stays plaintext-only. That check exists now, on the emitted C rather than on whether a link succeeds, so it asks the same question on a machine that happens to have libssl installed. Putting the extern back turns exactly one of eighteen checks red.
The dependency is opt-in again, by importing the module that declares it:
import "http/http.mere";
import "http/tls.mere"; // this line is what links OpenSSL
Found while trying to measure something else entirely — a plaintext test server would not build. The measurement it interrupted is worth recording too: the native `http_serve` handles one request at a time. Four concurrent requests against a handler that sleeps 400 ms take 1.66 s, not 0.4 s. The accept loop is accept → read → handle → write → close, sequential, with Connection: close hardcoded. That is the largest remaining gap for serving a real web application, and it is now a number rather than an impression.
v0.1.338 — 2026-08-28
Twenty-one gates in CI had not run for weeks, and CI said "skipped". component_parity invokes scripts/build-component.sh, which declares #!/usr/bin/env bash and uses set -o pipefail — and it invoked it with sh. On macOS /bin/sh is bash, so it was green here; on Ubuntu /bin/sh is dash, which has no pipefail, so all three component builds failed. Because it is a step in the middle of the workflow, every step after it — proto, graphql ×4, http2, hpack, grpc, window, html tokenizer, infer scaling, sourcemap, debug info, downstream_check, qemu_virt, and the new tls_server_check — was reported as skipped, which reads like "nothing else was wrong" and means "nothing else was asked".
Two fixes, because the second is the general one:
- the call goes through the script's own shebang. Reproduced by putting
dashon
PATH as sh: the old form fails 3 of 3, the new form passes 3 of 3.
- every gate step now carries `if: ${{ !cancelled() }}`, so one red gate no longer
turns the rest into unknowns. A red CI should be one known failure plus everything else measured, not one known failure plus everything downstream unmeasured. Setup steps are deliberately not marked: a failed toolchain install should make the gates below it fail loudly rather than silently not exist.
This is the same shape as v0.1.271's _GNU_SOURCE: a difference between the development machine and the CI machine, verified on one platform, believed on both.
And running them showed `downstream_check` had been wrong about four repos. With the rest of the workflow finally executing, it reported that mpng, mq, mere-blog and mere-ruby "no longer compile against this compiler". They compile fine — the CI step cloned each repo and never resolved its dependencies, so it was reporting a missing .mere_modules as a language regression. It passed locally because the development checkouts already have one. Fixed by running mere install after each clone, and by cloning with --recurse-submodules (mere-ruby carries contrib/mgz as a submodule, and a --depth 1 clone leaves the directory empty). Verified by cloning all four into an empty directory: three fail without mere install and pass with it, the fourth needs the submodule.
The if: sweep needed a second pass for the same reason a fix usually does: the first one matched run: sh scripts/ and missed the two steps that set an environment variable first — downstream_check and qemu_virt, one of which was the failure and the other the step it was hiding. 44 of 44 now.
mere install ended up inside the gate rather than beside it: putting it in the workflow file would have left the CI copy and the local copy able to drift apart again, which is the whole shape of this entry. The script now resolves a repo's dependencies when it has a mere.toml and no .mere_modules — and only then, because an existing one is the developer's and mere install would rewrite their mere.lock. Verified in both directions: fresh clones into an empty directory compile, and the development checkouts are untouched afterwards.
With that, `mbigfmt`'s deferral stopped being true. It was listed as deps — "needs mere install first" — and now the gate runs one. A skip outlives its reason silently, so the row was retired instead of left saying something false: 13 repos checked, 0 deferred.
_A Mere program can answer a TLS connection._ tls_server_init (load a certificate chain and key, once per process) and tcp_accept_tls (handshake on an accepted fd) join tcp_starttls / tcp_starttls_verified on the C backend. After the handshake, tcp_read / tcp_write / tcp_close route through the SSL* unchanged, so a plaintext handler becomes a TLS handler by inserting one line.
The gap was invisible because nothing was in its place. Every web program in this project served plaintext, and no README, doc or comment recommended a proxy or mentioned TLS for serving at all — mere-blog's README says ./mere-blog # serves :8080 and moves on. There was no workaround to notice and no advice to read past. The client half had existed for a long time, which made the surface look complete: grep tcp_starttls finds TLS, and TLS is there.
(The first draft of this entry, and the three commit messages under it, said these programs "told the reader to put nginx in front". They do not. That sentence was invented to explain the gap and then written down five times as though it were the record — the commit messages are pushed and stay wrong. What actually happened is quieter and is the more useful observation: a capability nobody asks for produces no workaround, and therefore no trace.)
test/tls/https_server.mere is a TLS-terminating HTTP server that refuses both of the available workarounds: it terminates TLS itself, and it reads its certificate paths and port from the environment (v0.1.337) rather than compiling them in. scripts/tls_server_check.sh drives it with curl and openssl s_client — two TLS implementations that are not ours and cannot be made lenient from this repository, and notably without `-k`, so the certificate presented is checked to be the one loaded.
Three defects in the gate itself, found by poisoning it, all of the same family — a check that reports something other than what it names:
- asserting curl's
%{ssl_verify_result} == 0stayed green under a poison that removed
the handshake entirely, because that field is 0 when curl never connected. Replaced by the assertion that gives verification meaning: the same request without the CA must fail.
- "the listener survived the failed handshake" was
echoed rather than tested, and was
printed by a run in which the listener had died. It now asks for another answer.
set -eplus akillon an already-exited server made the script stop after check 2
and exit 0 — five PASS lines, no summary, green CI. The gate now asserts how many checks ran.
And one in its bounds: with both key checks removed the server started successfully with a mismatched key, and a foreground run that was only ever fast because it was expected to fail became a hang. Both foreground runs are bounded now — a wrong answer should be a FAIL, never a wait.
Recorded because it is the opposite of what poisoning usually shows: the subject was fine and the instrument was not.
contrib/http gains http_serve_tls port cert key handler — the same accept loop with the handshake in front of it. The handler is byte-identical to the plaintext one, because every byte of a served connection now goes through one __conn_read / __conn_write / __conn_close triple instead of calling read / write / close directly in six places. `http_send_file` was the sixth: it writes straight to the connection rather than returning a body, so converting only http_serve would have left file downloads putting cleartext into a TLS socket — a half-fix that looks finished because every page still renders. Poisoning exactly that site turns exactly one of the thirteen checks red, which is why it is worth a check of its own.
Declaring http_serve_tls is what links OpenSSL; a program that never mentions TLS is unaffected. Declared-but-unreachable is an error message, not cleartext on a port the caller believes is encrypted.
And the other host had been broken for months. contrib/http has two implementations — the C runtime and scripts/run_http_server.js — and adding http_serve_tls to the second one meant starting it, which nothing in CI had ever done. It could not instantiate any Mere module: it hand-copied run_wasm.js's env imports under a comment saying it reused them, so when v0.1.277 added __lang_float_of_str_ok to the refusal enumeration only the maintained copy got it. The same paraphrase had __lang_float_of_str calling bare parseFloat where run_wasm.js calls a Mere-specific parser — so the two hosts disagreed about what a float literal is, quietly, for as long as both existed. scripts/mere_parse_float.js is now one copy that both require, and three checks run the Node host: that a module instantiates at all (the assertion that was false), that it serves, and that it terminates TLS.
Not on every backend, and unchanged in that respect: TLS is C-backend only, client half included. Plain Wasm also still answers None to env_var on every host, including Node, which has an environment — so a program configured through the environment takes its defaults there. Recorded rather than fixed: making it a host import would mean every browser host must bind it or fail to instantiate, which is the same failure class as the one above.
v0.1.337 — 2026-08-28
_A native binary could not read its own environment._
env_var existed on the Wasm component backend and nowhere else. So mere -c and mere -ll — the single-binary deployment the web dogfood advertises as its best story — produced programs configurable only by editing their source. mere-blog hardcodes its database host, port, user and name, and that is not the app being lazy: it is the only thing the backend allowed.
C and LLVM have it now. env_var reads yes | yes | yes across the compiling backends, with RV32I refusing as it should — there is no environment on bare metal.
Both backends got the same thing wrong first, and the second one had already been written down. getenv hands back a bare C string; this project's str carries a length header in the bytes before the pointer, so the environment's own pointer is not a str. C compiled, answered the unset case correctly, and died with "out of memory" on the first program that found the variable. __lang_str_of_cstr copies it into the region — which also means a later setenv cannot move the string out from under the value.
Then LLVM: the payload slot of an option holds a pointer to the value, not the value (Phase 25.0 boxed variant payloads). Storing the string pointer directly made the reader dereference the string's first eight bytes as an address. It matched Some fine and crashed the moment anything touched the binding — so the arm that ignores it passed, and str_len v segfaulted. Same shape as Q-069's LLVM half, made again inside a day.
test/parity/env_var.mere asks two questions whose answers are the same on every machine: a name nobody sets is None, and PATH is set and non-empty. Comparing PATH's value would compare machines. Both arms, because a lowering that always answered None would pass the first alone.
Plain Wasm diverges and is pinned: a browser or worker host has no environment, so mere -w answers None to everything. The component backend does have it. The divergence is the host's, not the backend's, and the pin retires itself if that stops being true.
dune test 2596/0, parity 137/0 with one declared divergence, host_matrix 0 MISSING / 0 nocompile / 0 error.
v0.1.336 — 2026-08-27
_A regression from v0.1.333, shipped in v0.1.335, that the gate built to catch exactly this could not see._
v0.1.333 made the C backend refuse a builtin it has no lowering for instead of emitting mu_<name> and letting a C compiler complain. Correct — and the guard that keeps a user's own binding of such a name on the ordinary path, user_shadows, answers for top-level functions, locals, captures and lifted inner fns. Not for a top-level VALUE binding. mere-ruby has
let int_max = 4611686018427387903 * 2 + 1;
let int_min = 0 - int_max - 1;
and stopped compiling. top_globals is asked about too now.
`downstream_check` said nothing, and that is the more useful half. It ran mere -t, and this failure is at mere -c — so the gate written three releases ago to catch upstream breakage in the dogfood watched an upstream breakage go past and into a release. It was found by building mere-ruby by hand.
The check is mere -c now: parse, resolve imports, type, and emit. All twelve repositories pass it, and 43 seconds covers all of them (41 of that is mere-ruby). Not -ll or -w as well: ten and seven of the twelve are legitimately refused by those backends for features they do not have, so requiring emission there would be requiring the wrong thing.
Poisoned by restoring the v0.1.333 behaviour: the mere-ruby row fails and names it.
v0.1.335 cannot build mere-ruby. It is the release cut an hour before this was found.
dune test 2596/0 (two new), parity 136/0, downstream_check 12/12.
v0.1.335 — 2026-08-27
_Around three dozen programs are written in this language, none of them has CI, and this repository could not see any of them._
An upstream language change breaks its own dogfood silently. The news arrives whenever somebody next opens that repository — which for most of them is not soon. v0.1.334 wired in the first one because the RISC-V differential needed it; this is the rest of the answer.
test/downstream/REPOS names twelve, one per language surface — TCP and actors (mkv), partial failure and time-as-messages (mraft), the bytes type (mpng), native CLI and package resolution (mq), subprocess (mk), interactive terminal (mrog), binary parsing (mwasm), web plus a typed model layer (mere-blog), the largest single program (mbrowse), the largest surface of the language (mere-ruby), the memory model as a library (mere-gcheap), and the emulator the RISC-V differential runs against (memu). A breakage lands on the row that names the feature.
The check is `mere -t <entry>` from inside the repo: parse, resolve every import, type the whole program. Deliberately not "the project works" — it does not run their tests, and each repo owns that question. It is the question this repo can answer, it is the failure a language upstream actually causes, and it costs about a second each.
Twelve check green today. A thirteenth row is recorded and not run: mbigfmt imports another repo through mere.toml, and fetching it is a network dependency this gate does not take on — so the row says that rather than being absent.
Poisoned by removing one builtin from the typer's environment: all twelve fail, each naming its surface. The mechanics were poisoned three more ways — a row pointing at a file that is gone, every repo missing (reported as 0 checked, 12 absent, not as success), and an unknown check value.
Said plainly in the file: this table cannot be derived, unlike the others here. The set of repositories lives on GitHub, not in this tree, so nothing notices a new one. Knowing which kind of table you are holding is the difference between trusting it and checking it.
dune test 2594/0, parity 136/0, downstream_check 12/12.
v0.1.334 — 2026-08-27
_The differential that gives the RISC-V arc its claim was never running in CI, and the gate reported the same line either way._
Three candidates were checked for being ungated. Two were already handled correctly, which is worth saying because the reflex after a run of findings is to expect more:
`bench_check` gates the deterministic half only — no wall clock, no peak RSS — and that is a stated decision with a reason (both measure the runner; RSS is quantized and stops reproducing above a few GB). All seven benchmarks carry an allocation band, so nothing deterministic is ungated. test/escape/ROUTES has four rows (not six) where a backend cannot read the program — all SAFE rows, all refused for documented codegen-subset reasons unrelated to escape, all recorded per backend and checked in both directions.
The third was real. scripts/qemu_virt.sh can diff each RISC-V image against the emulator written in Mere, and CI never set `MEMU` — so the headline claim of the whole arc, that the same bytes behave identically on QEMU and on a CPU we wrote, was true (it passes) and gated nowhere.
Worse than not running: it reported the same `3 passed, 0 failed` either way. A gate that checks half of what it can check, and says the same thing about both, reads as having checked all of it.
Now: the missing half is named on its own line, the summary says which halves ran, and CI checks out 284km/memu and points MEMU at it. If that checkout fails the run continues on the QEMU half and says so, the same way component_parity and socket_parity skip loudly rather than silently.
The summary's parenthetical is a claim, so it is only made when true — the first version of it said "agreed on every image" next to a nonzero failure count. Poisoned both ways: a wrong expected output fails as (qemu), and an emulator that disagrees fails as (memu disagrees with qemu) on all three images.
dune test 2594/0, parity 136/0, qemu_virt 3/3 on both emulators.
v0.1.333 — 2026-08-27
_Q-071: the six nocompile rows were two mechanisms, and both were the failure being handed to a C compiler._
nocompile is host_matrix's worst category: the backend emitted C, and a C compiler then rejected it — so the diagnostic reaches the programmer from a tool they did not invoke, about a name they did not write. Six rows had it. Classifying them before fixing anything (the point of classifying is to decide whether they can be bundled) split them two ways, not six and not one.
Four rows, one mechanism: `mu_<name>`. int_max, int_min, env_var and random_float have no C lowering at all, and the reference path fell through to c_safe_name — the mangling for a user binding. So the emitted C referenced mu_int_max, which nobody declared. LLVM and Wasm name their own gap; C now does too, via the same Typer.initial_env derivation as v0.1.332, guarded by user_shadows so a program's own let env_var = ... stays on the ordinary path. Q-064 recorded this shape one release before this one found the rest of it.
Two rows, a different one: a tag that named a function nobody emits. ty_tag erases an unconstrained type variable to "int", on the documented grounds that nothing ever inspects such a value — correct where a tag names a representation, and wrong here, where of_json_<tag> names a function whose definition add_type_and_deps declines to emit for a non-concrete type. The registrar and the emitter were applying different rules to the same type, so let _ = of_json "1"; 0 called of_json_int and nothing defined it. Both now use the registrar's rule, and the message of_json already carried for a missing target type is the right one for an undetermined one. (of_json s : int) is unaffected.
host_matrix: 0 MISSING, 0 nocompile, 0 error — every hole in that table is now a backend saying so itself, plus two bare-only cells that are an honest third answer. It has never read that way before.
dune test 2594/0 (six new, poisoned two ways: removing the backstop fails exactly the two mu_ tests, removing the concreteness check exactly the two of_json ones), parity 136/0.
v0.1.332 — 2026-08-27
_The matrix whose job is to say which backend has which builtin was missing a backend, and every hole it did show was a diagnostic blaming the user._
The RV32I backend has no flag-gated runtime sections — it emits instructions, not text blobs — so the Q-069 audit had nothing to find there. What it has instead is a column that was never in `docs/host-matrix.md`. A record whose whole purpose is "which backend has which builtin", with the fifth backend not in it.
RV32I now reads 48 yes, 100 refused, 2 bare-only, 0 MISSING.
Two things had to be true first.
RV said `unbound variable` for a name the language has. That is the failure shape host_matrix calls MISSING: a backend hole reported as a user typo. It names its own gap now — and the set is derived from Typer.initial_env rather than kept as a second list, because a second list drifts from the first.
`csr_read` / `csr_write` are neither. RV32I has them, on --bare only, where the program is handed the machine instead of a host. Calling that refused would be the matrix lying in the flattering direction and yes would be lying in the other, so the cell says bare.
Then the new column showed the same defect in two old ones. int_max, int_min, stdin_byte, bytebuf_new and the CSR pair reported unbound variable on LLVM and Wasm — six rows, all of them builtins added after those backends' hand-written "no lowering yet" lists were written. codegen_wasm.ml's own comment says this happened at v0.1.216. The lists stay, because they can say more (they know the scope), but the same Typer.initial_env derivation now stands behind them, so a builtin added tomorrow is named as a backend gap instead of as the user's mistake.
MISSING is 0 across every backend for the first time. The 6 nocompile rows — emitted C that a C compiler then rejects — are a different category and still open.
dune test 2588/0, parity 136/0, host_matrix 150 builtins / 0 MISSING / 6 nocompile / 2 bare-only / 0 error.
v0.1.331 — 2026-08-27
_The same audit on C and LLVM: eighteen more sections, one with no program, and it worked._
v0.1.329 enumerated the Wasm backend's hand-emitted runtime sections and measured which the corpus reaches. The C metrics defect in Q-068 had been found by accident — through a Wasm program — so the same measurement was owed to the other two backends.
C has eight gated sections, LLVM ten. Compiling all 135 parity programs with -c and -ll and looking for each section's marker left exactly one with no program: C's read_lines_helper. Probing it found read_lines correct on the interpreter and C, and refused by name on LLVM and Wasm (Phase 5.1 / 6.1 MVP) — the honest half of a partial implementation, which parity records as UNSUP. test/parity/read_lines.mere covers it, with a blank line and no trailing newline in the payload, since that is where a line splitter's off-by-one lives.
Running total: five sections across three backends had no program, and one of the five was broken. Writing the other four down matters as much — a file that recorded only the defects would make the ratio unreadable.
test/parity/SECTIONS now carries all three backends (29 rows), keyed c: / llvm: where a section name would otherwise collide — logger_used names a different piece of text in each. section_coverage compiles the corpus once per mode and each row's program in the mode the row claims, so a component-only section declared wasm fails on both counts. Extractor and table agree at 29 in both directions, which is the check that the gate is looking at all.
parity 136/0, dune test 2588/0, section_coverage 29 sections, component_parity 3 programs.
v0.1.330 — 2026-08-27
_The three sections that had no program now have one, and none of them was broken._
v0.1.329 left test/parity/SECTIONS with three rows marked UNREACHABLE: the parity harness compiles with plain mere -w, and env_var, read_stdin and the socket FFI helpers only exist under -w --component. Recording the reason was honest, but "cannot be reached by this harness" is a statement about the harness.
scripts/component_parity.sh builds each as a component and runs it against the interpreter: env_var with one variable set and one deliberately unset, so both arms of the option are exercised and the answer does not depend on the machine; read_stdin with a fixed payload — which is also why parity could not host it, since read_stdin blocks until EOF and a harness feeding nothing would hang rather than fail. The socket program opens nothing: declaring the externs is what turns the section on, so it is built, validated and run, while talking to a live peer stays socket_parity.sh's job. Those are two different questions and the file says so.
All three were correct. After Q-068 and Q-069 the expectation was more of the same, and it is worth writing down that it was not: an untested section is a risk, not a defect.
SECTIONS gains a mode column (plain / component / none) and section_coverage compiles each row's program the way the row says it is reached, so a component row mis-declared as plain fails on both counts. none survives as a status with no rows in it — the next genuinely unreachable section should be a row with a reason rather than an absence, and the gate still fails if one becomes reachable.
Zero unreachable sections. parity 135/0, dune test 2588/0, section_coverage 11 sections (3 component-only), component_parity 3 programs.
v0.1.329 — 2026-08-27
_Q-069: len on a list did not work on two of four backends, and the audit that found it is now a gate._
Q-068 ended with a general worry rather than a finding: runtime sections emitted as literal text are never type-checked, and one that no program reaches is never validated, linked or run either. So the sections were enumerated — eleven of them are gated behind a *_used flag in the Wasm backend — and each was checked against the corpus by compiling all 134 parity programs and looking for the section's marker in the output.
Four were reached by nothing. Three of those cannot be reached by this harness at all and are now recorded with the reason (two are --component-only; one needs a live peer). The fourth was $mere_list_len, and probing it found len on a T list broken on two backends, three separate ways:
Wasm read the tag as `i32.load offset=0` and the payload at offset 4 — the layout from before values widened. A variant's tag is `i64.load offset=0` and its payload `offset=8`, and a tuple field is `8 idx. * **LLVM** registered len-on-a-list in vec_to_list_instances, the table that *also* drives which vec_to_list helpers get emitted. One table, two questions: a program with a list and no vec emitted a vec_to_list helper dereferencing %mere_vec_<T>, a struct type emitted only for real vecs. It did not compile. * **LLVM again**, behind that one: the helper loaded the payload slot *as* the tuple. Phase 25.0 boxed the payload behind a pointer, and the code that does this correctly is written out at the list-fold site three thousand lines away. It segfaulted.
Interp and C were right the whole time, which is why nothing looked wrong.
scripts/section_coverage.sh and test/parity/SECTIONS keep it from happening again. The flag list is read from `lib/codegen_wasm.ml` at run time, so a new gated section with no row fails on the commit that adds it; COVERED rows must be reached by the program they name; and UNREACHABLE rows must still be unreachable, so a section that becomes testable is promoted rather than left looking covered.
parity 135/0, dune test 2588/0, section_coverage 11 sections / 125 programs compiled / 3 declared unreachable.
v0.1.328 — 2026-08-27
_Q-068: the capability runtime sections were hand-written for a compiler two width changes ago, and no program in the corpus reached them._
mk_logger / mk_metrics and the closures they hand out are emitted as literal WAT and C text rather than compiled from Mere. That text was written when a Mere value was an i32 on the Wasm backend and an int on the C one. Both have since changed, and the text did not.
Three defects, in increasing order of how well they hid.
- Wasm did not validate.
(local $tmp i32),(i32.const %d)for a string
offset and (i32.const 0) as an i64 result, against __lang_str_concat : (i64, i64) -> i64. A module that does not validate cannot run, so this one announced itself.
- Then it validated and printed the wrong thing. The Logger record still
used 4-byte fields, and a record field is read with i64.load offset=(8 * idx). So logger.warn returned the error closure and logger.error read past the end. The closure a field points at is two i32 slots (env, fn_idx) — two widths in one structure, and only one of them had moved.
- C compiled a function pointer of the wrong type.
__mere_metrics_record_inner_fn(void* env, int n) against a closure typed int (*)(void*, long long) — int here predates v0.1.41, which made Mere's int 64-bit on that backend.
Why it survived: the capability sections are reachable only from a program that uses a capability, and no Wasm-compiled program in the corpus did. examples/toy_sql.mere does — and it had not lexed since the interpolation rule changed, so it was examined by no backend at all. Fixing the lexer (v0.1.327) is what exposed this.
test/parity/logger_metrics_caps.mere is nine lines and touches the whole surface: all three logger levels, both metrics entry points, and therefore both widths. It is what found the C defect after the two Wasm ones were fixed. Poisoned by moving the warn field back to offset 4 — the module still validates and the output is wrong, which is the half that needed a runtime comparison rather than a type check.
examples/toy_sql.mere now agrees byte for byte on all four backends, which is what the README had been claiming since before it stopped lexing.
parity 134/0, dune test 2588/0, escape_check 16 routes / 0 holes, first_run_check 10 README commands.
v0.1.327 — 2026-08-27
_An interpolation error was reported in a file the compiler never read, an example nobody could run, and an install script offering an artifact nobody builds._
A `{...}` fragment inside a string is re-tokenised from scratch, so its tokens came back positioned at line 1, column 1 of the fragment — and those positions went straight into the outer report. let c = "x:{id=1,name='A'}" on line 3 announced its parse error at 1:3, which is a position in the file the reader is looking at and not the one the compiler read. A string literal cannot contain a newline, so the whole fragment sits on one line and the correction is a column shift; the inner tokens are retagged onto the enclosing string's position.
`examples/toy_sql.mere` did not lex, and the README advertised it. "{" opens interpolation, so show_row and ten expectation strings full of rows:[{id=1,name='Alice'}] were being read as code. It runs now: 59 tests, 0 failures, and the interpreter, C and LLVM agree byte for byte. Wasm miscompiles it — __lang_str_concat is called with [i64, i32] where the type says [i64, i64], so the module does not validate (Q-068, open). The README says which backends are claimed.
Finding it was its own lesson: the diagnostic pointed at line 1 of a comment that lexes fine on its own, so bisecting prefixes rather than reading the message is what located line 1053. And the first scan for sibling cases used {[a-z_]*=, which missed {users.uid= — a pattern narrow enough to look thorough. Extracting the string literals and asking each one whether it holds an unescaped brace found the last case.
`scripts/install.sh` offered `mere-macos-x86_64`, which the release matrix does not build. An Intel Mac got a 404 and a message asking whether a release had been published — releases exist; that platform is not in them. The script routes it to build-from-source now, and first_run_check holds the two files to each other in both directions.
`release.yml` never looked at what it published. Three checks now do: the tag against lib/version.ml, the built binary against the tag, and — the one that would have caught install.sh serving a binary 260 versions old — the uploaded binary, fetched back the way a user gets it, asked for its version and made to run a one-line program. gh release create ... || true also swallowed every failure, not just the expected "already exists".
dune test 2588/0, parity 133/0, escape_check 16 routes / 0 holes, first_run_check 10 README commands.
v0.1.326 — 2026-08-26
_Q-067: the region can be named in a type's declaration instead of in its application — and the table that found it._
mentions_region_in_value walked a TyCon's type arguments. type Box = { items: Vec[R, int] } and type hold = HVec of Vec[R, int] are nullary type constructors, so there were no arguments to walk: the region is named in the declaration, which lives in the record and constructor registries. An outer binding therefore read Vec[__heap, Box] with R nowhere in it and the store escaped.
Both dangled. Burning the arena with 4,000 strings and reading the container back gave 2,054,847,098 on the C backend against 3 under the interpreter — the same garbage value and the same two-backends-disagreeing shape as Q-053.
The check now consults both registries when it meets a nominal type. seen is not an optimisation: type t = T of t names itself in its own declaration, and without it the walk does not answer wrongly, it does not answer.
v0.1.323 recorded that recursive variants were already safe. That was true, and it was about a variant of `str`, which is deep-copied on the way out. A variant of a container is a different claim. test/escape/variant_of_str.mere and test/escape/variant_payload.mere are the two programs that separate them, and the first is in the suite so that fixing the second cannot swallow it.
How it was found, which is the part worth keeping
Not by hitting it. test/escape/ROUTES enumerates every way a region-bound value can reach something that outlives its block, and scripts/escape_check.sh runs it. A table has to have every row filled in; a bug report is complete at one. Filling it in produced these two holes and one entry-point asymmetry (the store check does not fire under the interpreter or mere -t, because those paths generalise each declaration while the compiling path desugars top-level lets into nested Lets — pinned in the table's interp column rather than left to be found again).
The gate is also what said the holes had closed: HOLE rows fail on being rejected, so the table cannot silently agree with whatever the compiler currently does.
16 routes, 0 open holes. Over-plug guards ran green throughout: dune test 2587/0, parity 133/0, selfhost all passed.
v0.1.325 — 2026-08-25
_Q-066: the call site looked the callee up under a name it is never emitted as._
A saturated call to a known N-ary function goes straight to its uncurried __direct twin instead of building a closure to apply the second argument. That never happened for a multi-instantiated function, and the reason turns out to be a single lookup: direct_fns was consulted under the source name, which is never a key for such a function. Its specialisations are in there — they are ordinary fn_decls with concrete parameter types, keyed by mangled_inst_name — and their __direct twins have always been emitted. Nothing could name them.
So every recursive call in contrib/json's rev_aux built a 32-byte closure environment: 800,000 of them, 25.6 MB, 18% of everything the JSON benchmark allocated.
The lookup uses the emitted name now — the instance's mangled name when the call's head type has resolved to a concrete arrow, the source name otherwise. A multi-inst call whose head type has not resolved still takes the closure path, which is the same fallback as before, narrowed to the case that needs it.
v0.1.324's library workaround is gone. rev_aux is one polymorphic function again, and the numbers are byte-identical to the two hand-specialised copies:
json alloc 112,911,240 B (same as the workaround, to the byte) json RSS 109.2 MiB json wall 105 ms
Which is the point: the workaround was faking what the compiler should do, and its comment said to collapse it back when the compiler learned to. It has.
Cumulatively, one 11.8 MB document went 169.2 → 159.0 → 139.8 → 112.9 MB across four releases, every step found by attributing allocations to return addresses.
The test asserts the CALL, not the definition, and that distinction is the whole lesson of v0.1.324. The first attempt at this fix computed the mangled name for the callee and left the lookup alone. It built, every gate stayed green, and the allocation moved by zero bytes — the guard could not fire, and nothing in the suite could tell. test/test_basic.ml now checks that the twin is reached: both instances emit one, and neither body still calls the one-argument curried entry. Poisoned by restoring the source-name lookup, which fails it and nothing else.
Nothing else in the benchmark suite moves — binarytrees, churn, crc32, matmul, startup and wordfreq have no multi-instantiated function on a hot path, and their allocation is unchanged to the byte.
v0.1.324 — 2026-08-25
_Half of a list reversal's cost was closure environments._
The JSON benchmark's biggest remaining allocation site was contrib/json's rev_aux, at 32% — and it is not the algorithm it looks like. Reversing an accumulator costs N cons cells and always did. What attributing the allocations to their return addresses showed is that half of that 32% was closure ENVIRONMENTS: 800,000 of them, 32 bytes each, 25.6 MB, one per list element.
Why: a call site can call a known N-ary function directly, without building a closure to apply the second argument — but only if the function is in direct_fns, and that table admits a function only when its parameter and return types are concrete. A polymorphic rev_aux : 'a list -> 'a list -> 'a list is not. So every recursive call built a closure. The uncurried __direct twin was being emitted for each instantiation the whole time and nothing could name it.
rev_aux is two monomorphic copies now, one per element type, which is what the compiler can exploit today:
json alloc 139.8 MB -> 112.9 MB (-26.9 MB) json RSS 134.9 MiB -> 109.2 MiB json wall 124 ms -> 107 ms
Cumulatively over four releases, the same 11.8 MB document went from 169.2 MB of allocation to 112.9 — a third of it gone, and every step found by attribution rather than by reasoning about where the bytes ought to be.
A compiler fix was tried first and reverted, which is the more useful half of this entry. The obvious move was to widen collect_direct so a multi-instantiated function resolves to its own instance's __direct twin — the call site already computes the mangled instance name for the closure entry, so the name was available. It built, every gate stayed green, and it changed nothing: the allocation was byte-identical. The reason is that the two conditions are mutually exclusive. direct_fns requires concrete parameter types; a function is multi-instantiated precisely because its parameters are NOT concrete. The widened guard could never fire — dead code that looked like an optimisation. Reverted.
The real fix is to register instantiations in direct_fns under their mangled names, which is a change inside the monomorphisation path rather than at the call site, and belongs in its own arc. Until then the library-level workaround carries a comment saying what it is working around and when it should collapse back into one function.
v0.1.323 — 2026-08-25
_Q-053: the escape check only ever looked at the way out it could see._
region R { } rejects a container returned as the block's RESULT. It said nothing about one STORED into a binding that outlives the block, because the check read the result type and these programs return nothing:
let m = map_new () in
let _ = region R { let v = vec_new () in ... map_set m "k" v } in
print (str_of_int (vec_len (map_get m "k")))
That type-checked. Containers are handles and copy-on-store copies the handle — by design; mkv's actor-shared Map depends on it — so m kept a pointer into the block's arena. With the arena burned over afterwards, the C backend read a length of 2,054,847,098 and then segfaulted, while the interpreter, which has no arenas, printed the right answer. Two backends disagreeing, with nothing watching. region R loop had the same hole, one iteration sooner.
The routes were enumerated rather than guessed, and the hole is exactly containers: Vec into a Map, Vec into a Vec, Map into a Map all dangled. str, recursive variants like list, and closures were already safe — all three are deep-copied on the way out, closures including their captured environment (v0.1.290).
The check now also scans the bindings that existed BEFORE the block. The unification that puts R into an outer binding's type happens inside the block, so by the time the body is typed the damage is already visible — there was no ordering problem to work around, which the Q-053 note had listed as the risk. The message names the binding and the type it became.
It walks everything except the inside of a function type, and region_loop_carry is what taught that: it binds, outside the loop, churn : Map[RL, str, str] -> int -> Map[RL, str, str] — a function that operates on the carry, which is how the construct is meant to be used — and the first version of this check stopped it type-checking. A function holds nothing; a closure's captured environment is deep-copied with it. A binding holds a region's memory when the name reaches it through a container, a tuple or a variant, and those are still walked.
Six cases in the suite: three that must be rejected, and the three that must NOT be — the function over the carry, a closure leaving a region, and a deep-copied value crossing outward. Each of the latter was rejected at some point while this was being written.
_And two things that came out of comparing against MoonBit._
`str_unescape` looked at nothing before allocating. A string with no backslash in it unescapes to itself, and it was copying it anyway: 1.12 million identical copies, 19.2 MB, 12% of everything the JSON benchmark allocated. Found by attributing allocations to their return addresses rather than guessing — the same method the Ruby workload's 8.6 GB was taken apart with. It returns s now, and so does str_trim when there is nothing to trim. str_replace already returned s for an empty needle, so this is an established shape rather than a new one.
json alloc 159.0 MB -> 139.8 MB (-19.2 MB, matching the attribution to the byte) json RSS 153.2 MiB -> 134.9 MiB
`docs/patterns.md` §8.4: a function whose body is a region. Not a new feature — region R { } is an expression and could always be a function's whole body — but nothing said so, and it is where the construct usually wants to be. It says plainly what had been left implicit: Mere does not reclaim by default, so a program that allocates in a loop and never writes region holds everything it ever allocated. The section carries the measurement (32 bytes in the default region with it, 192,032 without), what cannot cross the boundary, and why the function boundary in particular — the cost of a region is entirely what has to cross it, and a function draws that line where the crossing set is already written down.
v0.1.322 — 2026-08-25
_Half of what a tree-shaped program allocated was empty nodes._
A boxed variant's NULLARY constructors carry no payload. Every Nil, every Leaf of a recursive type is indistinguishable from every other one — and each was a fresh region allocation. On the binarytrees benchmark that was 50.3% of everything the program allocated: 7.4 million empty nodes, 170 MiB, for values that differ in nothing.
They share one immutable static per (type, tag) now.
binarytrees 191 ms 338 MiB -> 107 ms 169 MiB with region 87 ms 53.6 MiB -> 68 ms 28.1 MiB json 169.2 MB alloc -> 159.0 MB (-6.1%) wordfreq 27.4 MB alloc -> 26.7 MB (-2.3%)
The naive row is now faster than hand-written C on that workload (183 ms) and level with MoonBit (102 ms), where before it was C's speed on a hundred times C's memory. Nothing about codegen changed: cutting the memory in half cut the time in half, because a 338 MiB working set does not fit anywhere near the CPU.
Where this came from. The MoonBit row added in the previous commit beat naive Mere on both axes, and reading its generated C is what named the reason — Leaf compiles to &moonbit_constant_constructor_0, a static, where Mere compiled it to a bump allocation. The other reason it wins there is a reference counter, and that one Mere should not copy: region R { } already answers the same question and answers it faster (68 ms against 102 ms). What Mere was missing was not a way to reclaim, it was a decent DEFAULT — and this is the half of that gap that costs nothing to close.
The symbol is keyed by tag number rather than constructor name, because the expression emitter and the struct-body emitter reach the name by different routes (raw, canonical, substituted) and the tag is the one thing both already agree on.
What was checked before touching it, since this hands out a pointer that lives in no arena:
Nothing mutates a variant node after construction — `__mcopy` and `__mdeep` write to the fresh copy they just made, and `vec_to_list` writes to the node it just built. Equality on a boxed variant is structural (eq_t compares tags, then payloads), not pointer identity, so two separately written Nils were equal before and are equal now for the same reason. The copiers were deliberately left ALONE. Making them return a shared node unchanged would have been faster still, but it would make correctness depend on "every nullary node is the shared static" — an invariant the `list_str` builtin helpers quietly break, one node per call. They copy it instead, which is correct without needing the invariant to hold anywhere.
test/parity/shared_nullary_regions.mere is the case written for the failure this could introduce and nothing else would have caught: an empty list, a non-empty list, and a user variant each leave a region R { }, the arena is then burned over with 2,000 fresh strings, and all three are read again. Plus the region-loop carry, and the two equality claims above. It does not detect whether the sharing is happening — reverting to allocation passes it, and should, because both are semantically identical. What detects that is the deterministic allocation bound in benchmarks/binarytrees/MANIFEST, whose FLOOR caught this change as a breach: a number that halves is either an improvement or a meter that stopped looking, and from the outside those are the same event.
v0.1.321 — 2026-08-24
_Q-063 closed: the last two backends that shifted._
v0.1.317 tombstoned map_delete on the C backend and v0.1.320 on the interpreter. LLVM and Wasm still found the entry through the index and then shifted every dense entry above it down one slot and rebuilt the index, so one delete cost O(live) on both. Measured before being written about — 40k operations, live set 500 → 4,000, against the C backend's flat 0.02 s:
LLVM 0.19 0.45 1.16 s Wasm 0.34 0.44 0.68 1.12 s (node's ~0.18 s startup included)
Both carry the same design as the C backend now: dead[] marks a vacated dense slot, live is what _len answers, idx_used counts occupied index slots (live PLUS vacated, because a vacated slot still lengthens every probe through it), the index slot becomes -2 rather than -1 so a probe walks past it to a key inserted after the deletion, set reuses a vacated slot before consuming a fresh one, and the dense arrays are squeezed in place — keeping insertion order — only when the dead outnumber the live.
LLVM 0.02 0.01 0.01 0.01 0.01 s (live 500 → 8,000) Wasm 0.20 0.19 0.18 0.19 0.20 s (node startup, and little else)
Flat on both. At the largest live set the suite measured, LLVM goes 1.16 s → 0.01 s.
Two things that were structural rather than transcribed. Where the C backend re-probes inline after a squeeze, both of these re-enter `set`: the squeeze rebuilds the index, so every register or local the probe derived is stale, and a self-call is the honest way to say so — the key is still absent and there is room now, so it is one level deep. And the LLVM iteration helper is shared with the LINEAR fallback runtime, whose struct is five fields wide; reading a dead pointer out of it would be a read past the end. The helper takes the same map_key_index_safe predicate that picks the runtime, so the two cannot drift apart.
Wasm's map_clear gained work rather than losing it: it now resets the tombstone counters and rebuilds the index, because leaving vacated slots behind an emptied map would make the next probe walk them.
RV32I is not in this list and is not an omission. Its map is an association list in the bare-metal prelude — _mget, _mset and _mdel all walk it — so delete being O(n) there is not a distinguishable defect but the shape of a runtime chosen for code size on a machine with no hash index at all. Tombstoning it would not change any complexity class.
Semantics are unchanged, which is the whole claim, and test/parity/map_delete_tombstone.mere is what holds it: all four backends run it and produce one transcript. Poisoned on both new backends. Vacating with -1 instead of -2 on LLVM fails that case and nothing else — 131 of 132 stay green — which is the third time that specific poison has been visible only to the test written alongside the design. map_len returning the high-water mark on Wasm fails five.
v0.1.320 — 2026-08-24
_The interpreter had the same O(live) delete, and it was hiding behind its own overhead._
v0.1.317 tombstoned map_delete on the C backend. The interpreter's version was the same bug in a different shape: Hashtbl.remove (O(1)) followed by List.filter over the insertion-order key list (O(live)) to keep that list exact.
It was diluted rather than dominant, which is why it survived — at 20k operations with the live set grown from 500 to 8,000, the run went 0.38 s to 1.92 s, where O(1) is flat and pure O(live) would be 16x. Enough of the cost was interpreter overhead that the shape did not stand out until the C backend's had been measured and named.
Same discipline as the C fix. The key list is append-only now: a delete touches only the hash table. So the list accumulates keys that are gone, and keys that appear twice (deleted, then re-inserted) — and every reader (map_iter, to_string, to_json_string) walks it newest-first keeping the FIRST occurrence of each key still in the table, which puts a re-inserted key at its new position, matching every compiled backend. map_maybe_compact rebuilds the list when the stale entries outnumber the live ones, so it stays proportional to the live set rather than to the number of writes, amortised O(1) per delete.
V_map carries a small record now (m_tbl / m_order / m_order_n) instead of a Hashtbl and a list ref: the length has to be a field, because a trigger that called List.length would put back the O(live) it was removing.
Readers get a fast path for the case that has no tombstones at all — when the list length equals the table's, it is already exactly the live keys, so iteration does not build a dedup table for a facility only churn needs.
Measured, 20k operations, live 500 → 8,000: 0.38/0.58/0.82/1.29/1.92 s becomes 0.23/0.23/0.23/0.22/0.19 s. Flat, and 10x faster at the top.
Two backends still shift, and they were measured before this was written rather than assumed. 40k operations, live 500 → 4,000, against the C backend's flat 0.02 s: LLVM 0.19 / 0.45 / 1.16 s and Wasm (minus node's ~0.25 s startup) 0.09 / 0.19 / 0.43 / 0.87 s — both doubling with the table. Their map runtimes are hand-written IR and WAT, so this is a larger change than the two already made and is not folded in here. Semantics are unaffected either way: test/parity/map_delete_tombstone.mere runs on all four backends and the ones that still shift produce the same transcript as the ones that do not.
v0.1.319 — 2026-08-24
_The deterministic memory meter could not see the construct that manages memory._
MERE_REGION_STATS reported the DEFAULT region only. A program that did its allocating inside region R { } or region R loop — which is to say a program managing its memory deliberately — read as a few hundred bytes. The number that exists precisely because peak RSS is quantized and stops reproducing above a few GB was blind to exactly the programs where memory is the point, and peak RSS was the only figure left for them.
The benchmark suite is where this became untenable rather than merely known: its churn row ran a region loop that moved 62 MB through 42 arenas and the deterministic column printed 0 KiB.
Named arenas are created and destroyed during the run, so there is nothing to walk at exit. Each is charged to its source name as it is RELEASED, into a 32-entry table keyed by that name. Pooled block arenas are recycled across sites, so the name is set at acquire and the running total is reset there too. The output gains one line per name and a total:
region-stats default: blocks=1 cap=4194304 alloc_total=184
region-stats named region R loop: arenas=42 alloc_total=62376872 peak_cap=3145728
region-stats named-total: sites=1 alloc_total=62376872
The default: line keeps its exact spelling — scripts/region_slack_check.sh greps for it.
What the number immediately corrected. The region-loop variant of the churn benchmark was being read as the cheaper program. It is not: it allocates 62.4 MB against the naive version's 43.4 MB, because it copies the live set into a fresh arena once per generation. What it does is hold 3.1 MB at peak instead of 60 MB of capacity. The construct does not make reclamation free; it bounds residency and charges a copy for it. Both halves of that are now measurable, and benchmarks/run.py prints both — with the churn MANIFEST gaining a peak_max bound, since peak arena capacity is a function of the program in a way peak RSS is not.
The benchmark gate also stopped bounding the wrong thing: it held the default region's allocation, which a program that moves its work into a named arena steps out of by construction. It bounds every arena now.
Two limits, stated rather than left to be discovered. An arena still live at exit is never released and so never charged — in practice a map's private arena between compactions. And past 32 distinct names, releases are counted and reported as a region-stats WARNING line rather than dropped, because an undercount looks exactly like an improvement.
Three unit tests asserted on __lang_region_block_acquire(), which now takes the site name. One of them was invisible: it dumps the whole generated C on failure, the dump carries the NUL bytes of the emitted string table, and grep declared the log binary and reported zero matches while the summary line said one test had failed. Finding it took instrumenting the harness to print names to stderr. A harness that prints its entire subject on failure can make the failure unreadable — which is worth more than the three-character fix.
v0.1.318 — 2026-08-24
_Five of the eight polymorphic helpers only ever worked under the interpreter._
id, pair, swap, const and flip were builtins with polymorphic schemes in the typer and implementations in the interpreter, and nothing in any code generator. So they worked under mere file.mere and nowhere else, and had since they were added.
What made it worth a release rather than a footnote is the shape of the failure, which was different on each backend and worst on the one people compile with. LLVM and Wasm refuse them by name — unbound variable: pair — which is a diagnosis. The C backend emitted a reference to an undeclared mu_pair and left the diagnosis to the C compiler, which reported it in terms of a symbol the author never wrote, at a line in generated code they had never seen. An unimplemented thing should fail in a way that says it is unimplemented.
They are prelude definitions now (prelude_stdlib.ml), which is to say ordinary polymorphic Mere code:
let id = fn x -> x;
let pair = fn a -> fn b -> (a, b);
let swap = fn t -> (snd t, fst t);
let const = fn a -> fn b -> a;
let flip = fn f -> fn a -> fn b -> f b a;
Every backend compiles that. Partial application and use in value position work because these are closures rather than a special case in an emitter, and there is one implementation instead of one per backend to keep in step. The same migration pow, lcm, divmod and assert already made. All eight helpers now run on all four backends; three did before.
const b is still evaluated: Mere is call-by-value, so the argument's effects happen before the call, exactly as the builtin's did.
The names were already taken — this changes where they are DEFINED, not whether they exist — but id and const are common enough that a program defining its own is not hypothetical, so shadowing is pinned in test/parity/prelude_shadowing.mere rather than assumed. Three unit tests found the one real consequence: a user's id used to mangle as id__int and now mangles as id_uq1__int__int, because the prelude binding has to be uniquified around. Those tests are about specialisation, so they were renamed off the prelude's name; the shadowing case is the one that is about shadowing.
test/parity/poly_helpers.mere covers the eight on every backend. It is written to stay OUT of parity's "unchecked on wasm" list, which took two rewrites: Wasm refuses a polymorphic function that is used at several types and passed as a value ("no single value form on Wasm yet"), so id is applied directly there and pair is kept at one type. A case that a backend refuses at emit time is a case that is not protecting that backend.
Two things found alongside and NOT fixed here, recorded rather than folded in: fst / snd in value position (let f = fst) are still applied-form special cases in the C emitter and are refused by all three compiled backends; and partial application of a polymorphic two-argument function emits uncompilable C — let g = pair "k" produces a tuple_int_int initialised with a const char*. The second is not about these helpers at all: an equivalent user-written fn a -> fn b -> (a, b) does the same thing, so it is the monomorphiser defaulting type variables, and it is the same emit-bad-C-instead-of-refusing shape this release is about.
The builtin counts in docs/stdlib-reference.md were wrong before this touched them (the header read 202 against an environment of 226; the alphabetical heading read 129 against a list of 127). They now name scripts/host_matrix.sh as the thing that counts, because it enumerates the environment instead of being maintained by hand.
v0.1.317 — 2026-08-24
_Deleting one key cost the whole table._
map_delete found its entry through the hash index and then SHIFTED every dense entry above it down one slot and rebuilt the index. One delete was O(live), so a table under churn — a bounded live set with a lot of traffic through it, which is what a session table, a connection map or a bookkeeping table is — was quadratic in the number of operations.
The new cross-language benchmark suite is what named it, and named it in the form that made it unarguable. Its churn sweep holds the operation count fixed and doubles the live set: an implementation whose delete is O(1) draws a flat line. Rust, Go, Node, Ruby and Python all drew flat lines. This backend drew 130 / 208 / 392 / 817 / 1659 ms, doubling with the table, and came out behind Python on a workload it should not lose. Isolating it — same operation count, same key, only the resident table size changing — showed map_set flat and map_delete alone linear.
Delete tombstones now. dead[i] marks a vacated dense slot, live is what map_len answers, and the index slot becomes -2 rather than -1: a vacated slot must not end a probe chain, because a key inserted after the deletion may sit further along it. Nothing moves. The dense arrays are squeezed — in place, in insertion order, allocating nothing — only when the dead outnumber the live, so the O(len) squeeze is paid once per O(live) deletes and each delete is amortised O(1). set reuses a vacated index slot before consuming a fresh one, and the load factor counts occupied slots rather than live entries, so tombstones cannot silently lengthen every probe.
The sweep is now flat: 40 / 42 / 41 / 44 / 42 ms. The benchmark's main row goes 1687 ms → 74 ms (23x), and at the largest live set 1659 ms → 42 ms (39x). Mere now sits between Go and Rust there instead of behind Python. Cumulative allocation moves 43,298,448 → 43,404,960 bytes (+0.25%, the tombstone arrays) and peak RSS is unchanged.
Insertion order is why this is a tombstone and not a swap-with-last. map_iter order is pinned to insertion order across all four backends (Phase 27.1), which rules out the cheaper fix. Everything here is therefore supposed to be invisible, and test/parity/map_delete_tombstone.mere is what holds that: it crosses both internal thresholds deliberately (the squeeze needs the dead to outnumber the live AND to number at least eight; the index doubles past a 0.7 load factor), and all four backends run it — the three that still shift have to produce the same transcript as the one that no longer does. Poisoned both ways to check it is load-bearing: map_len returning the high-water mark fails it and three others, and vacating with -1 instead of -2 fails only this case, which is the subtle failure this design has and nothing else in the suite could see.
The C backend only. The interpreter has the same shape — map_delete filters the insertion-order list, O(live) per call — and it is measurably there (20k operations, live set 500 → 8,000: 0.38 s → 1.92 s, where O(1) would be flat and O(live) would be 16x), but diluted by interpreter overhead rather than dominant. LLVM, Wasm and RV32I still shift. Those are the remaining half of the question and are recorded as such rather than quietly counted as done.
v0.1.316 — 2026-08-23
_Every delete rebuilt the index into fresh memory._
map_delete shifts the dense arrays down and reindexes — at the SAME capacity, through a reindex that always allocated a fresh index buffer in the map's region. So a workload that deletes once per event (a bookkeeping table that records an entry per call and removes it on return) bled an index-sized allocation per delete, forever. Attribution of a large Ruby run's 8.6 GB default-region total put this single site at 3.15 GB — the biggest emitter, ahead of copy-on-store. v0.1.298's map_clear fixed the MASS deletion case and left the one-at-a-time case as it was.
reindex now reuses the buffer when the capacity is not changing (growth still allocates, as it must). The C backend and the Wasm backend — whose bump memory never shrinks, so it bled linear memory the same way — both carry the fix; the LLVM backend's delete does not reindex per call and needed nothing. On the probe (2,000 live entries, 50k set+delete churn) the default region's cumulative allocation drops from 820 MB to 1.0 MB.
The reuse has one caller it must not serve, and the parity suite caught it hanging: map_compact reindexes at a possibly-unchanged capacity while MOVING THE MAP TO A FRESH ARENA, so reuse would keep the index in the arena freed two lines later — dangling slots that probe forever. Compact zeroes idx/idx_cap before its reindex, which forces the fresh allocation it was always paying for. (Wasm's compact lowers to clear and never frees, so only the C backend had the interaction.)
The probe is test/regionstats/delete_churn_probe.mere, held by scripts/region_slack_check.sh via MERE_REGION_STATS: alloc_total must stay under 16 MiB, and above the probe's known floor in the other direction so a broken meter cannot pass as a fixed allocator. The attribution harness that found the site is bench/def_sites.sh in the mere-ruby repo (two-level return-address attribution, ICF disabled so the linker cannot fold the names).
parity 129/0; dune test 2576/0; host_matrix ok; region_slack both probes PASS.
v0.1.315 — 2026-08-23
_The optimizer was part of the language._
A dense dot product — the inner loop of an MLP dogfood — produced different BITS depending on how the emitted C was compiled: -O2 disagreed with the interpreter on 4 of 10 outputs by an ulp, and -O0 disagreed with both. clang contracts a*b+c into fused multiply-add by default on arm64, which rounds once where the interpreter (and the LLVM and Wasm backends) round twice. So a Mere program's float answers depended on a flag the language never mentions, passed to a compiler the user chooses.
Nothing in the repo could see it. Every gate that compiles C does so at -O0; the float parity that exists compares formatted output or quantized image bytes, both of which absorb an ulp. The blind spot was structural: the divergence lives exactly where no gate looked — unformatted bits, under optimization.
The emitted C now carries #pragma STDC FP_CONTRACT OFF: float evaluation is strict per-operation IEEE, at every optimization level, matching the other three backends. (gcc parses the pragma and ignores it — there, -ffp-contract=off is the user's knob; clang, which this project builds and tests with, honors it.) The new gate compiles one dense dot product at -O0 and -O2 and requires bitwise-identical output from both AND from the interpreter; reverting the pragma fails it with the two answers in the output. The probe's LCG is masked to 48 bits on purpose — the interpreter's int is still 63-bit OCaml native against the compiled backends' 64 (Q-037's open second boundary), and low bits survive both wraps, so the gate stays about floats rather than inheriting that older hole.
v0.1.314 — 2026-08-23
_Sound, but never a callback._
Native audio output did not exist — the browser backend could beep through a DOM binding, and that was all of it. The audio_* externs add it on the socket family's terms: SDL2 audio (the build line win_ already established), fixed s16le mono, samples crossing as arena bytes like every other device's data.
The design choice is the push model. Audio APIs usually hand you a callback that a device thread calls on a hard deadline — which for Mere would mean a host-owned thread calling back into the language, a whole boundary design of its own (the --lib arc met it from the other side). SDL_QueueAudio inverts the direction: the device drains a queue, audio_queued says how much is left, and the PROGRAM's job is to keep the queue fed before it empties. The hard deadline is still real — missing it is an audible gap — but it is now a deadline the program meets with ordinary code (a low-water loop of synthesize-and-queue), not a foreign thread inside the runtime.
The gate runs headless under SDL's dummy driver and holds the queue to its contract: pending bytes are visible, the queue drains to exactly zero in bounded time, a closed device answers -1, a nonsense rate answers -1. What the samples should BE is deliberately not this gate's claim — the dummy driver eats them, and SDL's disk driver was measured rewriting them (stereo upmix, chunk padding) — so byte-correctness belongs to the consumer, which checks its rendered PCM against an independently computed expectation.
v0.1.313 — 2026-08-23
_Every server so far either blocked or spawned._
mkv serves one connection at a time from a single loop; mhttpd spawns a thread per connection. Between those two shapes sits the one most production servers actually use — one thread, many connections, a readiness loop — and Mere had no way to write it: no poll, no nonblocking mode, nothing that answers "which of these fds is ready?".
The io_poll_* externs answer it, in the socket family's own conventions (ints and coded negatives, native C backend). A pollset is a registered interest set — registration persists across waits, the epoll/kqueue model, so the caller's work per wait is O(changes) and a future kqueue/epoll implementation changes no API. The v1 implementation is poll(2), portable everywhere the socket family already runs. One ready event packs into one int (fd8 + read/write/err bits), the convention midi_read established. `io_set_nonblocking` puts any fd of this runtime's into nonblocking mode.
Nonblocking mode is only usable because the failure codes already existed — almost. tcp_read has distinguished "try again" (-1) from "the peer is gone" (-2) since v0.1.226; tcp_write and tcp_accept still returned a raw -1 for everything, and a nonblocking writer against a full socket buffer needs the difference. Both now use read's codes. A polite client never fills the buffer, which is why write's went uncoded this long: the gate's probe fills it on purpose (32 KB writes against a reader that never reads) and pins the -1. The transcript checks both directions — an empty wait invents no events, a deleted fd's events stop — because a readiness API that fabricates readiness passes every positive-only test.
v0.1.312 — 2026-08-23
_It worked in every polite test._
mere_lib_shutdown freed the default region; mere_lib_init was a pthread_once, and once cannot be re-armed. So a host that shut the library down and then called an export again reused freed memory — and got the right answer anyway, in every test, because nothing had reclaimed the pages yet. A use-after-free that works is strictly worse than one that crashes: nothing names it. It took the second host (a deliberately abusive ctypes harness — the first, polite host could never see it) plus AddressSanitizer to make it loud.
The once is now a mutex and a flag. Shutdown is idempotent; a call after shutdown re-initializes — module init runs again, module state starts fresh, which is a meaning "call after shutdown" can keep rather than an accident it must avoid. Shutting down while another thread is mid-call remains the host's race to avoid, and the header says so. The gate grew a lifecycle section built with ASan where the compiler has it: early shutdown, a call that must re-initialize and answer correctly, a double shutdown — so "works by luck" fails loudly from now on.
v0.1.311 — 2026-08-23
_The process used to die before the leak mattered._
Each --lib wrapper call now runs inside its own region — the same acquire/release machinery as region R { } blocks, acquired after setjmp records the active-stack depth so a failure's unwind releases it too. Values were already handled: strings, cons cells and variant nodes follow the thread's current region, and the marshalling copies results out to malloc before the release. The leak was the containers. A Vec or Map whose region variable erased to __heap was pinned to the default region (v0.1.31: a container carries identity and may outlive any block — and the process dies with the program anyway, so nothing ever reclaimed it). A library breaks that last assumption: the process outlives every call. Measured, the pin cost 512 bytes per call for a small Vec and 1.5 KB for a small Map — 226 MB of default region across 100k calls of three small exports, in a program whose visible footprint should have been constant.
The fix follows from what the boundary already promises: a call is a transaction. Results are copied out; containers cannot cross. So in lib mode __heap containers follow the CURRENT region — the per-call region during an export call (reclaimed at return), the default region during module init (module state persists). Copy-on-store already covers the seam: a call-local value stored into module state is copied into the store's region, and the gate proves it by overwriting the host's buffer and reading the value back 50,000 calls later. After the change the same 300k-call run reads alloc_total=0 against the default region and a 1.5 MB resident set — was 264 MB of capacity.
The gate is two-sided and grew a thread section: module init builds a map so a zero reading would name a broken meter, two runs at different call counts must report byte-identical default-region stats, and 8 host threads hammer the same three exports concurrently (per-thread call regions, per-thread fail jmpbufs — v0.1.310's other half) with every value checked. TSan-clean; reverting the container redirect makes the gate fail with the growth numbers in its output.
v0.1.310 — 2026-08-23
_A library must not decide the host process is over._
An uncaught fail inside a --lib wrapper called exit(1) — correct for a standalone program, and exactly the thing a loaded library may never do to its host. Every wrapper now arms the same setjmp guard try_or uses (the same save/restore, the same active-region unwind from v0.1.301), and a failure comes back across the boundary as MERE_FAIL plus the failure's own words, malloc'd into the caller's err buffer and released with mere_lib_free. The library stays usable after its own failure — the gate calls the failing export again and expects the right answer.
Getting the words across took a slot: the fail_ helpers build their messages in stack buffers, and longjmp abandons the frame the message lives in, so `__lang_fail_impl` copies it into a per-thread buffer before jumping.
Per-thread is the other half of this change, and it is not lib-only. The jmp_buf and its armed flag were single process-wide globals, so a fail in a spawned thread could longjmp into a jmp_buf that some OTHER thread's try_or had armed — a jump into a foreign stack. try_or's save/restore protects against re-entry on one thread; only _Thread_local protects against a neighbour. Both are thread-local now, in every mode: a thread with no try_or of its own takes the uncaught path (print + exit 1), which is what the interpreter does.
v0.1.309 — 2026-08-23
_The boundary spoke three words of six._
v0.1.308's library boundary could pass int, bool and float; str and bytes — the types most functions worth exporting actually take — were skipped with a manifest apology. They cross now, as a mere_buf (pointer, length) pair in both directions, and the two directions deliberately do not mirror each other. IN, the host's buffer is borrowed: the wrapper copies it into the region and never looks at it again after the call returns. OUT, the host is handed a malloc'd copy it owns and releases with mere_lib_free — a region pointer never crosses the boundary in either direction, because the region's whole point is that a later call may reclaim it. A str copy carries one courtesy NUL past .len for C callers; bytes are length-only, and the gate round-trips a buffer with embedded NULs to hold that distinction (a boundary that secretly spoke C strings would truncate it). A bytes signature also forces the bytes runtime into the output even when the body never touches a bytes builtin — a passthrough only names the type.
mere --header <file.mere> prints the boundary's C header: the ABI types (include-guarded so two Mere libraries can share a host file), the lifecycle, and one prototype per export. The header is not derived a second way — it is read off the same wrapper signatures the emission just built, so the two cannot drift; the gate's host now compiles against the generated header instead of hand-written declarations, which is the header's actual test.
v0.1.308 — 2026-08-23
_A Mere program could only ever be the whole program._
Every backend emitted exactly one shape: a standalone executable with its own main. There was no way to hand a Mere function to an existing system — no library output, so "write one function in Mere" was not a thing a host process could do. mere -c --lib <file.mere> now emits a shared-library translation unit instead: no main, and a small C ABI boundary in its place — mere_lib_init / mere_lib_shutdown / mere_lib_free, plus one uncurried wrapper mere_<stem>_<fn>(args..., T* out, mere_buf* err) per exportable entry-file function. Exports are decided from the resolved tables the emission already had: entry-file functions (the prelude and imports carry their own file names) whose boundary types are scalars the C side can spell without marshalling — int / bool / float, unit results, unit parameters erased and passed as 0. Everything not exported is listed in a manifest comment WITH the reason (str/bytes marshalling is the next slice; a silently missing export reads as "covered").
A library must not do what main does. The standalone prologue installs a SIGSEGV handler, rebuffers stdout, and registers an atexit hook — all process-global policy that belongs to the host, not to a loaded library — so lib mode does none of them. Module init (the entry file's top-level bindings and its trailing expression) runs once, under pthread_once, inside the first wrapper call; a host that never calls mere_lib_init is still correct.
"Everything else is static" is a linkage claim, so the gate checks it with a linker's eyes rather than by grepping the C text: scripts/lib_check.sh builds the emitted C with -fPIC -shared, asserts the .so's exported symbol set is EXACTLY the boundary, links a real C host against it, and checks the values — calling a wrapper twice to prove init does not re-run. The internal symbols it closes off (mu_*, *_as_value, lifted functions, __lang_argc) were all external before, which a program never notices and two libraries in one host immediately would.
v0.1.307 — 2026-08-23
_One giant allocation taxed every block that came after it._
When an allocation missed the current block, the C backend's region allocator chained on a block of TWICE THE CURRENT CAPACITY and made it the bump target — even when the miss was a single giant request (a grown map index, a big string). That did two bad things at once: the remainder of the current block was stranded, and the giant became the base that every later doubling inherited, so a few hundred MB of data could sit inside several times as much capacity.
Now an allocation bigger than a quarter of the current block gets its own exact-size block, chained in BEHIND the bump block: freed with the region like any other block, invisible to map_recycle's roll-back walk (the seed is still the oldest block), but never the bump target and never the doubling base. Small allocations keep filling the block they were filling; doubling stays proportional to small traffic. On the forcing probe (one 256 MiB string in a stream of small cells) capacity drops from 1.99x of allocation to 1.0007x. The LLVM backend has the same switch-and-double design and is not changed here — same C-first precedent as the tail-call work (v0.1.230 → v0.1.267).
The meter landed with the fix: MERE_REGION_STATS=1 prints the default region's block count, capacity and cumulative allocation at exit (through main's epilogue or through exit(), whichever comes first). Capacity and allocation are functions of the program — unlike peak RSS, which is quantized to powers of two and stops reproducing above a few GB — so scripts/region_slack_check.sh holds their ratio to a bound, checks the meter against the probe's known floor in the other direction, and was watched to FAIL (1.99x) on the old allocator before it was trusted.
The meter's first real reading corrected a recorded number: the large Ruby workload whose default region was written down as "17.2 GB of capacity around ~2 GB in use" in fact hands out 9.0 GB cumulatively (cap 17.18 GB = a normal doubling tail, mostly untouched pages that never reach RSS). Its peak RSS moved within run-to-run noise under this change (5.7-7.6 GB across repeated runs of both binaries — above a few GB, RSS does not reproduce). What is left of that workload's footprint is the 9 GB itself: a collection problem, not an allocation-policy one.
Also wired thread_leak_check.sh and virtual_clock_check.sh into CI: both landed with v0.1.304/305 but were never run there — a gate that exists and does not run is a claim, not a check.
parity 129 passed / 0 failed; host_matrix ok; dune test 2561/0; thread_leak 7/0; virtual_clock 6/0; region_slack PASS.
v0.1.306 — 2026-08-22
_The bytes between the quotes were never looked at._
of_json accepted invalid UTF-8 inside a string on every backend that decodes JSON, because none of the three hand-written parsers processes \uXXXX escapes — a decoded string is exactly the raw bytes between the quotes, and nobody ever examined them. Go 1.27's encoding/json/v2 refuses such input by default (v1 silently rewrote it to U+FFFD, a lossy edit nobody asked for). Refusing is the only answer that neither invents bytes nor loses them. This is the second half of the Q-057 strictness slice; the first half (duplicate keys) landed in v0.1.303.
The rule is the PARSER's, not str's: utf8_len counts an invalid byte as one unit on purpose, because a str already in memory has no better answer. Bytes arriving as JSON do. The validator (shortest form, no surrogates, max U+10FFFF) went into the three parsers — the interpreter's parse_string_lit, the C runtime's __mj_pstr, the Wasm runtime's $__mj_pstr — so keys and nested strings are covered without writing the rule per type. lib/json.ml is untouched: it parses the language server's wire format, not user data.
test/parity/json_invalid_utf8.mere puts the accepting cases first, and they are multi-byte, not ASCII — a validator refusing every byte over 127 would fail loudly instead of passing a file that only asks whether rejection happens. Both sides of every boundary the validator special-cases are asked: U+D7FF then the surrogate just past it, U+10FFFF then F4 90, E0 A0 then the E0 9F overlong, C2 then C0/C1. The comparison is on of_json_opt returning None, not on a message (the backends word failures differently).
parity 129 passed / 0 failed; host_matrix ok; dune test green.
v0.1.305 — 2026-08-22
_The last thread to park advances the clock._
MERE_VIRTUAL_CLOCK=1 gives the interpreter Go 1.27's testing/synctest rule: WHEN EVERY LIVE THREAD IS PARKED, NOTHING CAN HAPPEN EXCEPT TIME, so the clock jumps to the earliest pending deadline instead of anybody waiting. A test whose timers add up to a minute finishes in milliseconds, and the ORDER in which they fire becomes a function of the program rather than of the machine's load. That order was previously untestable: a case asserting "the 10-second timer fires before the 30-second one" was a flake waiting for a loaded CI box.
This could not be built on top of v0.1.304's instrumentation, and the reason is the design constraint of the whole slice: that instrumentation marks a thread blocked and THEN parks it on the channel's own condition variable, and a clock that decides "everyone is parked, advance" inside that window signals a condvar nobody is on yet. The wakeup is lost — a flaky test, the disease a virtual clock exists to cure, built into the cure. So under the clock every wait parks on ONE scheduler condvar, and the blocked-count increment happens under the SAME lock as the park. The last thread to park is the one that advances the clock.
The first working version then span: the advancing thread holds the scheduler lock through its whole loop, and the deadline it just expired is still registered by its sleeping owner, so "advance to the earliest deadline" set the clock to where it already was, forever, while the owner starved for the lock the advancer never released. The fix is its own match arm: an already-expired deadline means its owner is woken-but-not-scheduled — the world is mid-step, not stuck — so the advancer parks (which releases the lock) and lets the owner run.
Routed through the scheduler: channel_recv / channel_recv_opt / channel_recv_timeout (whose deadline is fixed once — losing a value to another receiver does not extend it, and 0ms stays a non-blocking try-recv), sleep_ms / sleep, join (which parks on the registry saying the worker finished — status is written before the live-count drops, so the broadcast cannot arrive early — then the real Domain.join returns without blocking), and time, which reads the virtual clock. Its base is the real time of day at startup: only DIFFERENCES are deterministic, which is what a test may compare.
If every live thread parks and no deadline is pending, the program can only deadlock, and the run fails naming the situation rather than hanging — under a virtual clock a hang is never what a test meant.
Not routed: par_map's internal joins and OS-level waits (stdin, sockets). A thread in one of those counts as running, so the clock conservatively refuses to advance past it. The C backend is untouched — deterministic-time tests are an interpreter workload — and the default run of every program is byte-identical, which scripts/virtual_clock_check.sh checks from both sides: each case's virtual waits total 30-65 seconds against a 15-second wall bound (a secretly real clock gets killed, not passed slowly), each passing case must produce byte-identical stdout twice, and the deadlock case is re-run WITHOUT the variable and must really block.
Also: an uncaught failure in a worker used to vanish unless somebody joined the handle — the domain stores the exception for a join that never comes. The thread registry now records what killed it, and the leak report says died: fail: boom, never joined instead of the false finished.
v0.1.304 — 2026-08-22
_The runtime could not say how many threads it had._
spawn built a V_thread and dropped everything else on the floor, so nothing knew a worker existed once its handle went out of scope. A thread parked forever on a channel nobody sends to therefore cost nothing and SAID nothing: the program printed its last line and exited 0 with the thread still blocked. That is the gap Go 1.27 closed by promoting its goroutineleak profile, and the first half of closing it here is simply keeping a registry — which threads exist, and what each is waiting for.
MERE_THREAD_REPORT=1 prints, at exit, the threads that were NEITHER JOINED NOR DETACHED, and what each was blocked on:
mere: 1 thread(s) neither joined nor detached at exit thread 1: blocked on channel_recv
WHICH THREADS COUNT WAS THE HARD PART, and the language already had the answer. A server's accept loop is supposed to block for the life of the process, so "still blocked at exit" cannot be the test by itself — and a diagnostic that calls every well-formed server a leak is a diagnostic people switch off. detach is the language saying "I am not going to join this one", so a detached thread is disowned on purpose and goes unreported. Nobody claimed it and nobody disowned it is the condition.
This is the ANSWERABLE half of what Go does, not the same thing. Go looks for goroutines blocked on primitives that have become unreachable, which takes a tracing collector to decide; regions cannot answer that. Not covered either: the main thread (never registered — and a main that never returns is a hang, which is visible without a report), and the C backend, which has its own runtime and its own pthreads.
Off by default and written to stderr, because Go's profile is something you ask for too and a diagnostic on stdout would change what every existing program prints. scripts/thread_leak_check.sh holds both halves of that: it checks that the default run says nothing, and that turning the report on leaves stdout byte-identical.
The gate is BUILT AROUND ITS NEGATIVE CASES. Three of its six programs must report nothing — a joined worker, a deliberately detached blocker, and a program with no threads at all — because those are what separate a working report from one that fires on everything. Each program carries its own expected answer on its first line (//! leaks: 1 blocked on channel_recv) so that adding a case does not mean editing the script, and so a program whose expectation stops matching fails instead of being quietly renumbered.
Also: scripts/bounded.sh is the wall-clock wrapper v0.1.302 put inside parity.sh, moved out so this gate uses the same one. Two gates needing a bounded runner and defining it twice is how the two drift.
v0.1.303 — 2026-08-22
_A repeated JSON key was resolved by accident, in three different ways._
of_json accepted {"id":1,"id":2} on every backend that decodes JSON and kept the FIRST value. Nobody decided that. It fell out of building an association list in document order and looking it up front-to-back — and Go's encoding/json v1 kept the LAST for an equally accidental reason, which is the whole argument: two implementations resolve the same bytes to different values, so there is no right one to pick. Go 1.27's encoding/json/v2 stopped picking, and this does too. of_json fails; of_json_opt answers None.
The check went into the three PARSERS, one per backend — the interpreter's parse_object, the C runtime's __mj_obj, the Wasm runtime's $__mj_object — not into the per-type decoders generated from a record's fields. That is why a duplicate nested three levels down is refused too, and why the rule did not have to be written once per type. (lib/json.ml is untouched: it parses the language server's wire format, not user data, and holds no of_json traffic.)
test/parity/json_duplicate_key.mere CHECKS THE ACCEPTING CASES FIRST. A change that rejected everything would have passed a file that only ever asked whether rejection happens, so the clean object, the reordered one and the whitespace- heavy one come before the seven duplicates. It also asks both orderings — the duplicate early and the duplicate late — because a decoder keeping the last would have agreed with one and not the other.
The comparison is on of_json_opt returning None, not on a message. The three backends do not word a decode failure the same way: the interpreter appends (at offset N), the C runtime does not, and the Wasm runtime has no message channel at all, only a flag. Comparing messages would have pinned that unrelated difference instead of this behaviour.
Still accepted: invalid UTF-8 inside a string. Go 1.27 rejects both, but that half needs a validator that does not exist here yet, and it has to be reconciled with utf8_len counting an invalid byte as one unit on purpose.
v0.1.302 — 2026-08-22
_Two dogfoods had already written this search, and both wrote it the slow way._
str_index_of searched from the front and there was no counterpart from the back, so callers that needed the LAST occurrence wrote their own. Two did: contrib/url (userinfo ends at the last @, because a password may contain one; the port follows the last :; a path's final segment starts after the last /) and contrib/path (basename / dirname / ext). A third copy sat in contrib/url/ipv6.mere for the last colon of an embedded IPv4 quad. Each carried a comment saying the builtin did not exist -- which is the useful part of the record: the workaround was written down as a workaround rather than quietly becoming the way it is done.
Every copy scanned with char_at, and char_at returns a one-character str. On the compiled backends that is an allocation per byte examined, so finding one index in an n-byte string allocated n strings. str_last_index_of is a memcmp scan in all four backends and allocates none.
The empty needle answers the haystack LENGTH, not 0. It occurs at every position including one past the final byte, and this function reports the last of them -- which keeps str_index_of s "" <= str_last_index_of s "" true for every s.
Both lengths come from the length header (__lang_str_size in C, the same in the LLVM transcription, __lang_strlen in Wasm), never from strlen: a str has been able to hold NUL since v0.1.264, and test/parity/str_last_index_of.mere searches a haystack containing two of them for exactly that reason.
Also: test/parity/path_helpers.mere is new. contrib/path had no test of its own -- its only caller is the docs-site build -- so the three functions rewritten here had nothing holding them. Before deleting the hand-written versions they were run against the builtin over a corpus of paths and separators and agreed on every pair.
Not on RV32I: that backend does not have str_index_of either, and the pair stays symmetric.
Adding it took SIX registrations, not four. Besides the typer and the four backends, contrib/typer/typer.mere carries its own builtin type environment and contrib/codegen/codegen_wasm.mere its own known-builtin predicate, dispatch, WAT helper and helper-name table -- and the self-host tests compile contrib/path/path.mere, so the moment that file used the new builtin the bootstrap failed with unbound var str_last_index_of. Twice, once per layer. A builtin is not added until all six agree it exists.
The parity gate can now hold a program that blocks
scripts/parity.sh ran every program with no time limit, so a case that blocked forever would have hung the gate rather than failed it -- nothing reported, no log, and every case after it unknown. All eight execution sites (the interpreter reference and the three compiled runs, in both the ordinary and the supposed-to-fail loop) now go through a perl alarm wrapper, and outliving the limit is its own outcome, HUNG, named in the summary rather than folded into a DIFF. timeout(1) is absent on macOS and ulimit -t measures CPU time, which a thread blocked on a condition variable does not spend; neither would have worked. The limit is generous (60s, MERE_PARITY_TIMEOUT) because it exists to convert a hang into a report, not to time anything.
That was the prerequisite for the first concurrency cases this suite has ever had. There were 128 parity programs and not one of them spawned a thread, while the concurrency checks in test_basic.ml name (interp) in their own titles -- so "the interpreter and the C backend agree about threads" was not a claim anything held. Three cases now hold it: concurrency_channel (spawn / send / recv / join / close / recv_opt, with every output made order-independent by construction), concurrency_elem_types (what a channel can carry -- int, str, a str containing NUL, bool, a record -- which turns out to MATCH on all four backends, not just the two that have the full channel API), and channel_unconstrained_elem, a regression case for the v0.1.293 copier fault reached through a channel instead of a trait, which is the entrance that fix did not gate.
v0.1.301 — 2026-08-24
_A fail now releases the regions it jumps over._
Since v0.1.31, a fail that longjmps out of nested region R { } blocks restored the current-region pointer and let the abandoned blocks' chains leak -- noted at the time as acceptable, because regions were rare and failures rarer. Both stopped being true: an interpreter that wraps every method call in a region and models every guest exception as a fail leaks about a megabyte per caught failure (measured: 2,000 rescued fails = 2.1 GB peak; flat at 4 MB once released).
The C backend now keeps a thread-local ACTIVE-REGION STACK: block acquire pushes, block release drops (searching from the top, because the region R loop swap releases the arena under the one it just pushed), and try_or's catch arm unwinds the stack to the depth saved at entry -- releasing every chain the longjmp skipped -- before restoring the current region. Nested try_or nests the saved depths. Thread exit frees the stack's storage along with the cached block regions.
The interpreter is unaffected (OCaml exceptions unwind, the GC owns memory). The LLVM and wasm backends keep their previous behavior -- same values, backend-specific footprint; they get the same treatment when something that runs on them does fail-heavy region work. New parity lock: region_fail_unwind (fail through nested regions, nested try_or, fail out of a region loop -- every backend answers the same bytes).
v0.1.300 — 2026-08-23
_The frame-pool primitive._
map_recycle m: semantically map_clear, and on the C backend it also winds the map's private arena back to its warm seed block, freeing only the growth -- so a map reused as a CALL FRAME recycles at every return and the next use mallocs nothing. Promotes an unowned map on first use. The consumer this was built for holds a pool of frame maps: take one at call entry, recycle and return it at call exit unless the frame was captured (a proc, a binding, a define_method body) -- which is how an interpreter's per-call frames, its single largest unreclaimed share, come back without the env-indirection surgery that a store-based frame representation would have cost (measured at a 518-function transitive closure).
Borrowed internals dangle across a recycle, same discipline as map_compact. The interpreter and wasm lower it to map_clear exactly, so parity holds by construction.
parity 121/0 (the compact probe exercises recycle-reuse-recycle on all three), dune test 2541/0.
v0.1.299 — 2026-08-22
_A trigger cannot see byte pressure through an entry count._
map_bytes m / vec_bytes v: the bytes held by the container's OWN arena -- 0 until the first compact promotes it, capacity rather than bump-used (within 2x, costs nothing on the hot path, monotone between compactions: everything a collection trigger needs and nothing more). Found the way most of this week's primitives were found, by a consumer: a store of 20 KB values under-collected on any count-based trigger, and a large program's collector STARVED outright -- its amortization term grew with total-slots-ever while its trigger counted entries, so collections effectively stopped and an 87 GB peak looked like "GC present, memory anyway".
The value is deliberately backend-dependent (0 where no per-container arenas exist: interpreter, wasm) -- so it is for TRIGGERS, never for output. The parity probe uses it only in output-invariant positions, which is the rule for any program that wants to stay portable.
Blocks carry their capacity in the header's padding field now, which is what lets bytes be summed without a hot-path counter.
parity 121/0, dune test 2541/0, ctest.
v0.1.298 — 2026-08-22
_Emptying a map one key at a time was quadratic, and worse in a private arena._
map_clear m: length to zero, index wiped, O(index) and allocation-free. The missing mass-deletion primitive: map_delete keeps the dense arrays shifted and REBUILDS the hash index into the map's region on every call, so deleting half of a big map key by key is quadratic time and linear-times-index-size ALLOCATION. Found by v0.1.297's first consumer: a collector that deleted a store's dead entries one by one watched the store's freshly-promoted private arena double its way to 68 GB -- every delete allocating a new index into the very arena the compaction was about to reclaim. The pattern that replaces it: collect the survivors, map_clear, re-set, map_compact.
All backends answer the same bytes: interp (Hashtbl.reset), C (len = 0, index wiped), wasm (len = 0, in both the linear and the hash map runtimes -- the first patch landed in a third, LEGACY runtime that is emitted for nobody, and the parity gate said MISCOMPILE before the push instead of after).
parity 121/0 (the compact probe now exercises clear on all three), dune test 2541/0, ctest, selfhost_check.
v0.1.297 — 2026-08-22
_A container can hand back its own dead bytes._
map_compact m / vec_compact v: semantically invisible -- same entries, same order, same handle -- and allocator-visible. On the C backend the container's internal storage (arrays and every owned copy) moves into a fresh private arena and the previous generation is freed, so overwritten values, deleted entries' copies and abandoned grown arrays actually return to the OS. The first call PROMOTES the container to owning its arena; containers never compacted never pay (a per-call frame map costs nothing new -- promotion is lazy precisely because an interpreter makes millions of maps and compacts five). The interpreter and Wasm answer with a no-op, which is correct: compaction is an optimization, not an observable, so parity holds trivially.
Why a primitive and not a library: the alternative was measured first. Moving a program's stores into a region-loop carry works when the stores are threaded (the v0.1.296 consumer did exactly that for four of them), but a store accessed AMBIENTLY -- global helpers closing over it -- costs the transitive closure of every function on the access path: 518 functions in the first real program measured, on top of ~2100 direct call sites. A container that can swap its own generation makes the entire question disappear: the helpers keep their shape, the handle stays valid, and reclamation becomes one call at a safepoint.
The discipline that makes it flat, measured: caller temporaries die in a per-iteration region block, container churn goes at each compact -- 100k writes churning ~5 MB on each side peak at 1.7 MiB. Compact-only reclaims exactly the container's share (16.3 -> 11.6 MiB on the same probe with the caller's temporaries left leaking, which are the caller's business).
One rule the types cannot enforce: pointers handed out by map_get / vec_get before a compact dangle after it (the same safepoint discipline as region loops -- compact where nothing borrowed is held).
parity 121/0 (new: map_compact -- interp, C and wasm agree; llvm refuses maps as before), dune test 2541/0, ctest, stack_overflow, selfhost_check, lsp_smoke.
v0.1.296 — 2026-08-22
_A loop that carries its state between arenas._
region R { } is LIFO, and one thing genuinely does not fit LIFO: long-lived state that is periodically compacted. A compaction copies the live set into a NEW space and frees the OLD one -- the new generation must outlive the old -- and every attempt to write that with nested blocks puts the copy in a region that dies first (or, iterated, in the enclosing region, monotonically). The ML Kit hit this wall thirty years ago; it is the reason "region inference degenerates to one global region" for non-LIFO programs.
region R loop x { body } is the hand-over-hand form. Each iteration runs in a fresh arena named R. The body sees x : option C -- None on entry, Some carry after -- and answers region_flow[C, D] (a new prelude type): Continue carry deep-copies the carry into the next arena and releases the current one; Done d copies d out and exits. The typing rule is the block's with one split: the carry may mention R -- a Map[R, ...] crossing arenas is the construct's whole point -- the Done value may not. No rigid type variables were needed, though the road here assumed they would be: nothing of arena N reaches arena N+1 except through the Continue copy, so the one name R honestly denotes "the loop's current arena" in every iteration. The construct the compaction story was blocked on turned out to be a loop, not a quantifier.
Continue's copy goes through a new family, __mdeep_<tag>: __mcopy with the one difference that containers are copied into the target region instead of passed as handles (a handle would dangle the moment the old arena is released). A carried Map is rebuilt entry by entry -- _set re-hashes and copies keys and values into the new arena, so the copy IS the compaction: dead overwrites simply stay behind. What __mdeep cannot copy is refused at emit time, by type: functions (captures hide behind a void it cannot see into), Channel, ThreadHandle, OwnedVec, StrBuf, ByteBuf.
Measured: 2000 cycles overwriting 50 map keys 1000 times each -- 2.4 GB of dead entries churned, 51 entries live -- peak footprint 1.6 MiB, 0.32 s. The same shape without the loop is the monotonic default-region growth every long-running Mere program has today.
Backends: interp and C; llvm / wasm / rv32 refuse cleanly (their reclamation is a LIFO bump rollback, which cannot express the swap). Errors are typed: a Done value mentioning R is a type error naming both rules; an uncarryable type in C names itself and why.
parity 120/0 (new: region_loop_carry -- interp and C agree, llvm/wasm documented-refuse), dune test 2541/0, ctest, stack_overflow, selfhost_check, lsp_smoke.
v0.1.295 — 2026-08-22
_One cached region is not enough once regions nest._
The per-thread cache of block-region structs was one slot deep. That is the right size for regions that strictly alternate, and the wrong size the moment they nest -- a loop iteration's region inside a statement's region, statement regions inside the body -- because releases then arrive out of cache order: the innermost release fills the slot, the next-outer release finds it full and frees its region entirely (struct, 1 MB seed block and all), and the next acquire at that depth buys them back. An interpreter built on this runtime paid a 1 MB malloc+free round trip per loop iteration -- ~200 GB of allocator churn over a 200k-iteration run, read from its own block counters. libc absorbed it well enough that wall-clock barely moved, which is why it went unnoticed until the counters said it out loud.
The cache is now a stack of eight. Deeper nesting than eight falls back to malloc, which is correct and merely slower. The spawn trampoline drains the whole stack at thread exit (same leak the one-slot drain existed for).
parity 119/0 (the first build broke it -- the thread-exit path still assigned to the old single slot; the gate caught it before push), dune test 2536/0, ctest, stack_overflow, selfhost_check, socket_parity, tcp_read_codes.
v0.1.294 — 2026-08-22
_Every formatted value leaked its scaffolding._
Both native backends build a formatted string in an asprintf buffer, copy it into the current region at the __lang_str_of_cstr boundary -- and never free the buffer. One vasprintf buffer is 160 bytes on macOS, so a million str_of_int calls leaked 48 MB. The list formatter was worse: it rolled its accumulator through asprintf on every element (__acc = __buf), leaking the whole previous prefix each step -- quadratic bytes in the list length.
Found from the outside: mere-ruby's new bench/region_reuse.sh asked whether a region whose block chain GREW past one block hands its memory back in a reusable form (the property a compaction loop stands on), and the answer looked like no -- peak footprint grew linearly, ~38 MiB per iteration, while the runtime's own block counters insisted everything was returned. A minimal C probe cleared libc (the same malloc/free pattern is flat at any iteration count), and leaks named the 300,000 vasprintf buffers.
The boundary now has an owning twin: __lang_str_take_cstr copies into the region and frees its argument. All 13 asprintf sites in the C backend and all 6 in the LLVM backend go through it; getenv/argv/literal sites stay on the borrowing one. The C list/json formatters free their rolling accumulator.
After the fix the original question answers itself: released grown chains are fully reusable through plain libc -- 1, 8, and 32 sibling ~64 MB regions all peak at 36-39 MiB. No runtime block pool is needed.
parity 119/0, dune test 2536/0, ctest, stack_overflow, selfhost_check, url/encoding parity, debug_info, wasm_sourcemap, lsp_smoke -- all run locally before push, plus leaks --atExit reporting zero on the probe and on a list-show program.
v0.1.293 — 2026-08-21
_Whoever names a copier decides that it exists._
v0.1.290 taught __mcopy for an arrow type to deep-copy a closure's env, which put copier CALLS in two places that had never emitted one before: a record's copier, for a field holding a closure, and an env copier, for a captured value. CI went red for four releases -- 16 of 119 parity programs, every trait_* one among them -- with the emitted C naming a function that was never defined:
v.mu_add = __mcopy_closure_int_closure_int_int(r, v.mu_add); / undeclared /
Two independent faults, each hiding behind the other's failure.
The first: `ty_tag` names types that `ty_is_concrete` rejects. ty_tag erases an unresolved type variable to int, so a capture of list (str, 'a) is named __mcopy_list_tuple_str_int -- while the collector that decides which copiers to emit walks the same type, finds it not concrete, and drops it. A name with no definition. Registration now goes through ty_as_tagged, which returns the type ty_tag actually names, including the region slot that ty_tag renders as __heap (erasing that to int would reintroduce the mpng P5 shape: one type under two names).
The second: a polymorphic record's fields were read three times, and two of the readers had a different answer. A trait dictionary Num__dict 'a is monomorphized at its instance -- for Num__dict m7 the struct field is closure_m7_closure_m7_m7 -- but the copier emitter read r_fields directly and named __mcopy_closure_int_closure_int_int, the wrong type under the wrong name, while the collector dropped the field entirely for holding a TyParam. Only the closure-typedef collector, which had hit this before, substituted the arguments. The rule now lives in one place, record_field_types_at, and all three read it.
Also: the inner-lifted half of the env-copier list is no longer filtered on captures = []. A closure with no captures has no env at all and needs no copier; an inner-lifted fn always gets an env struct and its use site assigns __copy unconditionally. Measured, not reasoned -- restoring the filter fails exactly one parity program (prop_list).
Gates: parity 119/0 (was 103/16), dune test 2536/0, plus ctest, stack_overflow, host_matrix, selfhost_check, url/encoding parity, infer_scaling, debug_info, wasm_sourcemap, lsp_smoke, window_check -- none of which CI had reached since v0.1.290, because it stops at the first failing gate.
v0.1.292 — 2026-08-21
_The copier moves into the env, and a closure is two pointers again._
v0.1.290 put the env copier in the closure struct, which made a closure three pointers. v0.1.291 answered the cost by emitting two different closure shapes -- three pointers for programs that use a region, two for those that do not. That is two rules where there should be one, and it left the memory win unavailable to exactly the programs that wanted it: an interpreter that adds a region block to reclaim its temporaries pays the pointer on every frame of its dispatch, and one of CRuby's bootstraptest pairs loses its stack (ld caps -stack_size at 512 MB on arm64, so there is no room to buy back).
The copier now lives in a header on the ENV: {__lang_region* __r; void* (*__copy)(__lang_region*, void*);} at the front of every env struct. __mcopy for an arrow reads it through the void* it already has. The closure struct is {env, fn} again, for every program, and the region-conditional emission and the typer flag it needed are gone.
An FFI adapter used to hold a borrowed pointer directly as its env, which a header cannot be read from. Those envs are now real structs -- a header plus the borrowed pointer -- allocated in the default region with __copy = NULL, which says "do not copy me, I am permanent" in the same field that says "copy me like this" for a generated env. One rule, read the same way everywhere.
Measured on an interpreter written in Mere: with a region block per statement, 200k plain method calls hold 1181 MiB against 1358-1996 before, the same number three runs out of three, and run in 1.16-1.20s against 1.43-1.58s. Its corpus is 157/157 either way.
Two sentences in the first version of this entry were wrong, and the correction is worth more than they were. They said the memory win now costs nothing and that CRuby's bootstraptest was back to its baseline. Neither holds: the pair v0.1.291 was written around still fails with the closure at two pointers, so the third pointer was not its cause -- or not its only one. That pair runs a thread with while true; // =~ "" end beside a regex loop, and whether it finishes depends on which limit trips first, the interpreter's step budget or the native stack. The same source overflows at -O1 and does not at -Os: a canary on the line, not a measurement of a change. The interpreter's own record now says so, so that its err=59 is not read as a fresh regression.
The reasons to prefer this shape stand without those sentences: one closure layout instead of two, no per-frame pointer, and an FFI env that answers the same question the same way a generated one does.
2536 tests pass across four backends.
v0.1.291 — 2026-08-21
_An env records its region, and a program that uses no region pays nothing._
v0.1.290 gave a closure a copier for its env and then copied unconditionally. That is neither idempotent nor bounded. An env that captures a closure copies that closure's env in turn, so a chain of them is a chain of copies, and a threaded test in an interpreter written in Mere overflowed the native stack: stack overflow (recursion too deep), where the same program passed before. Copying also duplicated identity — two copies of one env are two mutable states, and a closure that writes through what it captured would write to the wrong one.
Each env now carries the region it was allocated in, and the copier returns its argument unchanged when that region is the destination. Copies inside one region — the common case, and every copy in a program with no region blocks — become no-ops, and a cross-region copy still walks only as deep as the chain that actually crosses. This is the elision the memory model deferred as needing type-level region tracking on values; one pointer per env does it at runtime.
Found by CRuby's bootstraptest, run against an interpreter written in Mere: 1697 pairs, of which exactly one moved from pass to error. Neither that interpreter's own 157-program corpus nor this repo's 2533 tests noticed, because the shape needs a closure captured inside a closure inside a thread. A wider gate is worth having even when it is somebody else's.
The idempotence fixed the copying, and did NOT fix what the bootstraptest pair was actually measuring: v0.1.290 made the closure struct three pointers instead of two, and a closure is passed BY VALUE through every frame of an interpreter's dispatch. One pointer per frame took that program over its stack limit -- and there is no room to raise it, because ld refuses -stack_size above 512 MB on arm64, which is exactly where the interpreter already was.
So the copier is now emitted only for programs that USE a region block. With no region block __lang_current_region IS the default region, an env cannot outlive its region, and the closure keeps its pre-v0.1.290 shape: two pointers, an env in the default region, a shallow copy. Charging a cost to programs that cannot benefit is the wrong trade, and the measurement said which programs those are. Verified end to end: the interpreter rebuilt without a region block is back to pass=1569 err=58, its baseline, and its generated C carries no copier field at all.
The flag that answers "does this program use a region" is set by the typer, not during emission -- the typer walks the whole program first, and a closure may be emitted before the region block that appears later in the file. It is reset per compilation, because the language server and the test harness both compile many programs in one process, and a leftover true would give a region-less program the shape of whatever was compiled before it.
v0.1.290 also shipped with version.ml left at 0.1.289: its changelog entry was written after the suite ran, so the version test's failure was not seen before the commit. Both are corrected here.
v0.1.290 — 2026-08-21
_A closure carries the copier that lets it leave a region._
Closure environments were allocated in the default (program-lifetime) region, and the reason was written where the deep-copy family is defined: "closures copy shallowly: their envs live in the default region". A closure's env is a type-erased void*, so nothing could copy it into another region — and a permanent allocation was the only answer that could not dangle. The cost is invisible in a small program and decisive in a large one: an interpreter written in Mere allocated 2548 closure envs' worth of default-region references against 1346 current-region ones, so every call leaked whatever it closed over.
A closure now has a third member beside env and fn: void* (*copy)(region*, void*), generated per env type next to the env's own typedef, where the field types are known. It region-allocates a fresh env, copies the struct across, then re-copies each captured field through that field's own __mcopy_<tag> — so a captured string is deep-copied rather than left pointing into the region that is about to be released. __mcopy for an arrow calls it when it is set, and a closure leaves a region block the way every other value does.
Two shapes deliberately keep the old behaviour, and get it for free from C's zero-initialization of omitted designated initializers: a closure with no captures passes env == NULL, and an FFI adapter holds a borrowed pointer it does not own. Both leave copy zero and are copied shallowly.
The test that asserted the old rule now asserts the new one, in the same three places it was checked: the allocation follows __lang_current_region, the closure literal carries .copy = __mcopy_env_..., and a captured str reaches the copier. 2533 tests pass across four backends; the change is C-only, since the other backends do not share this representation.
What this does NOT do on its own: an interpreter with no region blocks has __lang_current_region == &__lang_default_region, so nothing moves until the blocks exist. It removes the reason they could not pay off.
v0.1.289 — 2026-08-20
_A record field is a C struct member, and one name answered two questions._
The other half of the keyword surface. v0.1.286 routed a closure CAPTURE's name through c_safe_name; a user record's fields never went through anything, so type t = { short: str } emitted const char* short; — a keyword where a declarator belongs. Same argument as v0.1.56's, which chose a uniform mu_ prefix over a reserved-word list because the list "was inherently incomplete and recurred six times", so the same prefix is used and there is one rule to know rather than two. Fields live in a per-struct namespace, so a prefix collides with nothing; what it buys is that no C keyword reaches a declarator.
Fourteen sites: two definitions (the struct body and the monomorphised one) and twelve uses — literal, update, field access, pattern binding, show_, to_json, from_json, eq, cmp, __mcopy, and a map key's equality and hash. Plus two hand-written runtime literals, for mk_logger and mk_metrics, which spell the field names out and had to spell them the same way.
One real bug came out of doing it, and it is the interesting one. from_json used the same fname twice in one format string — once as the C designator and once as the JSON KEY. Prefixing both renamed every key in every serialised record, which four parity programs said immediately. The designator is an identifier and the key is what the document says; one name, two questions, and only one of them wanted namespacing.
And the loud-failure property earned its keep. Eleven assertions in the suite pin this struct's spelling from various sides and all eleven went red at once, which is exactly what v0.1.56 argued a uniform prefix would buy: a path that forgets it breaks every record rather than only the unluckily-named ones. Separating "the assertion is stale" from "the codegen is broken" was one measurement — the four JSON parity programs compile and run — and after that the assertions were bookkeeping.
Two sites were missed on the first pass and both were found by gates rather than by reading: the cmp line, because the replacement written for it was character-identical to the original and the patch skipped it, and mk_logger's runtime literal. host_matrix named the second in one line. mk_metrics is still nocompile on the C backend and was before this — a function pointer type mismatch, unrelated.
Verified on Darwin: dune runtest 2531/0, parity 119/0, ctest 14/0, host_matrix with no change, selfhost ok, stack_overflow ok. Verified on Linux in a container: fifteen programs through the C backend and eleven through LLVM, including a record with short, long and inline as field names, all matching the interpreter.
v0.1.288 — 2026-08-20
_The LLVM backend names a stack overflow on Linux, which took dlsym and a computed array length._
Three releases and two wrong shapes to get here, all of them the same underlying fact: LLVM IR has no preprocessor, so a runtime detail that differs by platform has to be selected at run time or not at all.
v0.1.285 made the Darwin-only pthread pair extern_weak and guarded the call, which stopped the link failure and left the bounds unknown off Darwin. v0.1.287 found that the handler was not even being installed there — stack_t, struct sigaction, the SA_* values and SIGBUS all differ — and selected seven measured constants off one icmp. What was left was the bounds, and the obvious move does not work: declaring pthread_getattr_np extern_weak breaks the DARWIN build, because Mach-O's linker refuses an undefined weak reference where ELF resolves it to zero. Measured, not assumed — it was tried and reverted.
dlsym(RTLD_DEFAULT, …) asks at run time and needs no reference at link time, which is what wanting an optional symbol actually calls for. Measured: it is in libc on both platforms, so no extra link flag, and RTLD_DEFAULT is itself platform-dependent — 0 against -2 — which the same icmp selects. glibc reports the LOW address and the size, the other way round from the Darwin pair.
platform C backend LLVM backend macOS/arm64 names it names it Linux/x86_64 names it names it Linux/aarch64 names it names it
So the divergence is gone and both the things that recorded it are gone with it. The parity pin for uncaught_stack_overflow is deleted, and scripts/stack_overflow.sh expects one answer everywhere instead of a per-platform one — if a platform stops naming the fault that is now a regression rather than a fact about the platform. The gate's second mode, for a backend expected not to name it, is removed rather than left: nothing reached it, and a branch nothing reaches cannot be told from a branch that is wrong.
One thing worth the line because it was caught late and cheaply: the three symbol-name constants had their array lengths written out by hand and two of the three were off by one. They are computed from String.length now. LLVM catches that, but only after the file is emitted, and nothing about the source made it visible.
v0.1.287 — 2026-08-20
_The LLVM handler installs itself on Linux, which needed seven measured constants._
v0.1.285 stopped the LLVM backend failing to link off Darwin by making the two Darwin-only pthread calls extern_weak and guarding them, and left the fault unnamed there. What it did not do was notice that the handler was not being installed at all: stack_t, struct sigaction, the SA_* flag values and SIGBUS are all different on glibc, so sigaltstack was handed a size where it expected flags, refused, and the early return meant no handler ever arrived. The process died with the shell reporting the signal.
The platform is a runtime fact here, not an emit-time one. IR has no preprocessor, and choosing when the IR is written would make -ll output specific to the machine that produced it -- which it has never been. pthread_get_stackaddr_np is already declared weak, so its address is non-null on Darwin and null everywhere else: one icmp and every constant below is a select.
Measured on macOS/arm64, Linux/x86_64 and Linux/aarch64 rather than recalled, because a wrong offset here writes a flag word into a signal mask and the handler simply never comes:
si_addr in siginfo_t 24 / 16 stack_t ss_size, ss_flags 8, 16 / 16, 8 struct sigaction sa_flags 12 (of 16 bytes) / 136 (of 152) SA_SIGINFO | SA_ONSTACK 65 / 134217732 SIGBUS 10 / 7
The observable difference on Linux is small and exact: Segmentation fault from the shell becomes segmentation fault from our own handler, and the exit status becomes 1 -- what the interpreter exits with -- instead of 139. So parity's pinned divergence for uncaught_stack_overflow moves from EXIT(139) to MSG: the process now fails the way the interpreter fails and only the message differs. The pin catching that is what it is for.
The name still needs the bounds, and that is one measured step away. The glibc pair is pthread_getattr_np + pthread_attr_getstack, and declaring them extern_weak does not work: ELF resolves an undefined weak reference to zero and Mach-O's linker refuses it outright, so adding them broke the Darwin build. dlsym with RTLD_DEFAULT is the portable way to ask for an optional symbol, and RTLD_DEFAULT is itself platform-dependent (0 against -2), which the same icmp can select. Not done here: it is a separate change with its own verification, and shipping it half-measured is how this release's predecessor got the layout wrong.
v0.1.286 — 2026-08-20
_A capture the walker forgot, twice, and the shape that needs three things at once._
The browser dogfood is the first program here to ask the C backend to compile something large, and it did not compile at all: 29 errors, in two families. The first was field names not going through c_safe_name and is fixed above. This is the second, and it is the one that had been invisible for a reason worth writing down.
`known` decides what is NOT captured, and `host_locals` is what rescues a name from it. known is builtins plus top-level names plus externs — things referenced directly in the generated C rather than through an env. host_locals is subtracted from it, which is how a frame-local that happens to share a builtin's name gets captured anyway. The top-level driver has always passed [f.param] for exactly this. Two places lost it.
A curried parameter's name was discarded. The walker's own Fun case read Ast.Fun (_, _, body), so descending through fn (cs) -> fn (id) -> ... recorded cs and forgot id — and id is the identity builtin, so it sat in known, and an inner let rec reading the parameter captured nothing. The lifted function then named an identifier that was never declared, reported by the C compiler thousands of lines from the decision.
And descending into a lifted body reset the list. walk_in_fn p [] fn_body threw away the enclosing frame's names, so one lift further in the outermost parameter looked like something already in scope. One level worked because the enclosing frame was the top-level one, whose parameter the driver had recorded.
Three things have to line up, which is why neither showed up before. The name shadows a builtin; it is a curried parameter rather than the first one, or read from a nested lift; and it is read from a lifted function. Any two of the three and the program compiles.
Two parity programs, one per defect, each verified to catch its own by reverting the fix. The second was written with an ordinary name first and passed with the fix reverted — it never reached the code it was for, and only poisoning said so. It names its parameter after a builtin now, and the comment says why.
MERE_LIFT_DEBUG=1 prints what this pass decided: per lifted function its host, parameter and captures, and per lift the body's free variables alongside the names known blocked. Kept rather than deleted after use — a missing capture surfaces as use of undeclared identifier in generated C, and the alternative to reading this is guessing which of the four skip conditions in the fixpoint fired.
The browser dogfood compiles now, and the measurement its north-star gate owed can be taken: a 559-byte page in 45 ms, of which 19 is loading the font and 18 is style and layout, at a 39 MiB high-water mark.
v0.1.285 — 2026-08-20
_"On every backend" was verified on one platform, and CI said so for four days._
v0.1.271 named the stack overflow, and its title says "on every backend". It was true on Darwin. On Linux it stopped the C backend compiling and the LLVM backend linking — every program, not an edge case — and parity has been red since the commit after it. Four consecutive days of failures, latterly hidden behind a stale version assertion.
Three platform-specific defects in one runtime block.
pthread_getattr_np is glibc's way to ask a thread for its stack, and glibc declares it only under _GNU_SOURCE. Without the macro the emitted C called an undeclared function, which clang 16 and later treat as an ERROR — and -w, which the parity harness passes, does not silence an error. The #define now comes before every header, which is the only place it works.
static char __lang_sigstack[SIGSTKSZ * 4] is a variable length array at file scope on modern glibc, where SIGSTKSZ expands to sysconf(_SC_SIGSTKSZ) rather than a constant. It is a literal 65536 now, comfortably above MINSIGSTKSZ on both platforms.
And the one that had no preprocessor to hide behind. The C backend picks between the Darwin pair and the glibc call with #ifdef; LLVM IR cannot, so the Darwin names went out on every target and nothing linked. They are declare extern_weak now and the call is guarded: on Darwin they resolve and the fault keeps its name, and elsewhere they are null, the bounds stay unknown, and the handler falls through to the plain segmentation fault it already reported for a fault outside the stack. That costs the NAME off Darwin and nothing else, and it needs no platform detection at emit time — so -ll output is still the same file wherever it was produced. A name derived from an assumed stack size would be wrong for any program linked with a bigger one, which is worse than no name.
What let it through is the shape worth keeping. The suite checks this feature by looking for stack overflow (recursion too deep) and sigaltstack in the EMITTED TEXT. Those assertions stayed green throughout, because a string is present whether or not the file it is in compiles. scripts/stack_overflow.sh runs a program that recurses until it dies and reads what it says, per backend, with the per-platform answers pinned rather than smoothed over. It is in CI.
platform C backend LLVM backend Darwin stack overflow (recursion too deep) stack overflow (recursion too deep) other stack overflow (recursion too deep) a plain crash, no name
Measured on Linux in a container, not inferred: the C backend now compiles and matches the interpreter on the first twelve parity programs, where before none of them built, and it names the overflow there for the first time.
v0.1.284 — 2026-08-19
_Generated inputs for the differential gates, and five NUL-length defects they found._
contrib/prop generates values; it does not compare them. A property test written with it is an ordinary parity case, so the other four backends are the oracle and there is nothing to commit as an expectation — the same shape as the browser dogfood's reftest, moved from pixels to values.
Three decisions, each a way this could have measured nothing. It does not use random_int: that builtin asks the host, five hosts would draw five sequences, and every line would report DIFF while the gate reported on itself. Values are a function of an index rather than a carried state, so a differing line in the diff already names its input — which is why there is no shrinker, the smallest reproducer is i and it is printed. Every intermediate stays below 2^45, because the interpreter's int is 63-bit and the compiled backends' is 64-bit.
The edge tables are not a draw. Every integer defect this suite has found sat on one of those values and a uniform draw reaches them with probability about zero, so a generator that only drew would be weaker than the hand-written gates it extends. NUL is in the byte pool on purpose.
Five defects on the first run, all one family
Code that was correct while a str ended at its first NUL, and silently wrong after v0.1.264 gave a str a length header.
| builtin | backend | what was wrong |
|---|---|---|
str_contains | C, LLVM | strstr — a needle beginning with NUL matched every haystack |
str_index_of | LLVM | strstr — answered 0 where the C backend's length-aware search answers -1 |
str_starts_with | LLVM | strncmp, and no length check on the haystack — the only one of the four starts/ends implementations across the two backends written that way |
str_join | C | the sizing pass asked __lang_str_size and the copying pass asked strlen, so an element holding a NUL was measured at its full length and copied only to the NUL. One function, two notions of how long a string is |
str_ptr, read_lines | C | found by sweeping every strlen in the backend after the first four, not by waiting for a gate to point at them |
Two recorded rather than fixed
Each pinned in its own case so it keeps being asked.
test/parity/prop_utf8.mere — UTF-8 character splitting. For "A" ++ chr 128 the interpreter says two characters and the three compiled backends say one, so utf8_at returns two bytes and codepoint_at then refuses. Index 0 is "A" and what follows cannot change that, so the interpreter is right. The span computation is prelude Mere, one source compiled by all four, which means the difference is under it — a separate measurement, not a guess to make while writing a test.
test/parity/llvm_loop_guard_global.mere — a loop whose bound is a top-level binding stops after one iteration on LLVM. All three ingredients are needed and each was removed in turn to check: the bound must be a top-level let and not a literal, the body must evaluate a try_or whose thunk calls a prelude function, and it must be recursive. A program folding this way would silently process the first element and report success.
The axes are kept apart
prop_int leaves multiplication out and prop_list leaves list_product out, both with the reason written down: their operands overflow, the interpreter wraps at 63 bits where the compiled backends wrap at 64, and a gate measuring the width axis while claiming the value axis reports the wrong one when either moves. codepoint_at moved out of prop_str for the same reason.
test_basic: "llvm: declares strstr" asserted how index_of is implemented rather than what it answers. It now asserts the memcmp and that strstr is absent.
Gates: unit 2527-0, parity 117-0 + failing 15-0 (5 declared divergences), ctest 14-0, selfhost 7-7, html_tokenizer ok.
v0.1.283 — 2026-08-18
_Two shadowing bugs, and a gate whose expectation was degenerate._
Both bugs had the same shape: the compiled backends resolve a name globally, and something shadowed it. The interpreter was right in both cases, which is what made them findable at all.
Q-045: a user binding that shadows a builtin the prelude calls
let show = fn (x: int) -> x + 1;
print (str_of_int (show 1))
Three lines, and it made the prelude fail to type — expected 'str', got 'int' at <prelude>:485 for a program that never mentions the prelude — on all four compiled backends. The prelude's pow calls the show builtin. Not show-specific: str_len and list_len are each called three times in there.
Two independent causes, and the first hid the second.
The desugared program was typed against the accumulated environment. Since desugar_program turns every Top_let into a nested Let, the expression rebinds all of them itself — so passing the accumulated environment added nothing except the one thing it must not: a binding visible to declarations that come before it. It now types against an environment holding only what desugaring drops, which is externs.
That turned the error into a refusal from the monomorphiser, because uniquify_toplevel_shadows never saw the user's show as a shadow — a builtin is not a top-level declaration, so the first binding of that name kept it. It is seeded with the builtin names now, so the user's binding is renamed and references before it — the prelude's — still mean the builtin.
The seed subtracts the prelude's own top-level names: the prelude deliberately shadows ten builtins (pow, divmod, assert, …) and those keep the names they have always had.
Q-046: a parameter named the same as a top-level binding
uniquify_inner_fns_expr renames inner fn bindings that collide with a top-level name, and left Fun parameters alone. So a lifted inner function took no parameter for its captured handler at all, and its body referred to the caller's global of that name.
It was invisible while both had the same type — the wrong binding happened to fit — and surfaced as C that would not compile only when the types diverged. contrib/http2 carried a naming workaround for it; the parameter is renamed now, so that workaround is a comment about history rather than a rule.
Parameters get a different treatment from inner fns and the distinction is the point: an inner fn becomes a symbol and must reserve its name, a parameter does not and only has to not be mistaken for one.
Three test assertions were pinning a symbol name
§30.0 checked that a user-defined is_alpha shadows the builtin by grepping the output for int mu_is_alpha(. The fix renames it to is_alpha__v2, so they matched the prefix instead. Measured before changing: the compiled program answers true and the builtin answers false for the same input, so the behaviour under test was unaffected — the exact name only said it by accident.
The line-break gate was comparing against nothing
linebreak_conformance reported 19338 of 19338 cases differing. want.txt had 19338 lines and not one break mark in it: under this machine's LANG=ja_JP.UTF-8, awk 20200816 compares ÷ equal to × — collation, not bytes — so every mark became ×.
The implementation was right and the expectation was degenerate. It runs under LC_ALL=C now. It passes on GNU awk and under the C locale, which is why CI was green and a Japanese-locale machine was not — the environment difference to suspect first when a gate disagrees with CI. The data is ASCII plus those two symbols, so byte comparison was what was wanted all along.
Verified: 28 CI gates, runtest 2526/0, parity 110/0 + 15/0 (two new parity programs), selfhost_check all passed, host_matrix ok.
contrib — GraphQL introspection and validation — 2026-08-18
_No compiler change. contrib/graphql, two new gates._
Introspection is answered by the ordinary executor
__schema, __type(name:) and __typename. The introspection types are ordinary SDL, generated from graphql-js into introspection_sdl.mere and appended to the document's own definitions — so __Type is an object type like any other and field lookup, nullability, null propagation, list handling and enum coercion all apply to it unchanged. One line, and there is no second executor.
The SDL is generated and committed, like the Unicode and HPACK tables: the specification fixes every name, type and nullability, and a 200-line transcription has a mistake in it.
Introspection is cyclic, so one value has to be lazy. __schema.types lists every type, each type's fields name types, whose fields name types; a strict value cannot hold that and eager construction does not terminate. What bounds the expansion is the query — getIntrospectionQuery() asks for ofType exactly nine levels deep. So gvalue has one non-data arm, GTypeRef of gtype, computed when asked for: the only place in this executor where a field is computed rather than looked up.
The strongest check is not a comparison. Our introspection result fed to graphql-js's own buildClientSchema and printed must equal printSchema(buildSchema(sdl)). It holds exactly when our answer carries the whole schema, it is blind to field order (which the specification does not fix), and nothing transcribes an introspection result. It passed before the JSON comparison did, and both differences that then surfaced were real:
- A built-in scalar belongs to a schema only if something refers to it. The
oracle's type map for type Query { a: Int } is Query Int Boolean String and not the other two — Boolean and String because the built-in directives refer to them, which falls out of walking the appended SDL rather than being special-cased.
- graphql-js 17 keeps an argument's default in `a.default.value`, not
a.defaultValue. Reading the old field silently dropped @deprecated's default.
`includeDeprecated` defaults to false, and the standard introspection query passes true — so every gate section using it agreed while that was missing entirely. It took a hand-written __type(name: "Colour") { enumValues } to disagree.
Validation, and how a partial validator is gated
19 of the specification's 32 rules. The gateable part is the design:
Every error carries the name of the rule that produced it, and the harness compares the set of rule names against graphql-js running each of its 32 rules individually — the oracle classifying its own output. Not the error list, for two measured reasons: graphql-js returns errors in visitor order interleaved across rules, so a partial validator could never agree about anything; and rule names are about thirty identifiers from the specification's own section titles, small enough that a shared misreading is not a real risk.
Three failures, not one: rejecting what the oracle accepts (the worst — a false positive fails a valid request), missing a rule we claim, and missing a rule we do not claim (DOCUMENTED-GAP, and the rule must be listed). A gap-list entry that never fires fails the harness as stale.
The claim list is checked in both directions: every rule reported must be on it. Without that, dropping a rule from the list while still implementing it left the harness green — poisoning found it, because the list is otherwise consulted only for rules that were missed. A wrong list silently weakens every check that reads it.
The document and the schema are separate arguments, unlike the executor. ExecutableDefinitions is the rule that a document being executed must not contain type-system definitions, so a combined list makes every schema violate it — the first version reported every SDL type as "not executable".
What the harness caught: we rejected { __schema { queryType { name } } } because __schema / __type / __typename are provided by the schema rather than declared in it; is not defined **by** operation versus is never used **in** operation, two prepositions and one helper that put the wrong word in one of them; and a fragment cycle is reported once, at the first fragment in document order.
A harness bug worth naming
Comments inside a node script passed in a double-quoted shell string may contain neither a backtick nor a double quote — one is command substitution, the other ends the string. Both mistakes were made in consecutive edits. The second turned the comparison into a syntax error rather than a wrong answer, which is the good failure of the two.
contrib — HTTP/2 flow control and gRPC streaming — 2026-08-18
_No compiler change. contrib/http2, examples/grpc_hello.mere, and two gates._
Both directions of the flow-control window, SETTINGS_MAX_FRAME_SIZE, and gRPC methods that send or receive more than one message. grpcurl now gets server-streaming, client-streaming and bidirectional answers from a Mere server, and a 300 KB request and a 300 KB reply both go through.
The limits bite in an order, and the first one is not a window. Measured against a python h2 client before any of the code existed:
| bound | bites at | what the client does |
|---|---|---|
wbuf capacity, unchecked | 16385 | nothing — wrote past the buffer, returned success |
SETTINGS_MAX_FRAME_SIZE | 16385 | refuses the frame outright |
| the flow-control window | 65536 | accepts no more DATA until it grants credit |
| inbound, the same window | 65536 | stalls at exactly 65535 with nothing from us |
A fix that added window accounting without splitting frames would still have failed at 16385. The wbuf row is the one with no symptom at all: mem_copy_bytes is a memcpy into a bump arena that records no allocation sizes, so nothing below could catch it. Demonstrated with a sentinel — 32 bytes into a 16-byte allocation changed the byte after it and reported 32.
Flow control makes the writer a reader. A send whose window is exhausted has to consume the WINDOW_UPDATE that reopens it, so send_data contains a reader.
An export list is not a coverage list
H2.read_settings read each SETTINGS entry's 32-bit value at offset i+4 instead of i+2 — two bytes past every six-byte entry, so the last entry of any payload read off the end. Wrong for as long as the file has existed, and nothing called it: the writer had a parity section from the start, the reader was exported, documented, and never fed a byte, and the server's own comment said its peer's settings were "read and ignored". The first caller found it on the first connection, at byte 42 of a 42-byte payload.
Enumerating every export and counting call sites turned up a second accessor in the same state — H2.error_code, zero callers and zero coverage. It happens to be right. http2_parity.sh now feeds every payload accessor.
What poisoning found that valid traffic could not
- A well-behaved client cannot see a server that ignores its window. Poisoning
the check to "a billion bytes" left the section green: a client that acknowledges as it reads has already granted more credit by the time the next chunk arrives. The gate now has a rude client that advertises 8192 and stops reading.
- `SETTINGS` adjusts a spent window by the delta; it does not reset it — and the
two rules agree exactly when nothing has been spent. The SETTINGS now arrives mid-response, driving the window to -57343; negative is legal, and a reduction after credit was spent is the only way to get there.
- The SETTINGS branch of the window-wait loop recursed under a comment saying it
returned to its caller. A peer may reopen a stalled stream with `SETTINGS` alone, and this server would have waited forever. Found only because the delta poison was undetectable for the same reason.
- One rule was written twice. A
WINDOW_UPDATEincrement of 0 is a protocol
error; the frame loop ignored it and the window-wait loop failed on it. The wrong one was on the path every ordinary frame takes.
- `SETTINGS_MAX_FRAME_SIZE` could not bind. Its minimum is the default, so with
a 16384-byte write buffer min(peer, capacity) was always the capacity and reading the setting could not change the answer. The buffer is 65536 now, which makes the term live rather than deleting it.
Two wrong expectations, both the same slip
The boundary in the field is not the boundary on the wire: a 16384-byte reply field is 16393 bytes of DATA once the gRPC prefix, the protobuf tag and a 3-byte varint are added, so it already needs two frames. That got the frame-count expectation wrong, and later a window grant that left the server nine bytes short.
Also h2's local_settings.initial_window_size = ... is silently ignored before initiate_connection — the client advertised 65535, the server correctly sent 65535, and the harness blamed the server.
Streaming is the list having more than one element
The handler takes the request's messages as a bytes list and answers RpcOk, RpcStream or RpcErr; H2Server.unary adapts a one-in-one-out handler. There is no separate path for client-streaming. Interleaving is not supported and is said so rather than approximated: the handler runs after the request half-closes.
examples/hello.proto gained five methods and contrib/proto needed no change — stream was already parsed, and our descriptor for the new schema matches protoc byte for byte.
v0.1.282 — 2026-08-18
_The FFI byte arena gets a bytes bridge._
mem_to_bytes / mem_copy_bytes, the arena's two directions for data that may contain a zero byte — which mem_to_str cannot carry, because it stops there, and which every binary protocol has. Native and Wasm-component both, and scripts/socket_parity.sh requires the two to agree.
What it replaced. contrib/http2/server.mere was reading arena → hex → bytes (two characters per byte, three passes) and writing with `mem_set_u8` once per byte. The write is the one that mattered: a response is now one FFI call instead of one call per byte of it. Same shape as v0.1.222, where file_pwrite could not take a bytes and built one boxed int per byte.
The Wasm helper is a few instructions because the arena is linear memory there. The one thing to get right is that a bytes pointer points at the length, where a str pointer points at its data with the length at ptr-4; a comment says so, and poisoning the layout both ways is caught.
Two mistakes, and the second is the interesting one.
First, the definitions went where the rest of the arena lives — and the generated C said unknown type name 'mere_bytes', because b->data needs the complete struct and that block is emitted earlier. So they moved into the bytes runtime.
Then they were emitted there unconditionally, and __mem — the arena — is only emitted when a program declares an arena extern. Every program that uses bytes without the arena stopped compiling, which is most of them; proto_parity said so on the next run. Two independent conditions, and satisfying one is what made the other easy to miss.
The gate for that turned out to need one more thing: bytes_runtime was a top-level constant, evaluated at module-initialisation time — before any program is parsed — so a Hashtbl lookup inside it would have been false for every program that ever declared the externs. It takes the flag as a parameter now.
Poisoning found a coverage gap too: reading the length with i32.load8_u instead of i32.load passed, because on a little-endian machine the low byte of a length under 256 is the length, and every payload in the corpus was shorter than that. There are 255-, 256- and 300-byte payloads now.
v0.1.281 — 2026-08-18
_IEEE-754 bit access, in two halves because one would not fit._
float_bits_hi / float_bits_lo / float_of_bits, and f32_bits / float_of_f32_bits for float32. All five on all four backends, closing the last thing that stopped a protobuf double or float field from being generated.
The API shape was measured, not chosen. The obvious design is one accessor returning the whole 64-bit pattern. It does not work: read as a signed int64, a double's pattern exceeds this interpreter's native int — which is OCaml's, and 63-bit — for a large share of ordinary values. -1.5, 1e308, inf and nan all do. A single accessor would therefore answer differently on the interpreter than on every compiled backend, for a literal as plain as 1e308, and the honest options were a DIVERGE pin or a different API. Two 32-bit halves are each below 2^32, so there is nothing left to diverge about — and they are also exactly what a wire format wants, since it writes the bytes anyway.
f32_bits narrows to float32 with the backend's own double-to-float conversion, which is round-to-nearest-even. That matters more than it sounds: 1e308 has no float32, and the answer is +inf rather than a wrapped number. Writing that rounding by hand is the part nobody should have to get right twice.
contrib/proto/wire.mere gains put_double / get_double / put_float / get_float on top, so the split stays inside that layer and neither a caller nor a generated codec learns about it. The generator's refusal of double and float is gone, and scripts/proto_gen_parity.sh covers them — 25 schemas now, including 1e308, 1e-308, float32's maximum, and repeated floats, all byte-identical to protoc.
Two mistakes worth recording, both mine and both in the checking rather than the code.
The Wasm arm shifted with i64.shr_s. The top bit of a double's pattern is its sign bit, so an arithmetic shift sign-extends it and float_bits_hi (-0.0) came back as -2147483648 — a value that then compares as negative and divides the wrong way, on that backend only. It is i64.shr_u, and the comment now says why.
Finding it took two detours. First I read a stale binary: dune build 2>&1 | head -2 had exited on SIGPIPE, so the compiler under test was the poisoned one from a mutation run that a timeout had killed mid-restore. Then I "verified" the restore with grep -c 'i64.shr_u', which returned 1 — from a different, pre-existing occurrence elsewhere in the file. A check that cannot tell the two states apart is not a check; the verification is anchored to the arm now.
test/parity/float_bits.mere pins every class — normal, subnormal, both zeros, both infinities, NaN, the range ends — on four backends. NaN is checked through its bits rather than through ==, because == on NaN is false by IEEE and comparing it the obvious way reports a failure that is the comparison's own rule.
2026-08-18 — A protobuf code generator, and two branches nothing reached
_contrib/proto/gen.mere + examples/protoc_mere.mere: a .proto in, Mere source out. examples/grpc_hello.mere now uses it, and its hand-written codec is gone._
mere examples/protoc_mere.mere examples/hello.proto ../contrib/proto/wire.mere \
> examples/hello_pb.mere
The codec in that example used to be a hand-written field walk and a two-line encoder. Both were correct, and both were a schema transcribed by hand — which is the thing a generator exists to stop.
scripts/proto_gen_parity.sh runs protoc's bytes through the generated codec and back: protoc --encode → decode_M → encode_M, byte-compared. 20 schemas. A round trip through the oracle's bytes catches more than it looks like, because a field the decoder ignores cannot be written back and shows up as a shorter byte string. The one thing it cannot see is a consistent swap of two fields holding equal values, so every field in the corpus has a distinct value.
Three representation choices, each forced by proto3 rather than chosen: a singular message field is a 0-or-1 list (message fields have explicit presence, so "absent" and "present but empty" differ — and it makes a recursive message expressible); an enum is an int and not a variant (a proto3 enum is open, and an unknown number has to survive a round trip); and a generated identifier never begins with the type name, because uppercase-leading is how this language recognises a constructor — M_get_a parses as one and is reported as an unknown constructor at its use site, naming something the schema author never wrote.
The encoder writes fields in field-number order, not declaration order, because protoc does.
Poisoning found two branches nothing reached:
- The repeated zigzag and fixed families were not in the corpus. It had singular
sint32 and repeated sint64, so the repeated-sint32 path was generated and never executed — breaking it changed nothing.
- protoc always writes packed, so the decoder's unpacked branch was never reached.
proto3 requires a decoder to accept both forms whatever the writer chose, and deleting that branch left the harness green. A decoder that only understands the encoding its own oracle emits works against exactly one kind of writer.
The second one needed input protoc does not produce, so the harness builds it — and the first hand-typed version had field 2's tag as length-delimited, which protoc rejected and the harness reported as its own bug, correctly. Hand-written test bytes are a transcription like any other, so they are computed from the values now. protoc is then asked whether the two spellings mean the same value rather than being told.
2026-08-18 — A .proto parser, and the bootstrap closing
_contrib/proto/parse.mere + contrib/proto/descriptor.mere: a .proto file in, a FileDescriptorSet out, byte-identical to protoc --descriptor_set_out across 72 derived files._
The bootstrap closes here, and that is why it is one harness rather than two. A descriptor set is itself a protobuf message, so the code that reads a schema is serialised by the code that reads wire bytes. If the varint encoder is wrong, this diff says so; if a descriptor field number is wrong, the same diff says so. One oracle checks both layers at once.
The comparison is bytes. A descriptor that decodes to the same text but different bytes is still a different descriptor to anything that hashes or caches it.
Three things were measured off the oracle before anything was written, and each would have been wrong from memory:
`descriptor.proto` is proto2, which reverses a rule the wire harness had just learned. A set field is written even at its default, so an enum value's `number: 0` appears on the wire — where proto3 omits a scalar holding its default and `v: 0` could not be swept at all. Same encoder, opposite rule, decided by the schema being encoded rather than by the encoder. `json_name` is not plain camelCase. my_field → myField is the easy one; a_b_c → aBC, trailing_ → trailing, __lead → Lead and num_2_x → num2X are the ones a guess gets wrong. The rule is "drop underscores, upper-case what follows each one". `type_name` carries a leading dot and is fully qualified, and a reference resolves from the innermost scope outward — so the same simple name is a different type at a different depth, and the wrong order still produces a well-formed descriptor.
The last of the 72 files to match differed by two bytes. rpc U (Q) returns (R) {} is not the same as rpc U (Q) returns (R); — the empty body sets MethodOptions to an empty submessage, so protoc writes 22 00 there and nothing for the semicolon form. Presence with no content is the whole difference.
The subset is refused, not skipped. oneof, map, import, option, reserved, optional, extend, group and proto2 are outside it, and protoc accepts every one — so the oracle cannot be asked whether they are wrong. The harness checks instead that each is refused by name, because a skipped construct produces a descriptor that is wrong where nothing looks. Getting that right needed two fixes the corpus found: an absolute type reference (.pk.M) is a different production from a name and needed its own reader, and a field option list met "expected ';'" — a syntax error about a file that is syntactically fine — until it was given a named refusal.
Ten poison mutations, none uncaught.
2026-08-18 — RPC statuses, and an expectation the wire corrected
_A gRPC handler can now fail._
bytes alone could not say "this failed", so an unknown method got a reply that looked like a success. A handler now answers RpcOk of bytes or RpcErr of (int * str), and a failure goes out as a status in the trailers — the request succeeded as HTTP and failed as RPC, so it cannot be expressed as a 4xx, and a failed unary call carries no DATA frame at all.
`grpc-message` is percent-encoded, which is neither optional nor cosmetic: the specification says so and grpc-go decodes it. Measured before implementing — caf%C3%A9 %25 done came back as café % done — so a raw % would have been read as the start of an escape.
And then the wire corrected the harness. The first expectation written for the raw trailer assumed a space and an apostrophe would be encoded. They are not: the rule is "bytes outside printable ASCII, plus % itself", the same as grpc-go's own encoder. The implementation was right and the expectation was a guess; the failing diff is what said which.
The DOCUMENTED-GAP that asserted the absence of statuses did its job on the way out: closing the gap made that assertion fail and report itself stale. Three routes are checked in its place — a method the schema declares and the server does not handle, a method that exists and refuses its input, and a message that needs encoding — and the python client now reads the raw trailers. That is the observation grpcurl cannot give: it prints the decoded message and the rendered code name, so whether the encoding happened is invisible there. Poisoning found nothing uncaught across five mutations, including lower-cased hex digits.
2026-08-18 — A GraphQL executor, and the one case out of 67 that mattered
_contrib/graphql/exec.mere plus scripts/graphql_exec_parity.sh against graphql-js's execute()._
The resolvers are the data. A field is resolved by looking its name up in the parent object, which is graphql-js's default. Both sides get the same schema, the same document and the same root value, so nothing about resolution is transcribed — two hand-written resolver sets would drift and the drift would look like an execution difference.
What that arrangement really exercises is null propagation, and it is why an oracle earns its place. A non-null field that resolves to null is not an error in that field: it destroys the nearest nullable ancestor. { nn } answers "data": null. An item error in [Int!] nulls the whole list with a path of ["ln", 1]. All of it was measured off the oracle before being implemented, because reading behaviour off a specification and reading it off a running implementation are different activities.
Every position is one of two things — GOk (value, errors) or GBubble errors — and the four cases in the specification fall out of those two constructors instead of needing exceptions the language does not have.
Of 67 corpus cases exactly one differed, and it was the interesting one. { o2 { nnx } } with o2: O! produced two errors: the real one about O.nnx and an invented one about Query.o2. The oracle produces one. When a non-null position is null because an error already propagated, no second error is raised — the check that would raise it is never reached, because the propagation went past it. The fix splits the job: raw completes a position with no absorbing, and complete decides what the position's nullability means. A list item goes through complete, so [O] nulls the item and [O!] nulls the list, which now falls out rather than being special-cased.
Poisoning the gate ten ways caught nine immediately. The tenth is worth recording as a coverage lesson: breaking the named fragment's type-condition check went unnoticed because the corpus only had the inline form. Two nearly identical branches, one of them unchecked. Three named-spread cases closed it, and the poison then failed as it should.
Error messages are compared verbatim — the specification does not fix the prose, so matching the reference implementation is the only way to compare them at all, and the oracle is pinned. locations are stripped from both sides: they need source positions the parser does not carry, which is a stated gap rather than an accident.
v0.1.280 — 2026-08-18
_A capture whose name has a dot in it._
An inner let rec that referenced a module-qualified name put that name into the closure's environment struct verbatim, so the C backend emitted
long long Wire.delimited;
and the generated C failed to parse. The error pointed at a line the program did not contain, which is the part that makes this worse than a refusal: the diagnostic was about generated C, not about the source that caused it. The interpreter was always correct, so nothing saw it until a dogfood compiled a wire-format decoder with Wire.delimited inside a field-walking loop.
The shape is narrow, and measuring which variants were affected is what located it. Four were fine — M.k at the top level, a top-level closure using it, that closure passed to a function, and an inner let rec capturing only its enclosing parameter. Only an inner let rec referencing a qualified name reaches the environment-struct emitter instead of being referenced directly.
The fix is that two paths now agree. The inner-LIFTED environment (__env_local->…) already ran capture names through flatten_module_dots; the anonymous-closure environment did not. Three sites — the struct's field declaration, the adapter's substitution, and the creation site that fills the environment in — now flatten too, with the substitution still keyed on the name as written because that is what Var n in the body says.
flatten_module_dots is the identity for a name without a dot, so every program that did not hit this emits byte-identical C. And every output that does change was previously invalid C — a dot in a field name has never been legal — so there is no program whose working output moved.
test/parity/module_qualified_capture.mere pins it on all four backends, including the variants that were not broken, so a future change cannot fix one path and regress the other.
2026-08-18 — A gRPC server in Mere, and two bugs loopback was hiding
_contrib/http2/server.mere plus examples/grpc_hello.mere. grpcurl asks a Mere server for a greeting and gets one._
$ grpcurl -plaintext -protoset hello.protoset -d '{"name":"mere"}' \
127.0.0.1:50079 hello.Greeter/SayHello
{
"message": "hello mere"
}
Everything under that line is Mere: the protobuf wire format, the HTTP/2 framing, HPACK, and the connection. TLS is absent by measurement rather than oversight — h2c is what grpcurl -plaintext speaks, 24 literal bytes and no negotiation.
scripts/grpc_parity.sh is the first harness here that does not compare bytes. Every layer underneath already has a byte-level gate; none of them answers whether a client that knows nothing about any of it gets a reply it recognises. Two clients, because they ask different questions: grpcurl (Go) over four connections including an empty proto3 request, and a python h2 client making three requests on one connection — the case grpcurl cannot express for a unary method, and the one that catches a per-stream HPACK decoder.
Poisoning it found two bugs that every other section passed:
- Removing the frame reassembly changed nothing. On loopback a request arrives
in a single read, so a server that parses whatever one read handed it works. The comment in server.mere had said reassembly "is not a simplification, it is a bug that happens to pass on a fast loopback" — and the harness proved the comment right by failing to catch its removal. The python client now sends one request split every seven bytes, so the 9-byte header itself straddles two reads.
- Removing the SETTINGS acknowledgement changed nothing, because neither client
blocks on it. The client now asserts the ACK arrives.
Both are the mirror of the dead guards found in frame.mere yesterday: there, code that nothing reached; here, behaviour that nothing observed.
Two harness bugs were worth more than they cost. The readiness check waited for the port by connecting to it — which the server counts, since it serves a bounded number of connections and then exits, so the probe ate one and the last real call got "connection refused" while the harness blamed the server. It waits for the server's own "listening" line now. And the unknown-method section printed its DOCUMENTED-GAP without calling anything: a claim, not a check. It calls a method the schema declares and the server does not handle, and requires the documented answer.
One compiler bug came out of writing the example and is recorded rather than fixed: an inner let rec that references a module-qualified constant makes the C backend emit the qualified name as a struct field — long long Wire.delimited;, which is not valid C. Six lines reproduce it, the interpreter is correct, and the error names generated C rather than the program that caused it.
2026-08-18 — HPACK, and three ways a gate can pass without checking anything
_contrib/http2/hpack.mere plus a generated table and a seven-section gate. The implementation went green on the first run; poisoning it is what produced the interesting part._
What had to be complete was measured. grpcurl 1.8.8's first HEADERS frame on a real connection uses static-table indices, Huffman-coded literals and seven dynamic-table insertions — so the decoder is complete (integers, both string forms, all five instructions, eviction) while the encoder is trivial: literal, new name, no indexing, no Huffman, which RFC 7541 permits and which a real client was confirmed accepting end to end. That asymmetry is why a gRPC server is reachable without writing a Huffman encoder.
The dynamic table is connection state and is threaded through the API rather than hidden in a global, because one decoder per connection is correct and two is not, and a parameter makes that visible at the call site.
The tables are generated from the same library the gate uses as its oracle, which is a hole a differential test cannot close: a shared transcription error is invisible. So one section decodes a header block captured from grpc-go 1.57 — a third implementation, in another language, that never saw either table. Nine header names and the table's own accounting come back matching.
Then the gate was deliberately broken, nine ways. Six were caught immediately. The other three were the point:
- Valid input does not test a refusal. Removing the Huffman padding check
changed nothing, because well-formed blocks have well-formed padding. Nine malformed blocks now exercise the refusals — the same lesson the GraphQL reject-list taught, arriving again in a different file.
- "Does it refuse" and "does it say what was wrong" are different questions.
Removing the index-0, EOS and string-length checks still refused, further downstream, with a message about something else — so the gate stayed green while three named diagnostics had been deleted. Each malformed case now asserts a substring the refusal must contain. That couples the harness to our own wording, which is the acceptable direction; coupling it to the oracle's would be brittle for no gain.
- A check whose input never arrives passes. The case file was written with two
columns while the reader expected three, so every expectation was the empty string and grep -q "" matched every message. An empty expectation is now a failure in itself — and so is a reject-list that checked zero cases, which is the rule the builtin matrix learned in another form: a gate reporting ok for no cases cannot be told apart from one that did not run.
One line survives poisoning on purpose. len2 < huff_min_len is a fast path, not a check — huff_count already answers 0 below the minimum length — and it is labelled so the next reader does not have to work that out. The contrast with the lshr mask deleted yesterday is deliberate: that one was a guard that could not fire, this one is an optimisation that provably cannot change an answer.
2026-08-18 — GraphQL's type-system half, and HTTP/2 frames
_contrib/graphql grows SDL; contrib/http2/frame.mere arrives with a gate against hyperframe. Both gates were then deliberately broken to see whether they could fail, and both had blind spots._
SDL — schema / scalar / type / interface / union / enum / input / directive, with descriptions, implements A & B, argument definitions and defaults. The derived corpus goes from 153 to 366 documents and the round-trip covers all of it unchanged, because graphql-js's print handles type-system definitions too.
Then the half that was missing: everything so far fed valid documents, and a parser that accepts anything passes all of it. A reject-list of 37 documents the oracle rejects found six real defects on its first run:
`{ }`, `type T { }`, `enum E { }` and `input In { }` were all accepted. Every delimited list in this grammar needs at least one element except the two that are values — `[]` and `{}` are legal, `{ }` as a block is not. One helper that names the production replaced seven call sites that each returned `[]`. 007 was accepted (an IntValue may not have a leading zero) and so was 1. (a fraction needs a digit). The lexer now also refuses a number followed by a digit, a . or a name start, so 1.2.3 is an error rather than two tokens.
A harness defect came out of the same exercise: set -e plus a crashing subject made the script stop mid-run having printed no verdict, hiding every later section. The subject now runs through a helper that never aborts and names the section it failed in. set -e stays for the preflight, where a missing oracle should stop everything.
Type-system extensions (extend ...) are refused rather than mis-parsed — reading extend type T as type T would produce a tree that says something the document did not — and the refusal is asserted, so the day extensions land that section fails and says the assertion is stale.
HTTP/2 frames — the preface, the 9-byte header, SETTINGS / WINDOW_UPDATE / RST_STREAM / GOAWAY payloads, and the gRPC message prefix. hyperframe is the oracle: 50 frames byte-identical on encode, the same 50 compared on fields on decode, with the sweep crossing type × flags × stream id × payload size (0, 1, 2, 255, 256, 16383, 16384 — every length-encoding boundary).
One byte has two independent confirmations, which is the only kind of agreement worth having: an empty SETTINGS frame serialises as 000000040000000000, and that is byte-for-byte what grpcurl 1.8.8 was observed sending on a real connection. hyperframe and grpcurl never met.
Poisoning that gate found two things the sweep could not reach:
The `lshr` mask was dead. Every shift here is immediately masked with 255 to extract a byte, and the bits an arithmetic shift copies in sit above bit 7, so a logical shift gives the same answer. Removing the mask changed nothing, which is the definition of dead code. It is not dead in `contrib/proto`, where the shifted value is the loop variable and a negative int64 would never terminate — the contrast is now written down in both files. The writer's stream-id mask was never exercised, because every stream id in the sweep already has the reserved bit clear. A section now writes 0x80000001 and requires the same frame as stream 1. The decode direction was already covered: 0x80000001 must read as 1, not 2147483649.
Both are the same shape — a guard nothing reaches is indistinguishable from a guard that is wrong.
2026-08-18 — Two protocols: the protobuf wire format, and GraphQL documents
_contrib/proto and contrib/graphql, each with a gate against somebody else's implementation._
`contrib/proto/wire.mere` — the Protocol Buffers wire format, both directions, below any schema. scripts/proto_parity.sh holds it to protoc's bytes: one message containing every wire type compared whole, swept value lists crossing every varint length boundary, the int64 edges, and protoc's own bytes read back and rewritten. Interpreter and C backend both.
Three things the harness had to be taught, each by being wrong first:
`bit_shr` is arithmetic, so a negative int64 shifted right never reaches zero and the encoder loops forever. Every right shift here is masked. 0 cannot be swept. proto3 omits a scalar holding its default, so protoc --encode answers the empty string for v: 0 while a schema-less layer writes 0800. That is a question one layer up, not a wire-format disagreement. The zigzag bound is half the varint bound, because zigzag doubles its input.
And one thing that is pinned rather than fixed: the interpreter's int is 63-bit (OCaml's native int) while every compiled backend's is 64-bit, so above 2^62 the same program gives different answers — bit_shl 1 62 is negative on the interpreter and bit_shl 1 63 is zero. Those are recorded as a DIVERGE pin, not a tolerance, so the day the interpreter becomes 64-bit that section fails and says the pin must be retired. Related: an integer literal above 2^62−1 cannot be written at all — it dies with an uncaught Failure("int_of_string") — so the edge values in the harness are computed from 2^62−1 instead. This is the int axis of the open question about builtin parity at particular values, and it is the same shape as the float axis before exponent literals landed in v0.1.260: the reason it was not measured is that it could not be written.
`contrib/graphql` — lexer, parser and printer for executable documents (operations and fragments; SDL is a separate grammar and is not here). scripts/graphql_parity.sh checks it against graphql-js by sending our output back through their parser:
theirs: print(parse(D)) -> A ours: print(parse(ours(D))) -> B assert A == B
print is a function of the AST alone, so equality holds exactly when the parses agree — and nothing transcribes an AST. The alternative, serialising both trees into a shared format, needs a serialiser for the oracle's tree that can hold the same misreading as the parser it checks; a differential gate that shares a bug with its subject reports agreement. The corpus is derived rather than written down: every value kind × every position admitting a value, every selection form, every operation shape, nested type expressions. 153 documents.
The gate was then deliberately broken to see whether it could fail. Dropping every directive in the printer: caught. Dropping field aliases in the parser: caught. Making the lexer call every number an IntValue: not caught — the printer emits a numeric literal's lexeme unchanged, so IntValue "1.0" prints 1.0, the oracle re-parses it as a FloatValue, and the two printed documents agree. A round-trip is blind to any distinction that prints identically. A fourth section now asks for the kind directly, and it is the one place that transcribes anything from the oracle's tree — a two-word vocabulary.
The printer refuses a block string whose value would not survive re-parsing rather than emitting an ordinary string: the dedent rule is applied again on the way back in, and downgrading it would round-trip the text while losing block: true from the tree — a wrong answer wearing the face of a parser bug.
Both gates are in CI, both run under dash before being put there, and both oracles are pinned and printed (protoc 27.3, graphql-js 17.0.2).
v0.1.279 — 2026-08-17
_The refusals that were left, and three bugs they walked into._
The enumeration's tail: bytebuf_get / bytebuf_set, write_file_bytes's byte range, len on a value that is not a list, and comparing two functions. Asking them turned up three things that were not about refusals at all.
A program that used a ByteBuf and no `bytes` value did not compile. The C backend emits the bytes runtime when bytes_used is set, and freezing a ByteBuf calls __lang_bytes_alloc — the comment above the emit even says so, so the order of the two runtimes had been thought about and the dependency had not. Third time a use-gate has emitted a call to a function it did not emit.
`len (Some 1)` emitted C that dereferenced `payload.Cons` on an option. The guard for the cons-walking branch asked whether the program contained Nil and Cons anywhere, not whether this type has them, so any program that also mentioned a list took that branch for every polymorphic variant. It asks the type now, and a type with no length is a clean codegen refusal naming the alternatives.
Comparing two functions was a runtime failure on the interpreter and invalid C on the compiled backends — == between two closure structs does not compile. The type is known where the comparison is written, so the typer answers there now, with the interpreter's own words.
The refusals themselves: ByteBuf's index failures and the byte-range check both printed and called exit(1), so try_or could not take them and the program stopped where the interpreter carried on. They are catchable failures now, with the messages they already had. test/parity/bytebuf_edges.mere is the gate — interp + C, since bytes/ByteBuf is those two backends by scope.
Also: __lang_str_count still returned i32 on LLVM after v0.1.276 widened it on C and Wasm — a count past two billion came back negative. The sweep missed exactly one backend of the three.
v0.1.278 — 2026-08-17
_The last of the oracle's refusals._
The bytes value has the same index surface as vec, and it had all the same problems one slice later than vec did:
| before | after | |
|---|---|---|
| C, LLVM | abort() — exit 134, and nothing try_or could catch | the interpreter's catchable failure |
Wasm bytes_get | trapped with no message, after truncating the index to 32 bits | checked at full width, and it says what happened |
Wasm bytes_slice | no range check at all — copied from wherever the arithmetic pointed | refused |
random_int on the Wasm host accepted any bound. There is no RNG wired there yet, which is a documented limitation — but "no RNG here" and "any bound is fine" are different statements, and the second one is a wrong answer. The bound is checked now even though the value it returns is still a deterministic zero.
That closes the enumeration started in v0.1.276: every refusal the interpreter raises is now raised by all four backends, with the same words, catchable in the same way. test/parity/refusals.mere covers the caught side and there are 15 programs in test/parity/fail/.
v0.1.277 — 2026-08-17
_Finishing the enumeration, and two things it turned up on the way._
v0.1.276 asked the oracle's refusals and fixed the nine it had probed. This is the rest of that list.
`bool_of_str` answered false for anything that was not "true" — on all three compiled backends, with a comment in the C source stating that this "matches interp". The interpreter has never done it: it refuses any word that is not one of the two. A claim in a comment is not a check.
`float_of_str` was three different wrong answers to the same program:
| input | interp | C / LLVM (atof) | Wasm (parseFloat) |
|---|---|---|---|
"abc" | refuses | 0.0 | nan |
"1.5x" | refuses | 1.5 | 1.5 |
"1_000.5" | 1000.5 | 1000.5 | `1.0` |
"inf" | inf | inf | `nan` |
"0x1p3" | 8.0 | 8.0 | `0.0` |
"" | refuses | refuses | nan |
atof reads a prefix and has no way to say "that was not a number"; on Wasm the host's parseFloat has the same problem plus a smaller idea of what a float is, and its failure value is NaN — which is indistinguishable from the float nan, so nothing could be refused there at all. Validity now comes back from the host separately, because one return value cannot say both.
The oracle is float_of_string (String.trim s), which accepts rather more than a decimal point: hex floats, inf / infinity / nan in any case, signs, and underscores as separators (1__0.5 is 10.5). All four backends now agree across all of it, and test/parity/refusals.mere checks fourteen spellings.
Two things the new gate found on its way in:
str_unescapeallocated its buffer for the input's length, and every escape
makes the output shorter — so the length header said the wrong thing, and str_len (str_unescape "a\nb") was 4 on C and LLVM where the interpreter and Wasm said 3. Printing the value wrote a trailing NUL, which is invisible in a terminal. It took a case that printed an unescaped string to see it.
- The Wasm host's worker env is a second import table for the same module, and
adding an import to the module without adding it there makes instantiation throw inside the worker — after which the main thread waits on a generator that will never yield. That is a hang, not an error: it cost twenty minutes of a parity run before I looked at the process list.
v0.1.276 — 2026-08-17
_Asking the question mechanically, after finding the same shape by accident twice._
v0.1.274 found string sizes narrowed to 32 bits; v0.1.275 found indices narrowed the same way, one layer down. Both were found by a probe that happened to ask. So the third time the question was asked from the other end: every refusal the interpreter can raise, enumerated from eval.ml, checked against what the compiled backends do with the same input.
Nine diverged, and always in the same direction — the compiled backends had no check at all and returned whatever the byte or pointer arithmetic produced:
| input | interp | C / LLVM / Wasm (before) |
|---|---|---|
chr 256 | refuses | a NUL — the cast to unsigned char was the domain check |
chr (0-1) | refuses | 0xFF |
chr (2^32+65) | refuses | "A" |
ord "" | refuses | 0 (it read byte 0 whatever the length was) |
ord "ab" | refuses | 97 |
str_repeat s (0-1) | refuses | "" |
str_unescape "a\q" | refuses | "aq" — an undefined escape became the letter |
owned_vec_get v 5 | refuses | abort(), which try_or cannot catch |
Every one is a value a program can produce by accident and then keep computing with. All four backends now raise the interpreter's own failure, with its exact words: chr: 256 out of byte range [0, 255], ord: expected single-char str, got length 2, str_repeat: negative count -1, str_unescape: unknown escape '\q'. test/parity/refusals.mere is the caught side; uncaught_chr_range and uncaught_ord_length join the uncaught ones.
The same sweep turned up three more narrowings and closed them: str_count's result, strbuf_len's result, and sleep_ms's argument were all int.
A note on method: the harness for the first measurement had a bug of exactly the kind these slices keep finding — when a backend failed to compile, the shell variable kept the previous backend's output, so one column of the table was a copy of another. It was caught because random_int has no LLVM lowering and the "LLVM" column reported values anyway. A measurement that cannot fail loudly is not a measurement.
v0.1.275 — 2026-08-17
_An index the collection does not have — answered, for years, with an element._
let v = vec_new ();
let _ = vec_push v 10;
let _ = vec_push v 20;
vec_get v 4294967297 // interp: refuses. C, LLVM, Wasm: 20.
An index is a Mere int, sixty-four bits of it, and every compiled runtime took one as thirty-two. The bounds check then ran on the truncated value, so 4294967297 was checked as 1 and a two-element vec cheerfully returned its second element and exited 0. All three compiled backends agreed with each other, and none of them agreed with the interpreter. No test caught it because every index in every test was small — the same blind spot the string widths had in v0.1.274, one layer down.
Out of range was its own mess, three ways at once:
| before | after | |
|---|---|---|
| interp | catchable failure naming index and length | unchanged — it was the oracle |
| C | fprintf + abort(): exit 134, and nothing try_or could catch | the interpreter's failure |
| LLVM | abort(), no message | same |
| Wasm | unreachable: a trap, no message | same |
char_at / substring, all three | read past the end and returned what was there | same |
Every backend now raises the interpreter's own failure, which names both numbers at full width: vec_get: index 4294967297 out of bounds (len = 2). On C and LLVM the message is built with snprintf — LLVM's needed its varargs signature spelled out at the call site, or the arguments arrive through the wrong ABI slots and the message prints numbers nobody passed. Wasm has no snprintf, so it builds the sentence from interned parts with its own show_int and str_concat.
test/parity/index_edges.mere is the caught side (nineteen positions, in range, one past the end, negative, backwards, and past what 32 bits can name); uncaught_vec_index, uncaught_char_at_index and uncaught_substring_range are the uncaught side.
What it found immediately. The Ruby subset's utf8_cp decodes a sequence's continuation bytes before checking they exist, and on binary data — where any byte ≥ 0xC0 looks like a lead byte — the last one sends it past the end of the buffer. It had been reading out of bounds on every SHA-1 call for as long as digest has existed, and nothing showed: the read returned the NUL terminator and the value was discarded, since only the sequence length is used. A check that refuses is how a read like that stops being invisible.
v0.1.274 — 2026-08-16
_A string the machine cannot hold, and two backends that answered with a number._
Asking for str_repeat "ab" 500000000000000 used to produce four different things, and the two most used backends produced a plausible integer and exit 0:
| before | after | |
|---|---|---|
| interp | Fatal error: exception Out of memory, exit 2 | out of memory, exit 1 |
| C | -1530494976, exit 0 | out of memory, exit 1 |
| LLVM | 2764472320, exit 0 | out of memory, exit 1 |
| Wasm | (no output), exit 1 | out of memory, exit 1 |
Two defects compounded to make that possible.
Mere's int is 64-bit; the runtimes serving it were not. The count reached C through an int parameter and LLVM through an explicit trunc i64 ... to i32, so a 64-bit value arrived as its low 32 bits and asked for a string the program could actually have — a different string, returned without complaint. The same narrowing was in str_len ((int) __lang_str_size), substring's indices, str_index_of's result and utf8_len's count. A 2.15GB string — one byte of address past what a 32-bit offset can name — reported a negative length and a negative match offset. Every string in every test until now was small, so the axis had never been asked.
The allocator never read what malloc answered. When it said no, the next line wrote through the null it returned, and the program died by segfault — the nameless death v0.1.271 removed everywhere else. It is a named, catchable failure now, on both backends, and the region's doubling no longer wraps size_t on its way to a request bigger than half the address space.
On Wasm the memory is a fixed 64MB and nothing grows it, so exhaustion arrives as an out-of-bounds trap. The host used to exit 1 in silence for every trap that was not fail; it now names that one out of memory and prints the engine's own words for anything else, so no trap is anonymous.
Gated two ways. test/parity/fail/uncaught_out_of_memory.mere holds all four backends to the same sentence on a request no allocator can satisfy — it is refused instantly, so it costs nothing. scripts/bigstr_check.sh is the 2.15GB measurement, deliberately not in the parity run: it costs 4.3GB of resident memory per backend, and a gate too expensive to run is one people stop running. The Wasm backend is not in it, because a 2GB string is not a value that backend can hold — asking it would measure the memory limit rather than the width.
v0.1.273 — 2026-08-16
_A gate that cached the thing it was testing._
scripts/selfhost_check.sh compiles a set of programs with both compilers -- the OCaml one and the self-hosted Mere-in-Mere one -- and diffs what the two binaries print. Run yesterday it reported 7 failures out of 7, every case, with the self-host side empty: the reading a person would take from that is that the self-hosted compiler is completely broken.
It was fine. The gate builds the self-hosted compiler into /tmp/selfmere.wasm only if that file does not already exist, and the copy sitting there was three days and one ABI change old -- built before str grew its length header. The gate was testing a compiler nobody had asked about.
CI never disagreed. A fresh runner starts with an empty /tmp, so CI always built the current compiler and always passed. The same commit was green on the machine nobody looks at and red on the machine someone is working on, and the red one was the wrong answer.
It builds every run now. The whole build is 260ms -- there was never enough here to cache. A sweep of the other gates for the same shape (reuse an artifact if present) finds none.
The failure report is also bounded now. When this failed it wrote 18MB of WAT into the terminal, and the part worth reading -- that one side was empty -- was the first line of it. Ten lines of each side and the paths, so the rest is one command away.
v0.1.272 — 2026-08-16
_Q-032 closed: the Wasm backend learned to unwind, and the last pin came down._
fail on this backend has nothing to unwind with. It sets a flag, returns a sentinel, and the try_or at the boundary reads the flag -- so the failure was caught, but everything between the fail and the catch still ran. A body that pushed three strings and failed after the second pushed the third here and nowhere else. That one line was the parity suite's last DIVERGE pin, standing since v0.1.246.
Unwinding, written out by hand: after a call, ask whether the callee failed and return at once if it did. The check goes in at emit_instr -- the one place every call passes through -- rather than at each emission site, which is how the sites that were forgotten in earlier sweeps would have been forgotten again. A return_call is exempt: it is a tail call, the frame is already gone. So is try_or's own call to the thunk, which is the frame that must not propagate.
The interesting half is the second question: which calls need to ask. A callee that cannot reach $__lang_fail cannot have set the flag, and the assembled module can prove that where the emitter could not -- so a pass over the finished module computes reachability and takes those checks back out. On a self-hosted compile that is 2,524 of 3,722 call sites:
| module | vs no unwinding | |
|---|---|---|
| no unwinding (before) | 241,915 B | — |
| a check after every call | 275,273 B | +13.8% |
| dead checks pruned | 253,067 B | +4.6% |
Everything unknowable keeps its check: indirect calls (the callee is a table index), imported functions (the host can re-enter the module), and any callee without a definition in the module. Removing a needed check would be a silent wrong answer, so every doubt resolves toward keeping it.
scripts/parity.sh also learned to fail on a stale pin. A pinned divergence that starts matching used to pass quietly, because a matching case never reads its .expected file -- the declaration would have stayed on disk saying something that had stopped being true. That is the exact failure the pin mechanism exists to prevent, so the gate now names the file and asks for it to be deleted. It is how this slice found out it was done.
v0.1.271 — 2026-08-16
_The failure that had no name._
Recursion deeper than the stack is the most common way a Mere program actually dies -- three of this week's findings were exactly that -- and until this slice it was the one failure the language never said anything about. Four backends gave four answers, and the two most used gave none:
| before | after | |
|---|---|---|
| interp | Fatal error: exception Stack overflow, exit 2 | stack overflow (recursion too deep), exit 1 |
| C | (nothing), exit 139 | same sentence, exit 1 |
| LLVM | (nothing), exit 139 | same sentence, exit 1 |
| Wasm | node's RangeError + hundreds of trace frames, exit 1 | same sentence, exit 1 |
The compiled backends carry a SIGSEGV/SIGBUS handler that runs on a stack of its own -- the stack that just overflowed has no room left to run a handler -- and it claims a stack overflow only when the faulting address is near the stack. Anything else keeps its own name: a segfault from some other cause is still reported as a segfault, because a diagnostic that guesses is worse than one that does not exist. The bounds are read from the thread rather than assumed, which is what makes the answer right for a program linked with a bigger stack -- a 512MB one, as the Ruby subset uses, overflows far below any 8MB guess and would otherwise have been misnamed.
The interpreter's limit is declared, not discovered. OCaml 5 grows the main fibre's stack by copying it, so finding the host's real ceiling costs 68 seconds and gigabytes of copying, and the depth it finds -- around forty million frames -- is two orders of magnitude past anything a compiled backend can reach. A program that recurses that far has already failed everywhere else. The interpreter now stops at 1,000,000 frames (MERE_MAX_DEPTH to move it) and says the same sentence in about a second. The count comes back down with the stack it was counting, so a program that catches a failure inside a loop does not drift upward into a depth it is not at.
test/parity/fail/uncaught_stack_overflow.mere is where the agreement is checked rather than claimed: 8 failing-program cases now, all four backends matching on exit status, prior output and message.
This also corrects v0.1.270's closing paragraph, which named the region as the cause of a crash nobody had measured. It was the stack.
v0.1.270 — 2026-08-16
_The other helper that could not survive a long list, and a sweep that says there is no third._
list_filter was fixed two slices ago for rebuilding its result through the return path. The obvious next question is whether it was the only one, and asking it the lazy way — grep the prelude for a self-call wrapped in a constructor — turns up exactly one more: `list_sort_insert`, which walks to the insertion point through Cons (h, list_sort_insert cmp t x) and so recurses once per element it passes. Inserting into a 50,000-long sorted list overflowed the stack on the compiled backends.
It accumulates and reverses now. The same sweep over the whole prelude afterwards finds no remaining wrapped self-call: list_map, take, concat, flat_map, zip, append, the merge-sort quartet and the utf8 helpers were all already in that shape, and the two that were not are both fixed.
Worth separating from the depth question: each list helper handles a 50,000-element list on its own, and a program that builds several such lists at once stops earlier on the compiled backends. That is the region running out, not the stack -- a different axis, and one this measurement deliberately does not mix in.
Corrected in v0.1.271. That last paragraph is a guess written as a finding: the crash was never measured, only named. It is the stack. The same program runs to completion under ulimit -s 65520, and under v0.1.271 it says so itself. The region was not involved.
v0.1.269 — 2026-08-16
_A match arm is a tail position too, which is where nearly every recursive list helper's tail call actually lives._
v0.1.267 gave the LLVM backend a notion of tail position and emitted musttail, and the measurement that immediately followed it — the collection value axis — said list_sum over a list still died at 100,000 there while C was fine at 200,000. The reason was narrow: only `If` had been made tail-aware. Every recursive helper in the prelude is written with match, so its tail call sat in an arm that branched to a join and phi'd — the one shape musttail cannot take.
An arm in tail position returns now, exactly as a tail-position If branch does. list_sum runs at 500,000 on LLVM, and a program that puts a 300,000-element list through every list helper — len, sum, rev, map, filter, fold — gives the same answers on the interpreter, C and LLVM.
The lesson is about where the measurement pointed. "Tail position" sounded like one concept and was implemented as one construct; the gate written an hour later found the other construct by asking a question that had nothing to do with tail calls.
v0.1.268 — 2026-08-16
_A missing map key is catchable everywhere, and list_filter survives a list longer than the stack._
The third value-axis parity case — after the integer widths and the string edges — asks about collections: the empty ones, the single-element ones, a key that is not there, and a list long enough to leave the small cases behind. It found two.
A missing map key was not catchable on any compiled backend. map_get wrote its own diagnostic and then abort()ed (C, LLVM) or executed unreachable (Wasm), so try_or (fn () -> map_get m k) d answered d on the interpreter and died with 134 or 1 everywhere else. It goes through the same failure path every other fail uses now: caught, it returns the default; uncaught, all four print the same sentence and exit 1.
`list_filter` recursed to the length of the list. Every other list helper in the prelude accumulates and reverses — list_map, list_take, list_concat, list_zip, list_append all do — and this one built Cons (h, list_filter t p), so filtering 100,000 elements overflowed the stack on a backend where mapping them did not. It accumulates now.
The case is sized at thirty thousand deliberately. Past that the axis stops being about values and becomes about stack depth, which differs per backend and per operation — list_sum alone survives 80,000 on LLVM and dies at 100,000 while C is fine at 200,000. A value gate that also measured the stack would report the wrong thing whenever either moved.
v0.1.267 — 2026-08-16
_The LLVM backend's tail calls are musttail, so a loop written as recursion runs in constant stack there too._
The C backend learned this in v0.1.230 — self tail calls became a goto — and the same measurement was left standing against LLVM: ten million iterations died with SIGSEGV at `-O0` while C ran them in constant space, and the emitted IR carried no tail marker at all. The note recorded it as measured rather than assumed, and said what the fix would cost: musttail is the marker LLVM guarantees regardless of optimisation level, but it requires the call to be immediately followed by a `ret` of its result, and this backend had no notion of tail position.
It has one now. emit_expr carries the same tail-position flag the C backend uses, cleared at entry and restored only for the sub-expression that stays in tail position. An If in tail position returns from each branch instead of joining through a phi, which is what puts a tail call next to its ret. A call in tail position whose prototype matches the enclosing function's is emitted as musttail.
Ten million iterations now run at -O0, and so does mutual recursion between two functions. MERE_NO_TAIL_CALL=1 turns the transform off, the same escape hatch MERE_NO_TAIL_LOOP gives the C side — "is it this change?" stays a one-variable question.
v0.1.266 — 2026-08-16
_The string values four backends were never asked about, and the two answers that were wrong._
A new parity case walks the axes the open question about builtin parity names next to the width one it already closed: the empty string, an index that is not inside the string, and a string big enough to leave the small cases behind. Every string probe in the suite until now used a literal a human types in the middle of the range, so a helper that divides by a length of zero, or reads one byte past the end, answered every question it was asked.
It found two, both on LLVM and both from this week's header migration:
- `str_repeat s 0` allocated its empty result directly, without a header — the
one case that computes no length, and so the one the mechanical conversion missed. str_len of it read whatever preceded the allocation.
- `str_trim` walked a pointer into the string and then asked for its length.
An interior pointer has no header of its own. This is the same bug the Wasm backend had in v0.1.262 — invisible on LLVM until today, because there was no header there to read wrongly.
The second one is the interesting one: the same defect existed in two backends, written years apart, and the gate that found the first could not see the second until the representation changed underneath it.
v0.1.265 — 2026-08-16
_What fail hands back on Wasm is an empty str, so the code it cannot unwind past survives to reach the try_or._
Wasm has no unwinding here: fail sets a flag and returns, and the callers check the flag on the way out. The value it returned was 0 — which is not an address. A consumer sitting between the fail and the try_or that treated it as a str read the length header at -4 and trapped, so try_or (fn () -> str_len (fail "b")) (-1) died with no output where the interpreter, C and LLVM all answered -1. Which of the two happened depended on what the consumer did with the value, which is the worst property a failure can have.
The sentinel is an interned empty str now. An empty str is a valid answer to every string operation, so the code between the fail and the catch survives, and the try_or default comes back on all four backends.
This does not make fail unwind, and the parity case still pins the difference that remains: statements after a fail in the same body still run, so a buffer that three pushes wrote to reads "onetwothree" there and "onetwo" everywhere else. That is the open half of the question, and it is where the pin now points.
v0.1.264 — 2026-08-16
_The LLVM backend's str carries its length, and the last pin comes off._
Its str was a bare pointer: str_len called strlen, == called strcmp, and literals were plain [N x i8] globals. So on that backend a zero byte was not a byte — str_len (chr 0) answered 0 against 1 everywhere else, and chr 0 == "" was true. The open question about it named the consumers: percent-decoding a URL and decoding Shift_JIS both produce arbitrary bytes, and LLVM alone gave a different answer for them.
It now lays a str out the way the other three do — [i64 len][bytes][NUL], value at byte0. Literals are { i64, [N x i8] } constants and the value is a constant getelementptr into the second field; every runtime helper that builds a string allocates through __lang_str_alloc; the ones that only learn their length at the end (replace, trim, unescape) call __lang_str_finish to write it. Strings that arrive from libc — asprintf for show, snprintf for floats — are copied into a header by __lang_str_of_cstr at the boundary. str_len, ==, str_compare and print all read the header, so a NUL is a byte on all four backends.
The trailing NUL stays. Everything else in this runtime still hands pointers to libc, and keeping the terminator is what let the migration be incremental rather than a rewrite.
Two things repeated from the Wasm slice a day earlier, which is worth writing down: the allocator had to be emitted unconditionally (it lived with the concat helper, so a program that only showed a value referred to a function that was not there), and every internal message global that gets concatenated — the fail prefix, the int_of_str message, the show constants — needed a header of its own.
nul_in_str has no pinned divergence left.
v0.1.263 — 2026-08-16
_The self-hosted Wasm backend carries the length header too, so the host can stop guessing._
The previous slice found that print on Wasm wrote up to the first NUL and could not do otherwise: the JS host runs modules from both compilers, and while the OCaml backend lays a str out as [i32 len][bytes][NUL], the self-hosted one still emitted the pre-header representation — a bare NUL-terminated buffer — while stamping the same ABI number as the compiler that carries a header. A number that says "you may read the length at ptr-4" is worth nothing if half the modules do not have one.
So the self-hosted backend lays strings out the same way now: literals carry a four-byte header in the data section, $__lang_strlen reads it instead of scanning, and every helper that builds a string — concat, substring, repeat, char_at, chr, strbuf_to_str, unescape, show_int/bool/str — allocates through one place that writes it. The digit buffer show_int fills right-to-left is copied into a str that has room for a header in front, the same answer the OCaml backend reached.
With both compilers agreeing, the host reads by length, and print on Wasm writes the whole value: the nul_in_str parity case needed a Wasm pin for exactly one slice. LLVM keeps its pin — its str is a bare pointer with no header anywhere.
Two things fell out of doing it. $__lang_str_alloc had to be emitted unconditionally rather than with the length helper: a module that showed a bool without measuring a string referred to a function it did not define. And the emitted module grew 17 bytes, which the size guard on the self-hosted codegen reports — four bytes per literal is the cost of a length that does not have to be searched for.
v0.1.262 — 2026-08-15
_An interior pointer has no header, and str_trim asked one for its length._
Chasing the Wasm half of the previous slice found something smaller and real. str_trim skips leading whitespace by walking a pointer INTO the string, and then called $__lang_strlen on that pointer to find out how much was left. The length of a Mere str lives in a header immediately before byte0 — so for an interior pointer that read takes the string's own bytes as a length. It is bounded by a NUL scan on the other backends, which is why only Wasm blew up on it, and why str_len (str_trim s) was the shape that showed it rather than print (str_trim s).
It takes the original length once now, uses it to bound the skip, and subtracts what the skip consumed. A NUL inside the value stays a byte through str_trim, on all four backends.
The Wasm host still writes up to the first NUL, and the reason turned out to be sharper than "something overstates its length": the self-hosted Wasm backend emits the pre-header string representation — its $__lang_strlen scans for a NUL and its strings carry no header — while stamping the same ABI number as the compiler that does. The host is shared between both kinds of module, so it cannot trust the header until the self-hosted backend carries one, or stops claiming the ABI that says it does. That is written down where the parity case pins it.
v0.1.261 — 2026-08-15
_print writes the string's length, and the two backends that cannot are pinned._
str became byte-safe in the v0.1.129 arc: the length lives in a header rather than in a terminator, and str_len (chr 0 ++ "X") has answered 2 ever since. print did not — it went out through puts, which stops at the first NUL — so the same value measured 2 and printed one character. Two answers about one value is a bug, not a choice, which is what the open question about this asked to settle.
The C backend writes by length now (print, print_err, print_no_nl), and matches the interpreter byte for byte.
The other two are pinned rather than fixed, each for its own reason, in a new parity case that prints NUL-carrying strings as well as measuring them:
- LLVM has no length header at all — its
stris a bare pointer, and
str_len (chr 0) is 0 there against 1 everywhere else. The NUL is not truncated on output; it was never in the value.
- Wasm has the header and
str_lenreads it, but the JS host writes what it
finds up to the first NUL. Making the host trust the header instead surfaced a second thing: one str on the self-hosted compiler's path arrives with a header that overstates its content by half a megabyte of zeros. Until that is understood the host keeps scanning, and the pinned case is where it will be noticed.
Pinning is the point. Dropping U+0000 from the case — which is what happened the first time this came up — would have made the file agree by not asking.
v0.1.260 — 2026-08-15
_Exponent notation, because the ends of the double range could not be written down._
1.7976931348623157e308 lexed as the float 1.7976931348623157 applied to a variable named e308. A probe that needed the largest finite double, the smallest normal and the smallest subnormal had to build all three out of powers of two — scaling by halves until the value stopped changing — rather than write them.
Now a literal with an exponent is a float, with or without a decimal point: 1e3, 2.0e-3, 4E+5. A digit has to follow the e (after an optional sign), so 1.5 e is still a float applied to a variable called e, which is the only thing this could have taken away.
The formatter had to change with it. string_of_float keeps 12 significant digits, which was enough for every literal that could be written before and is not enough now: formatting 1.7976931348623157e308 would have written a different number back. It emits the shortest form that reads back as the same double.
And then the notation paid for itself immediately. A new parity case walks the float values four backends had never been asked about — the ends of the range, the subnormals below the smallest normal, both zeros, NaN — none of which could be reached before, because every float literal in the suite was one a human types. It found a real divergence on the first run: `nan < 0.0` was true on the interpreter and false on all three compiled backends. The interpreter routed every ordered comparison through OCaml's total compare, which sorts NaN below everything; the operator is IEEE, where every ordered comparison with NaN is false. Sorting still uses the total order — that part was deliberate — but < on two floats is now the float comparison, and the four backends agree.
v0.1.257 — 2026-08-14
_The tokenizer switches its own state after a start tag, which is a shortcut with a stated boundary rather than a guess._
<title>t</title> used to tokenize as a start tag, a bare t, and an end tag, because the content of a title is RCDATA and the standard has the tree builder decide that. A tokenizer handed a whole document cannot ask.
For HTML the decision is a pure function of the tag name, so the tokenizer makes it: title and textarea go to RCDATA, style and its relatives to RAWTEXT, script to script data, plaintext to PLAINTEXT. The exception is foreign content — <title> inside SVG is an ordinary element — and until this backend knows about foreign content the two answers are the same. That is why this is a shortcut and not a mistake: the boundary is known and written down where it is taken.
The conformance suite still passes 1,900 of 1,900, and the browser dogfood's tree construction went from 83 of 189 to 88.
v0.1.256 — 2026-08-14
_All 228 labels the Encoding Standard defines, generated, replacing a hand-written list of the four this directory can decode._
The previous slice added utf-16 to that list. This is the same gap at its full size: a page declaring iso-8859-2 read as a page declaring nothing at all, because the list only had labels for encodings there is a decoder for.
Those are two different questions. label_of answers what a label names; whether this directory can decode the result is what decode answers, and it already answers None. Conflating them made the absence of a decoder look like the absence of a declaration, which the HTML standard treats differently — it is the difference between "use the default" and "use the encoding the page named".
The table is generated from encodings.json and carried the way the Unicode tables and the HTML named references are (Q-028): fixed-width records in one sorted string, binary-searched by index arithmetic. 228 labels for 40 encodings, 8,208 characters.
The comment above the old list had already written down what would happen — "it cannot discover a label we forgot to list, and that gap is real and is recorded rather than implied". It took a program that asks about labels rather than about decoding to walk into it.
v0.1.255 — 2026-08-14
_A label the table forgot, found by the program that asks about labels._
label_of had no utf-16, utf-16le or utf-16be. The comment above that table already said what would happen — "it cannot discover a label we forgot to list, and that gap is real and is recorded rather than implied" — and this is the gap, found by the browser dogfood asking a page what encoding it claims to be in.
Listing them is right even though nothing here decodes UTF-16: label_of answers "what does this label name", which is a different question from "can this directory decode it", and decode already returns None for a name it does not implement. Leaving them out made <meta charset=utf-16> look like a page with no declaration at all — and the HTML standard answers that differently from a page declaring UTF-16, which cannot be true (the declaration is written in ASCII) and reads as UTF-8.
v0.1.254 — 2026-08-14
_1,900 of 1,900. The vendored tokenizer suite passes entirely, with no exemptions._
Three fixes, and two of them were the same mistake in different clothes.
Form feed was not whitespace. _is_space compared against a \012 escape the lexer did not read as U+000C — so the literal was four ordinary characters and form feed silently stopped being a space character everywhere the tokenizer looked for one. It is chr 12 now: whether the language spells it \f, \014 or \x0c is a question this file does not need an opinion about, and getting it wrong is silent. Twelve cases.
The bogus-doctype state was inventing quirks mode. It hardcoded the force-quirks flag rather than carrying the one the token already had, so <!DOCTYPE a PUBLIC'''' followed by junk — a complete, correct doctype with a parse error after it — was reported as a quirks-mode document. Reaching a recovery state does not by itself mean the thing recovered from was fatal. Sixty-two cases, and the same shape as the fix two slices ago: the recovery path was discarding what it had rather than keeping it.
And a NUL starting an attribute name went through unreplaced, because that one path appended the character without the substitution every other one had. Four cases.
v0.1.253 — 2026-08-14
_Character references, named and numeric: 1,807 of 1,900 becomes 1,822 — and the last exemption bucket comes out of the harness._
The 2,231 named references are one string of fixed-width records, generated from the standard's own entities.json, sorted, and binary-searched by index arithmetic. That is the shape Q-028 settled for the Unicode tables, applied again: 44 characters per record — the name space-padded to 32, then two code points as six hex digits each. Space pads rather than NUL because a Mere str cannot carry a NUL through the compiled backends, and because space sorts below every character a name uses, so the padded order is the plain order and a short name needs no special case. 98,164 characters, and it compiles and runs on the C backend as well as the interpreter.
Matching is longest-first, because ∉ is one reference and not ¬ followed by in; — a search that stopped at the first match would be a different tokenizer. Inside an attribute value a match that did not end in ; is left alone when the next character is = or alphanumeric, because ?a¬=b is a query string far more often than it is a negation sign.
And the harness lost its last exemption. It had a bucket that excused cases needing character references from the failure count, back when there were none. It came out the moment they were implemented: a bucket that exists because a feature is missing hides real failures as soon as the feature arrives. Every one of the 1,900 cases is now a pass or a failure, and the 78 remaining are named one by one.
v0.1.252 — 2026-08-14
_The tokenizer's other starting states: 1,598 of 1,704 becomes 1,807 of 1,900._
`Html.tokenize_in` takes the state to start in and the tag that opened the element. An element whose content model is text — title, textarea, style, script — puts the tokenizer in RCDATA, RAWTEXT or script data, and only the tag that opened it can end it. That is why the starting state is a parameter and not a guess: the tree builder is what knows which element it is inside, and a tokenizer that guessed would be wrong exactly where </ appears inside a script.
The three text-like states share one shape and one implementation: character data until </ plus the opening tag's name plus a space, a slash or >. PLAINTEXT is the degenerate case that never ends.
And the case count went up, which is the point. A case listed under several initial states is several cases — the suite writes it once and means it for each. Running only the first reported a number smaller than what was being checked, so the harness expands them: 1,704 entries are 1,900 cases.
v0.1.251 — 2026-08-14
_Three more of the tokenizer's rules: 1,505 of 1,704 becomes 1,598. Sixty-two of those came from one line._
Newline preprocessing — CRLF and a lone CR both become LF before the machine sees them. The standard puts this in a preprocessing step rather than in the states for the reason it shows here: otherwise every state has to say it.
U+0000 becomes U+FFFD inside markup — names, attribute values, comments, doctypes. It is a parse error and the character is replaced rather than dropped, because dropping it changes how many characters a later consumer counts.
And the one that was worth sixty-two cases: the bogus-doctype state was throwing away what had already been parsed. <!DOCTYPE a PUBLIC"" followed by junk has a public identifier that is present and empty; discarding the state on the way into the recovery path reported it as missing, which is a different token. Recovery paths keep what they have — that is what makes them recovery rather than restart.
v0.1.250 — 2026-08-14
_An HTML tokenizer, measured against the standard's own suite from the first run: 1,505 of 1,704._
contrib/html/tokenizer.mere the state machine, with the standard's state names
test/data/html5lib/*.test vendored html5lib-tests (tokenizer)
scripts/gen_html5lib_testdata.sh how they got here (maintenance, needs network)
scripts/html_tokenizer_conformance.sh the gate
Written as the standard writes it: one function per named state, so a line can be found in the specification by searching for its state. Not a regular expression or a lookahead scanner, because the recovery rules are what make HTML parseable at all and they are stated per state — <, </, <! and <? all have defined behaviour when what follows them is not what it looked like, and that is where the bugs are. Two of the first three the suite found were exactly that shape: a repeated attribute keeping the last instead of the first, and a comment ending in a lone dash keeping it instead of dropping it.
The pass count is pinned exactly, not as a floor. A floor lets a regression hide behind a new pass. The harness prints the first ten failures in full, so what is missing is in the output rather than only in a document — and what is not covered yet is counted by category rather than skipped silently: character references (53), non-Data initial states (111), U+0000 replacement (57).
The suite is vendored rather than fetched by the gate, the same decision the UCD conformance files got: a gate that needs the network fails for reasons that have nothing to do with the code.
Placement: a conformant tokenizer is a library — the same shape as contrib/url and contrib/encoding — so it lives here. Tree construction is where a browser starts and is not.
v0.1.249 — 2026-08-14
_A window, its pixels, and its input — promoted from a probe to a capability. The interesting part is that it can be checked without anybody looking at a screen._
lib/codegen_c.ml the SDL2 runtime, emitted when a win_* extern is declared
contrib/window/window.mere the typed side: window, event, show, capture, poll
test/window/window_check.mere draw, show, read back, compare
scripts/window_check.sh new gate: SDL's dummy driver, no display needed
Six externs and no language feature. win_open / win_size / win_blit / win_readback / win_poll / win_close. Pixels cross the boundary as a flat arena offset and everything else is an int — the contract the socket family established — so the runtime is conditional C the way PortMidi's is, emitted only when a program declares one of these. Build with sdl2-config --cflags --libs.
`size` asks the renderer, not the window. They are different numbers on a HiDPI display: a 640×480 window has a 1280×960 renderer, and the pixels are in the second one. This was recorded as friction 3 when the capability was a probe — a readback comparison against the window's size compares against a number the pixels are not in — so the capability returns SDL_GetRendererOutputSize and nothing else.
`Window.show` composites a `canvas` with the same arithmetic as `Canvas.to_ppm`, so what a program puts on the screen and what it writes to a file are the same image by construction rather than by two similar loops.
The gate is a readback, and it needed one more thing to be evidence. Draw a known pattern with contrib/raster, show it, read the window's pixels back, compare — 3072 pixels, 0 mismatches, under SDL's dummy video driver, which has a software renderer and a real event queue but no display. That runs in CI and does not open a window on your desktop.
But `show` writes the image into the same arena block `capture` reads back into, so a readback that did nothing at all would hand back exactly what was written and every pixel would match — a gate that passes while testing nothing. capture poisons the block first. Checked by making the runtime's readback a no-op: 3072 of 3072 pixels then differ.
Not gated: that an event ever arrives. poll is checked only for answering Nothing on an empty queue. Delivering a real key or click needs either a display or a way to inject one, and a test-only extern that pushes events would be checking the scaffolding rather than the capability.
v0.1.248 — 2026-08-14
_Adding the last two math builtins turned up a silent wrong answer in every LLVM program that used floats: f_pow 3.0 2.0 was 3.0, because the prelude's integer pow had taken libm's symbol._
lib/codegen_llvm.ml Mere top-level names are prefixed `mu_`, as C's always were
lib/codegen_c.ml exp / log beside sqrt
lib/codegen_wasm.ml the same, as host imports
scripts/run_wasm.js + 5 __lang_exp / __lang_log; and str_of_float follows C's %g
test/parity/exp_log.mere new: identities within a tolerance, not digits
`exp` and `log` were the last of the family that began with floor / ceil / round in v0.1.243: names in the typer's environment with a type, so a program using them type-checked everywhere and then mere -c emitted a call to a symbol the C compiler had never heard of, while LLVM and Wasm said "unbound variable". They are three lines on each backend. MISSING 8 → 6, nocompile 10 → 8.
Then the new test failed on LLVM, for a reason that had nothing to do with it: `f_pow 3.0 2.0` came back `3.0`. This backend emitted Mere top-level names into the IR's global namespace unprefixed, so when the prelude grew an integer pow in v0.1.245 that became define @pow — which is libm's symbol, and @llvm.pow.f64 lowers to a call to pow. Every f_pow on this backend has been calling the integer power since. The C backend has prefixed with mu_ since it was written; this one now does too.
Neither internal linkage nor renaming the intrinsic helps — both were tried and measured. The collision is the name, and the name was in a namespace shared with the C library: write, exit, time, free and every other libc symbol were the same accident waiting for a program to name a function after one.
The rename is the loud kind of change: a site missed by the prefix fails at link time with an undefined symbol rather than computing something else. Four such sites turned up and all four were tables keyed by the source name — the free-variable analysis, the lifting pass's host, the shadowing guard's position lookup, and the debug info's DISubprogram(name:), which shows in a debugger and must stay what the program calls it. Emitted names are for the IR; source names are for everything that reasons about the program.
And the Wasm host printed floats by a different rule. str_of_float there emulated C's %g with JavaScript's toPrecision, which is not that rule: %g goes exponential when the decimal exponent is below -4 and toPrecision stays decimal down to 1e-7, so exp -10 printed as 0.00004539992976248485 on Wasm and 4.5399929762484854e-05 everywhere else. The host implements the actual rule now, exponent padded to two digits as C does.
What the test asserts, and why it is not digits. A transcendental function is not required to be correctly rounded by anybody: exp -10 differs in the last bit between libm and JavaScript, and a gate comparing the digits would be reporting the C library's build options. So exp_log.mere prints exact values only where the answer is exact in binary floating point and asserts everything else as an identity within a tolerance — log (exp x) = x, exp (2 log 3) = f_pow 3 2, exp (0.5 log 2) = sqrt 2. That still fails an exp that returns its argument or a log wired to log10, both checked by reverting the fix and watching the gate go red.
v0.1.247 — 2026-08-14
_x / 0 was four different things, and three of them were not failures. The gate built last slice is what made fixing it a two-line test._
lib/codegen_c.ml __lang_idiv / __lang_imod: a checked divisor
lib/codegen_llvm.ml the same, as IR functions
lib/codegen_wasm.ml the same, with the message interned per program
test/parity/fail/uncaught_div_zero new
test/parity/fail/uncaught_mod_zero new
scripts/parity.sh the message is the first line of stderr, not the last
What it did before. The interpreter raised division by zero. The C backend emitted a bare a / b, which is undefined behaviour in C: this machine's arm64 quietly answered 0 and the program carried on printing, while an x86-64 build of the same source raises SIGFPE. LLVM emitted sdiv, undefined in IR and therefore something the optimizer may assume never happens. Wasm trapped — a defined failure, but a silent one, with no message at all. A wrong answer, a crash, or a silent death, depending on the backend and the CPU.
All four raise now, with the interpreter's messages (division by zero and modulo by zero — it distinguishes them, so the others do too), catchable with try_or. It costs a branch per division. INT_MIN / -1 is the other undefined case in C and IR and wraps now, matching what the interpreter already produced.
The `-rv` backend is the exception, and it was measured rather than assumed: built for QEMU's virt board and run there, 17 / 0 is -1 and 17 % 0 is 17, the RISC-V specification's non-trapping answer. That backend targets bare metal, where there is no stream to write a diagnostic to and no process to exit, so the platform's answer is the answer. Making it raise would mean deciding what a machine-mode trap means for the kernel that runs on it, which is its own piece of work.
And the harness needed one more fix to see any of it. Its notion of "the message" was the last line of stderr, which works for a fail call — that carries no location and renders on one line — and not for a failure the interpreter can point at, which renders with --> file:line and the source under it. The message is the first line. A single-sink backend is the mirror image: there the diagnostic is the last line of the program's own output. Two ends, because the two files are different files.
v0.1.246 — 2026-08-14
_The parity harness compared stdout, so the failure surface of the language was the one part of it four independent implementations were never held to. No parity test used fail — and that was not an oversight: none could have passed._
scripts/parity.sh failing programs compared on exit + stdout + message; DIVERGE
test/parity/fail/*.mere new: five uncaught failures, four backends
test/parity/failure_caught.mere new: try_or over every failure kind
lib/codegen_c.ml exit 1, not abort; the `fail: ` tag moves to the builtin
lib/codegen_llvm.ml a stderr at last: diagnostics, print_err, print_no_nl
lib/codegen_wasm.ml the tag; int_of_str names its input; print_err refuses
What an uncaught failure did, before this slice. The same program exited 1 on two backends and 134 (SIGABRT) on two others. It wrote its diagnostic to stderr on two and stdout on two. It tagged the message fail: on three and not on the fourth. And int_of_str "abc" named the offending input on two backends and said only int_of_str: not a valid int on the other two — the same failing program telling you two different things, and the version that omits the input is the one you cannot debug from. Five differences, none of which any test could see.
All four now write one line to stderr and exit 1, with the message the program raised. LLVM's panic path had been using puts — a backend that refused print_err for having no stderr lowering was writing its own diagnostic to stdout. write(2, ...) was already declared for print_bytes; declaring it unconditionally made the panic path correct and made print_err and print_no_nl three lines each, so both stopped being refusals on that backend.
The `fail: ` tag belongs to the `fail` builtin, not to the printer. Tagging where the diagnostic is written tagged the backend's own failures too, which the interpreter does not: hence fail: int_of_str: ... against int_of_str: .... Moving it to the builtin makes the message comparable verbatim, which matters more than it sounds — see below.
On Wasm, `print_err` now refuses. It wrote to the same host sink as print, so a diagnostic landed in the program's own output and nothing said so. The JS host ABI has one sink (env.puts); giving it a second one is a change to every host that instantiates a module, which is a deliberate change and not a side effect of a panic message. Nothing in the repo used print_err, so there was nothing to break — and a refusal names the missing thing where a silent stdout write named nothing.
`test/parity/fail/*.mere` is the gate: programs that are supposed to fail, compared on exit status, the stdout written before the failure, and the message. Five of them. The harness takes only the interpreter's envelope off the message (it names the source file it is running; a compiled binary has none) and compares the rest byte for byte.
The first version of that comparison stripped the `fail: ` tag as well, and it made the harness unable to see the difference this slice had just fixed: boom and fail: boom both normalized to boom, so removing the fix still passed. The control experiment caught it — revert a fix, confirm the gate goes red. A normalization is a place a gate stops looking, and the fix was to make the thing consistent by construction instead of normalizing it away.
And a DIVERGE state, because the caught-failure test found a live one. fail on Wasm sets a flag that callers check on the way out; it does not unwind, so statements after it in the same body still run. Inside a try_or thunk that is an observably different result, not just extra output — a buffer that reads onetwo on three backends and onetwothree on the fourth. A gate with no place to say "these two legitimately differ here" loses the first real divergence it finds, along with everything else that test was checking. So a divergence is declared by a file next to the case (failure_caught.wasm.expected) holding that backend's output exactly: the known difference is pinned rather than tolerated, any other change is still a failure, and the day that backend learns to unwind, the declaration breaks and says so.
Two smaller things found on the way. Passing a test/parity/fail/ file as an explicit argument ran it as an ordinary test and reported it as one the interpreter could not run — arguments are partitioned by path now, the same way the defaults are. And __lang_str_concat("fail: ", msg) in the C backend hung the program instead of printing: this backend's strings carry a length header before byte 0, a raw C literal has none, so the concat read whatever preceded the constant as its length. Same trap that made str_replace return "" in v0.1.233.
v0.1.245 — 2026-08-14
_Ten builtins that existed only on the interpreter became ten definitions in the language. Moving them exposed three bugs that had nothing to do with them, two of which were wrong answers from builtins every test called correct._
lib/prelude_stdlib.ml sign incr decr square cube sum_range pow lcm divmod assert
lib/eval.ml -105 lines: the ten builtins those replace
lib/codegen_llvm.ml string globals minted once; int_of_str returns i64
lib/codegen_c.ml gcd, random_int, file_size widened to long long
test/parity/int_width.mere new: every int builtin, with arguments above 2^31
test/parity/show_json_same_program new: show and to_json in one program
The ten. sign / incr / decr / square / cube / sum_range / pow / lcm / divmod / assert were in the typer's environment and in eval.ml and nowhere else: they type-checked on every backend and mere -c then emitted a reference to a name the C compiler had never heard of. Defining them as Mere source in the prelude gives all five backends the same one from one place, which is cheaper than five codegen cases and cannot drift between them. int_max / int_min are deliberately not among them — they cannot be a portable literal while int is 63-bit on interp and 64-bit elsewhere.
"The same as the builtin" was wrong four times out of ten, and only outside the range a small test looks at. sum_range and pow recursed where the builtin was closed-form and square-and-multiply, which is a stack overflow rather than a slow answer at a million terms. lcm multiplied before dividing, which overflows for operands whose lcm fits but whose product does not — on interp that made lcm 3037000493 3037000493 answer 13. And divmod left its documented zero failure to /, which is not one thing across the backends: bare `x / 0` raises on the interpreter and returns 0 on C and LLVM, so the failure would have depended on which backend you built with. It has its own check now; the divergence in / itself is still there and is not yet under any gate, because the parity harness compares stdout and does not compare failures at all.
The prelude definition shadows the builtin — measured, not assumed. Making eval.ml's sign return 999 changed no answer, so the ten builtins were unreachable rather than merely redundant, and 105 lines came out. The typer declarations stay: that is the set host_matrix.sh generates its questions from, so deleting a declaration would have removed the question rather than answered it.
`@.s_true` was defined twice. The first symptom of any of this was to_json_composite failing to compile on LLVM with redefinition of global '@.s_true'. The show emitter and the to_json emitter register their string constants in separate blocks, and they want some of the same names — s_true, s_false, s_lbracket, s_rbracket. Any program using both `show` and `to_json` hit it, which no test did until a prelude helper's error path called show and gave every program a show emitter. Minting is idempotent by name now, and a second mint with different content fails the compile instead of resolving last-wins.
Then the value that found `lcm` found a real one: `gcd 3037000493 3037000493` was `1257966803` on the C backend. The generated runtime declared static int __lang_gcd(int, int) while int is 64-bit there, so both arguments were truncated to their low 32 bits. Nothing was missing and nothing failed to compile — the builtin was recorded as present and correct on all four backends, because every probe of it had used a one-digit literal. The gap was in the values, not in the list of names.
test/parity/int_width.mere is that gate, and it found the second one while being written for the first: LLVM's int_of_str parsed with strtoll and truncated to i32, so int_of_str "3037000493" was -1257966803 and the largest int was -1 — the return width left behind when int widened to i64. random_int and file_size on C are the same shape and are widened too, though neither is observable from a parity test (one is random, the other needs a 2GB file). Sweeping both backends for the rest: every remaining narrow helper returns a status code or a count bounded by a string's length.
An intermediate has a width too. With those fixed, one line still diverged: sum_range (0 - 3037000493) 0 has an answer both int widths hold and a naive Gauss intermediate only the 64-bit one does. Halving inside the product — exactly one of the two factors is even, since an odd count means the endpoints share parity — makes the function portable over the whole range its result can hold.
v0.1.244 — 2026-08-13
_The harness whose job is to ask which backend has which builtin was asking about a set somebody remembered. Now it asks for the set too — and the answer is 144 builtins, not 50._
bin/mere.ml --dump-builtins: every name in the typer's environment, with its type
scripts/host_matrix.sh probes generated from those types; a new `nocompile` state
docs/host-matrix.md 50 rows -> 144
`mere --dump-builtins`. One line per name in Typer.initial_env, name<TAB>type. The compiler is the authority on its own environment, which is the same argument host_matrix.sh already made for the answers — it just had not been applied to the questions.
Probes are synthesized from the type: one literal per argument for int, float, str, bool and unit, and a bare mention for a non-arrow. Anything needing a value a literal cannot make — a File, a Vec, a Channel — falls back to a hand-written override, of which there are 28. 70 of the 214 names are not synthesizable and are counted and named, so the part of the environment this harness cannot see is a number rather than a silence. A synthesized probe that does not type-check is dropped and counted too.
And a new state, `nocompile`. yes used to mean "the backend emitted code", which is not the same as working: floor, ceil and round emitted fine and the C compiler then failed on an undeclared identifier. So for C the emitted source is now handed to a compiler, and the row says nocompile when that rejects it. That state is exactly the blind spot those three sat in for as long as they had been in the environment.
The matrix went from 50 builtins, 0 MISSING to 144 builtins, 18 MISSING, 20 nocompile. The old number was not wrong — it was the answer to "is there a hole among the 50 somebody listed", which reads like the answer to a different question.
Of the 20 nocompile rows, twelve are pure integer functions — cube decr divmod incr int_max int_min lcm pow sign square sum_range assert — which want defining in the prelude as Mere source, where all five backends get them at once. exp and log want libm cases beside sqrt. Neither is done here: this change is about being able to see them.
Two rows that reported error were probe artifacts rather than defects — map_new () alone leaves its key and value types unresolved, and codegen correctly refuses. Both now have overrides that pin the types, so the matrix reports 0 error.
v0.1.243 — 2026-08-13
_A rasterizer, and the three math builtins it turned out no compiled backend had._
contrib/raster/canvas.mere premultiplied pixels, one blend, rects and clips
contrib/raster/path.mere antialiased polygon fill, curves, strokes
lib/codegen_c.ml floor / ceil / round
lib/codegen_llvm.ml floor / ceil / round
lib/codegen_wasm.ml the same three, refused loudly
`contrib/raster`. A pixel buffer, source-over compositing, and one antialiased polygon fill that rectangles, glyph outlines, borders and strokes are all expressed in terms of. Nothing here opens a window, and that is the point: turning a document into pixels is checkable by comparing pixels, which needs no display.
Premultiplied alpha, so source-over is src + dst*(255-sa)/255 on every channel with no division by the result and no special case for a transparent destination. And a*b/255 is exact at both ends — the obvious (t + t/255)/255 returns 256 for 255*255, which overflows the byte into the next channel of the packed colour. The first smoke test drew an opaque black canvas instead of a white one, which is a good way for that to be found.
Coverage is separate from alpha and multiplies it, and the parity test asserts it: coverage 128 with a solid colour and coverage 255 with a half-alpha colour must land on the same pixel.
Coverage, not sampling. Each pixel row is cut into slices; per slice the edges are intersected, the crossings sorted, and the spans added to a per-pixel accumulator exactly in x — an edge at x = 3.25 puts 75% into pixel 3. Only y is quantized. Geometry is float and coverage is integer, both measured rather than assumed: doubles print identically across backends, and an integer accumulator cannot drift.
And then: `floor`, `ceil` and `round` did not exist on any compiled backend. Emission succeeded and the C compiler then failed on an undeclared mu_floor, so nothing short of a program that used them could notice — and nothing did, for as long as they had been in Typer.initial_env. sqrt, sin, cos and tan were all handled; these three were simply never added. Now on C and LLVM, verified identical to the interpreter including the negative half-way cases (round (-2.5) is -3 on all three).
The Wasm backend refuses all three, deliberately. f64.floor and f64.ceil are instructions and looked like a five-line addition — but putting the names in that backend's eta-expansion list sent it into an infinite expansion, and a bare floor call never finished emitting. round is worse than absent there: f64.nearest rounds half to even where C and the interpreter round half away from zero, and doing it properly needs a scratch f64 local, which that backend declares per function. A backend that says "no" is one a caller can work around; one that hangs, or that quietly rounds differently, is not.
A survey, since one missing builtin implies others. Fifteen names in Typer.initial_env emit as undeclared identifiers on the C backend: ceil cube decr exp floor id incr int_max lcm log pow round sign square sum_range. Three are fixed here. The rest are recorded rather than fixed, because scripts/host_matrix.sh — the harness whose entire job is to ask which backend has which builtin — has a hand-written case list and covers none of them. Generating that list from the typer's environment is the actual fix and is its own change.
contrib/raster runs on interp and C. The framebuffer is a ByteBuf, which the LLVM and Wasm backends do not have, so they refuse at emit time and the harness records UNSUP rather than a failure — C is what a native renderer targets and the interpreter is an independent second implementation, so the gate still compares two.
v0.1.242 — 2026-08-13
_A bug both hosts had, in a place the gate was only compile-checking. Backend parity could not see it, because being wrong the same way twice looks like agreement._
lib/codegen_c.ml mem_get_u32be: long long, unsigned
scripts/pg_env.js getUint32, not getInt32
lib/eval.ml the byte arena, so the interpreter can run these programs
scripts/ctest.sh compare when both sides can run, compile-check when they cannot
`mem_get_u32be` sign-extended. It returned a C int, which widens into Mere's 64-bit int with the sign — so 0xFF008080 came back as -16777088. Every opaque pixel has alpha 0xFF, so anything touching pixels hit it. It was also undefined behaviour rather than merely wrong: q[0] << 24 on an int promoted from unsigned char overflows a signed int once q[0] reaches 0x80. Now long long, computed unsigned.
The JS host had the identical bug — getInt32 where getUint32 was meant. Which is the interesting part: backend parity could not catch this, because both hosts were wrong the same way. That is the same shape as an exhaustive test file derived from the rules it tests, and it is worth naming: agreement is only evidence when the things agreeing are independent.
So what did find it? A probe that opened a window, wrote a known pattern, blitted it and read it back — when a pixel it wrote did not compare equal to the pixel it read. And what let it hide was scripts/ctest.sh: a program containing an FFI declaration was compile-checked only, on the grounds that a bare extern has no linkable symbol. True of an arbitrary name, false of the native FFI set — mem_*, tcp_*, str_ptr and the rest get static definitions emitted, so those programs link and run. Their answers were never compared.
The rule is now: compare when both sides can run, compile-check when they cannot. Requiring the interpreter to run it too is what keeps this from firing on a program that would open a socket — an extern the interpreter mocks is an extern somebody thought about.
Which needed the interpreter to have the arena at all, and now it does: the same bump allocator over a fixed buffer, the same capacity, the same first offset, so the two agree on arithmetic as well as on values. That also makes contrib/db/pg.mere runnable on the interpreter, and it is a prerequisite for the raster work: a framebuffer is an arena.
test/ctests/mem_arena_u32.mere covers the round trip at 0x7FFFFFFF, 0x80000000, 0xFF008080, 0xFFFFFFFF and the byte order, on both sides. Reverting the fix turns it red.
v0.1.241 — 2026-08-13
_The one algorithm here with both kinds of gate pointed at it — which is what makes the difference between them concrete rather than theoretical._
contrib/unicode/normalize.mere NFC and NFD
contrib/unicode/nfc_table.mere generated: ccc, decompositions, and the derived inverse
scripts/gen_normalize_tables.sh derives all three
scripts/normalize_conformance.sh 20,034 UCD cases x 6 assertions
scripts/unicode_parity.sh + 8,755 inputs vs node's String.prototype.normalize
`Normalize.nfc` and `Normalize.nfd`. é can be one code point or two, and the two spellings are the same text. Anything that compares text — an origin check, a cache key, a search — has to pick one, and a renderer that draws both spellings differently is drawing the same text two ways.
Three things carry the weight and only the first is obvious. Canonical ordering sorts each run of combining marks by class, stably, because marks of equal class must keep the order they were typed in. A decomposition is not automatically a composition — four kinds of mapping are excluded from the inverse, and the generator applies and counts all four rather than assuming them: 1,035 singletons, 4 non-starter decompositions, 81 script-specific exclusions, and 3,833 compatibility mappings that are not canonical at all. 2,081 canonical mappings in, 961 primary composites out. And blocking: a mark reaches the last starter only if nothing between them blocks it, which is why q + dot-below + dot-above composes nothing while d + dot-below + dot-above composes only the first.
Hangul is arithmetic rather than table lookup — 11,172 syllables that would otherwise be entries. Canonical mappings are stored pairwise and applied recursively, because the longest one in the UCD is two code points and a pre-expanded table would need variable-length values to buy a recursion a few levels deep.
Both gates, and the reason neither substitutes for the other. The UCD's conformance file is exhaustive in ways no independent implementation is sampled for — canonical-order permutations, PRI #29's chained composites, the closure of every composite — but it is derived from the same rules this code reads, so a shared misreading would agree with itself. node's normalize is independent but not exhaustive. So: 20,034 cases × 6 assertions against the file, and 8,755 inputs against node, the latter derived from the generated tables so a row nobody thought to test still gets one.
Six assertions per conformance line rather than two, because NFC(c1) == c2 alone would pass an implementation that is wrong about already-normalized input — which is the common case in real text.
Both green on the first run, which for once is worth saying: the previous slice's gate found three defects, and the difference is that this algorithm's hard parts (ordering, blocking) were measured against node before the gate existed rather than reasoned about.
Also: the claim in contrib/unicode/README.md that a layout engine wants East Asian Width "for advance widths" was wrong and is corrected. A renderer with a font takes advances from the font's metrics. EAW is for terminal-style layout and for a fallback when there are no metrics, which moves it down the list rather than up it.
v0.1.240 — 2026-08-13
_The first gate here that is not an independent implementation — and it caught three defects a careful reading of the rules had not._
contrib/unicode/linebreak.mere UAX #14, forty-four rules in the standard's order
contrib/unicode/lb_table.mere generated: 2,175 ranges, four UCD properties
scripts/gen_linebreak_table.sh derives the table
scripts/gen_linebreak_testdata.sh vendors the conformance suite
scripts/linebreak_conformance.sh 19,338 cases, all agreeing
`LineBreak.opportunities` says where a line is allowed to end. Not where it should — that is the layout engine's decision, made with widths — but where the text permits one.
Three things make UAX #14 long, and none of them are the rules themselves. LB9 and LB10 are a preprocessing step: a combining mark takes the class of the character before it, so the unit the rules see is a base plus its trailing CM/ZWJ run, except after a hard break or a space where LB10 makes the leftover an AL. "even after spaces" appears in six rules, each needing the last non-space class as well as the immediately preceding one, all six sitting before the rule that breaks after a space. And some rules look further than one character either way — LB25 needs a number state and two of lookahead, five rules need what came before the previous character, LB30a needs a count.
Four UCD properties in the table, because several rules are written in terms of things other than Line_Break: LB15a/15b test General_Category Pi and Pf, LB19a and LB30 test East_Asian_Width, LB30b tests an unassigned Extended_Pictographic. LB1's resolution happens in the generator, which lets the rules read the way the standard writes them.
The gate is a different kind, and the difference cuts both ways. UAX #14 has no oracle in node: Intl.Segmenter has no line granularity and Intl.v8BreakIterator is gone. So this runs the Unicode Consortium's own conformance file instead — weaker, because it is derived from the same rules the implementation reads and a shared misreading of the prose would agree with itself; stronger, because it is exhaustive over the pair table, every class against every class with and without an intervening combining mark and space, which no hand-written corpus would reach. It is vendored under test/data so it runs offline and cannot drift from the table's version.
It earned its keep immediately. Three defects survived a careful reading of the rules:
positions were reported per unit rather than per code point, so every case containing a combining mark was one short — 48% passing, and the pattern was uniform enough to name the cause before reading a second failure; LB8a was called unreachable in a comment, and is not. A ZWJ at the start of text has nothing to fold into, so LB10 turns it into an AL — but LB8a comes before LB10, so ZWJ × still applies. Then the fix needed a second correction: the flag means "this unit's last code point is a ZWJ", so folding a combining mark on top of one clears it. The suite distinguishes those two readings in 24 cases; LB19a's last line tests the character before the quotation mark, not the mark itself. Three cases, all of them CJK text with curly quotes.
None of those would have been found by a corpus somebody wrote by hand, which is the argument for the exhaustive-but-not-independent gate rather than against it.
v0.1.239 — 2026-08-13
_A grapheme cluster is what a reader calls a character, and it is what a renderer has to advance by. Four of UAX #29's rules are not local, and those four are the whole difficulty._
contrib/unicode/grapheme.mere UAX #29 extended grapheme clusters
contrib/unicode/gcb_table.mere generated: 1,631 ranges, three UCD properties folded into one
scripts/gen_unicode_tables.sh derives the table from the UCD
scripts/unicode_parity.sh 8,509 inputs against node's Intl.Segmenter
`Grapheme.clusters`. á is one cluster, 👩👩👦 is one, 🇯🇵 is one, \r\n is one. Code points are not the unit and neither are bytes — cursor movement, selection and glyph advance all break visibly when the wrong one is used.
Most of UAX #29 is local: read the class of the code points on either side of a position and decide. Four rules are not, and the walk carries exactly four pieces of state, one per rule. GB12/13 needs how many regional indicators precede rather than whether one does (🇯🇵 is one cluster, 🇯🇵🇯 is two). GB11 needs whether the ZWJ was itself preceded by ExtPict Extend* — the ZWJ alone does not say. GB9c needs whether a Linker appeared between two Consonants, which is why the InCB property is in the table at all. And GB9b is decided by the left character, the only rule that looks that way. breaks_between is written in the standard's own order so it can be checked against it line by line.
Three UCD properties folded into one class per code point, from three different files because that is how the UCD is arranged — and the folding is only sound because of three facts the generator asserts rather than trusts: every Extended_Pictographic code point has Grapheme_Cluster_Break=Other (all 2,848), InCB=Consonant is disjoint from the non-Other breaks, and InCB=Linker/Extend live inside Extend or ZWJ.
The table shape carried over unchanged from the JIS slice, which is the point of having settled it: 1,631 ranges as a fixed-width hexadecimal literal, fourteen characters each, with Other as the unstored default. A lookup is a binary search rather than an index this time, and nothing else about the decision changed — including the reason for hex, which is still that a str is strlen-based on the LLVM backend.
The Unicode version is pinned to the oracle's, before being bitten rather than after. Intl.Segmenter follows whatever node's ICU implements; a table of a different vintage would differ from it for reasons that are neither a bug nor interesting. Both the generator and the harness assert process.versions.unicode, so a node upgrade fails with one line instead of a page of diffs. That is the node 22/24 lesson from url_parity applied in advance.
The oracle here is worth more than the others in this repository. It is not a second reading of a specification this code also reads — it is ICU, which is what browsers ship. 8,509 inputs, none hand-picked: every ordered pair from 22 class representatives, every triple from 8, every quadruple from 7 (which is what reaches ExtPict Extend ZWJ ExtPict), runs of 1..8 regional indicators and 1..5 of each repeating shape, and both ends of every one of the 1,631 ranges in the generated table — so a shifted range shows up as a segmentation difference rather than waiting for a character nobody tested. All agreeing.
The corpus is generated once and written to a file both sides read, because writing the same list of code points twice in two languages is how the two lists come to disagree.
v0.1.238 — 2026-08-13
_The first table in this project too large to write as code — so it is generated, and the question of what a 35KB literal does to five backends is answered by measuring it._
scripts/gen_jis_index.sh derives both JIS indexes from the Standard's own files
contrib/encoding/jis_index.mere generated: 2 x 8,836 slots as fixed-width hex
contrib/encoding/jis.mere Shift_JIS and EUC-JP
scripts/encoding_parity.sh + 196,608 sequences, compared a different way
The table question, settled by measurement. A 35,344-character string literal compiles and runs identically on interp, C, LLVM and Wasm, and mere -rv emits a 37,633-byte RV32I image from it — so the worry about what a large literal does to a backend's rodata is answered rather than assumed. (RV32I verified at emit; running it needs an emulator that lives elsewhere.) No startup expansion, no data file, no compression: a slot is four O(1) char_at reads.
The encoding of the table is decided by the LLVM backend's `str`, not by size. Raw 16-bit values would be half as long, and are unusable: a str there is strlen-based and a raw table is full of 0x00 bytes (U+00A2 is 00 A2). Fixed-width hexadecimal is NUL-free by construction, and 0000 doubles as the hole sentinel because U+0000 is not a mapping either table produces. Two open questions meeting in one design decision is worth writing down.
`contrib/encoding/jis.mere`. Shift_JIS and EUC-JP. Four details that are easy to get plausibly wrong: Shift_JIS's lead offset is two numbers (0x81 below 0xA0, 0xC1 above, because the single-byte katakana range sits in the middle of what would otherwise be one contiguous lead range); there is a private-use window past the end of the table (pointers 8836..10715 are U+E000 onwards, the vendor extensions the encoding grew); an unmapped pair puts an ASCII trail byte back (82 40 is U+FFFD then @, the same rule as UTF-8's); and EUC-JP has two tables and a three-byte form, a 0x8F lead selecting JIS X 0212 for the pair that follows it.
The tables are derived from the Standard, not from node — and that is a change of oracle with a measured reason. node's shift_jis is ICU's CP932. Its index jis0208 is identical to the Standard's in all 8,836 slots, but its index jis0212 maps 21 pointers the Standard does not (from pointer 7708, the small Roman numerals), it remaps three single bytes in a cycle (0x1A→U+001C→U+007F→U+001A), it treats 0x80 as an error where the Standard returns U+0080, and its error recovery consumes a malformed sequence whole instead of putting an ASCII trail back. A browser implements the Standard, and accepting 21 code points the Standard does not is the same failure mode contrib/url guards against — agreeing with an implementation instead of a specification, in the permissive direction.
So the generator reads the Standard's published index files and pins each file's `Identifier:` hash, which is the oracle-version lesson applied to a data file that states its own version.
The gate keeps node, and asserts something exact rather than something weaker. Strict equality would be asserting ICU. Instead: the two implementations never disagree about which character a byte sequence is — 75,547 inputs that both call characters, all agreeing — and every remaining difference must be either error handling (U+FFFD on at least one side) or the named three-cycle. Anything else is a table or pointer bug and fails. The sweep is exhaustive over all three two-byte spaces (196,608 more sequences, 268,032 in total).
That framing was not chosen up front. The first run reported thousands of differences; each class was then measured and named, and what fell out was that the disagreements are entirely about error handling and never about identity. A gate that says that is more useful than one that says "equal".
v0.1.237 — 2026-08-13
_A decoder's interesting behaviour is all in its error cases, and the Encoding Standard specifies how many U+FFFD a malformed sequence produces — which is not one per byte._
contrib/encoding/decode.mere UTF-8, windows-1252, and the label table
scripts/encoding_parity.sh 71,424 sequences swept against node's TextDecoder
`contrib/encoding/decode.mere`. Bytes off a wire into a str. Decoding never fails — every malformed sequence becomes U+FFFD — so decode_utf8 returns a str rather than an ?str. The only thing that can fail is recognising a label, and decode returns ?str for a different reason: the label named a real encoding that is not implemented yet, so a caller can tell "unknown encoding" from "known but unsupported" and say so instead of guessing.
Two facts here are counter-intuitive enough to be worth stating outright.
ascii means windows-1252. So do latin1, iso-8859-1, us-ascii and ansi_x3.4-1968. The Standard folds them into one encoding on purpose, because that is what the deployed web already did — so a page that calls itself ascii and contains byte 0x80 has a euro sign in it, and an implementation that "helpfully" treats ascii as 7-bit produces U+FFFD where every browser produces €.
And the replacement count is specified: F1 80 80 41 is one U+FFFD then A, because those three bytes were a valid prefix and are one error together, while E0 80 80 is three, because E0 requires its first continuation in A0..BF — so 80 is not part of the sequence at all, and each remaining byte is then reconsidered on its own and fails on its own. Getting this wrong changes how many characters a page has, which changes every offset after it, and is invisible on valid input.
That same bound mechanism does all the other rejecting with no separate checks afterwards: ED requires 80..9F, which is exactly what keeps the surrogates unrepresentable, and F0 requires 90..BF while F4 requires 80..8F, bounding the range at both ends.
The gate sweeps rather than samples, because the error cases are invisible on valid input: every single byte (256), every two-byte sequence (65,536), every three-byte lead × first continuation (4,096), every four-byte lead × first continuation (1,280), and every byte through windows-1252 (256) — 71,424 comparisons against node's TextDecoder. The two-byte sweep is exhaustive; the three- and four-byte sweeps are exhaustive in the dimension that carries the logic, with the rest held valid. Neither side builds an input as a string literal — both loop over the byte — so there is no escaping layer to get wrong. Labels are checked rather than derived (40 of them, including three that must not resolve), and that gap is printed as a SKIP: the harness cannot discover a label nobody listed.
A new, measured consequence of the LLVM backend's `str`. A decoded 0x00 is U+0000, and because that backend's str is strlen-based, str_of_codepoint 0 yields the empty string — the character does not truncate the text, it vanishes, and every offset after it shifts by one. 41 00 42 decodes to length 3 on interp, C and Wasm, and to length 2 on LLVM. Silent loss is worse than truncation for anything that then indexes the result, so this is documented as 0x01..0xFF-safe there for now, and test/parity/encoding_decode.mere omits 0x00 with the reason written at the top rather than asserting the bug.
Shift_JIS and EUC-JP resolve as labels but have no decoder yet. They need the JIS X 0208 index — 6,879 code points — which is the first table in this project too large to write as code, and that question deserves settling once rather than per encoding. meta charset sniffing is likewise absent on purpose: it is a scan for tags and belongs with an HTML tokenizer, and doing it in two places is how the two come to disagree.
v0.1.236 — 2026-08-13
_An IPv6 address has many spellings and exactly one canonical form, so the serialiser is as much of the answer as the parser._
contrib/url/ipv6.mere eight 16-bit pieces, in and back out
scripts/url_parity.sh + 53 IPv6 literals
`contrib/url/ipv6.mere`. [0:0:0:0:0:0:0:1] and [::1] are the same host, and until now the host field carried whichever one was typed — a comparison on the text would have called them different origins. Both now parse to the same eight pieces and serialise to [::1].
The serialiser's rules are narrow enough to get plausibly wrong while still looking right on [::1], which is what the corpus is weighted towards: the longest run of zero pieces is compressed, ties go to the first run ([1:0:0:1:0:0:1:1] is [1::1:0:0:1:1]), and a run of exactly one zero is left alone ([1:0:2:3:4:5:6:7] keeps it). :: may stand for a single piece on the way in — [1:2:3:4:5:6:7::] is 1:2:3:4:5:6:7:0, which then serialises without any :: at all.
A trailing dotted quad occupies the last two pieces and is strict dotted decimal: four parts, no leading zeros, no hex. That is deliberately not the multi-base parser a bare host uses, so [::0x7f.1] and [::01.2.3.4] are not addresses while http://0x7f.1 is 127.0.0.1. Two IPv4 parsers in one file looks like duplication and is not — they are answering different questions, and the comment says so at both.
Implementation note worth keeping. The Standard's IPv6 parser is a pointer-walking state machine over a mutable eight-slot array. This is the same algorithm expressed by splitting: once on :: (three parts means two of them, which is one too many), then each side on :, with the zero fill computed from the two lengths. A trailing quad is folded into two hex groups before the split, so nothing downstream has to know about dots — and anything else containing one fails on its own, because . is not a hex digit. No mutation, and the failure cases fall out rather than being enumerated.
v0.1.235 — 2026-08-13
_Resolution against a base, hosts checked one byte at a time — and the first place the oracle turned out to be wrong, which the harness now says out loud._
contrib/url/host.mere resolve; forbidden host code points; domains decoded
contrib/url/percent.mere decode_strict
scripts/url_parity.sh + 272 resolve pairs, + 190 single-byte hosts, + DIVERGE
`Url.resolve base input`. What every link in a page needs: the reference wins wherever it says anything and the base fills in the rest. A relative path replaces the base's last segment and is then normalised, so ../d against /a/b/c is /a/d; a reference that says anything at all about the path drops the base's query, and one that says nothing keeps it.
The rule worth stating separately: a reference carrying a special scheme that names the base's own scheme is relative, not absolute. http:d against http://h/a/b/c is http://h/a/b/d; https:d against the same base is https://d/. Getting that backwards turns a same-origin relative link into a request to a host the page named.
Forbidden host code points, and domains are decoded. A domain is percent-decoded before anything else looks at it, so %41 is a and %2e is a label separator — http://0x7f%2e1 is 127.0.0.1. An allowlist that inspects the host before decoding sees an opaque name where there is an address. A malformed escape means it is not a domain (http://a%b/ is not a URL, hence Percent.decode_strict), and the forbidden code points are checked after decoding, so http://a%2fb has a / in its host and is rejected. An opaque host is neither decoded nor folded, but the forbidden points still apply to it.
Two new derived gates. Resolution is checked as a cross product — 8 bases × 34 references = 272 pairs, each compared as href — because the interesting cases are combinations, and picking pairs by hand is picking the ones already thought of. Hosts are checked one byte at a time: http://aXb/ and foo://aXb/ for every X in 0x20..0x7E, 190 comparisons, with neither side building the string as a literal (both loop over the byte, so there is no escaping layer to get wrong). That is the component where being too permissive is worst, and it caught the missing forbidden-code-point check and the missing domain decode together.
And the oracle was wrong once. The Standard's no scheme state admits a reference against an opaque base only when the reference's first code point is #. node v24 accepts any reference that merely contains one — new URL("?q#f", "mailto:x@y") is mailto:x@y?q#f, and new URL("e?q2#f2", "mailto:x@y") invents a path segment and gives mailto:x@y/e?q2#f2. We follow the Standard, and the harness prints these as DIVERGE lines with the count and both answers rather than dropping them from the corpus.
Overruling the oracle here and not elsewhere is a judgement, so the reasoning is recorded next to it: the spec text is explicit (unlike the ^ in the path set, where the prose did not name it and the implementation did), and the divergence is in the permissive direction — accepting more than you should is the failure mode this gate exists to find. A gate needs somewhere to say "the oracle is wrong here" or the first time it happens the answer is to quietly delete the test.
v0.1.234 — 2026-08-13
_The rest of the URL, and two places where the previous slice's simplifications turned out to be wrong about where a delimiter stops mattering._
contrib/url/path.mere dot segments, query / fragment split, per-piece encoding
contrib/url/host.mere the authority rules corrected; href
scripts/url_parity.sh 95 inputs vs node, nine fields each
`contrib/url/path.mere`. Splitting rest into path / query / fragment, resolving . and .., and encoding each piece with its own set. The fragment starts at the first # and the query at the first ? before it; neither delimiter is special inside the fragment, so #a?b is a fragment of a?b.
A dot segment is matched on the whole segment and includes its percent-encoded spellings case-insensitively — ., %2e, .., .%2e, %2e., %2e%2e. So /a/%2E/b is /a/b, but /a/..%2f is left exactly as it is: ..%2f is a segment that begins with two dots, not a dot segment. And a trailing dot segment leaves an empty segment behind, which is what keeps the trailing slash: /a/b/.. is /a/, not /a.
The path is never decoded. %41 stays %41 and %2f stays %2f in whatever case it arrived in — decoding an escaped slash into a separator is a path-traversal bug with a long history.
An opaque path — no authority and no leading / — is not segmented and gets only the c0_control set, which is why mailto:a b keeps its space. foo:/a/../b is foo:/b, but foo:a/../b keeps its dots.
Two things the previous slice had wrong, both found by widening the oracle corpus.
A special scheme's authority needs no slashes, or any number of them. `http:h`, `http:/h`, `http:\\h` and `http:///h` all have host `h`. The old code required `"//"` and rejected the rest, so it saw a relative reference where there was in fact a host — which is the shape of a real allowlist bypass. A backslash also ends an authority: `http://u\p@h/` has host `u`, because the `\` comes before the `@` ever does. Backslash folding belongs to the path, not the whole URL. The old code folded \ to / across everything after the scheme. http://h/a\b?c\d#e\f has path /a/b but query c\d and fragment e\f, both keeping the backslash and both leaving it unencoded.
`href`, and three booleans. has_authority, has_query and has_fragment are fields rather than being folded into the strings, because presence and content are separate state: foo:// and foo: have the same empty host but only one has an authority, and http://h/? has an empty query that still serialises its ?. href is what needs them, and round-tripping is the honest test of a parse — a dropped field or an invented delimiter shows up there and nowhere else.
The gate is now one section of nine fields over 95 inputs (scheme|user|pass|host|port|path|search|hash|href), up from five fields over 24. search and hash are derived from the two bools to match node's shape, where the delimiter is carried when the component is non-empty and dropped when it is present but empty.
test/parity/url_host.mere grew to match, and examples/url_parse_demo.mere prints every field of a parse. Both hold all four backends to one output.
One find worth writing down for anyone else generating Mere source from a shell: a { in a string literal opens interpolation, so a corpus entry like mailto:a{b has to be escaped as \{. The compiler's error message says so exactly, which is the only reason this cost one run instead of an afternoon.
v0.1.233 — 2026-08-13
_Two gates, two classes of bug: the oracle found three places the parser was too permissive, and the backend-parity test found that str_replace had never worked when compiled to C._
contrib/url/host.mere scheme / userinfo / host / port, WHATWG
scripts/url_parity.sh + authority section, 24 inputs vs node, per field
lib/codegen_c.ml str_replace: allocate a str, not a raw buffer
test/parity/string_ops.mere str_replace across four backends, with lengths
`contrib/url/host.mere`. The scheme and authority half of a WHATWG URL: cleaning, the scheme, userinfo / host / port, and origin. rest hands the path, query and fragment over untouched for the next slice. It returns ?url_parts rather than failing, because rejecting input is half of what a URL parser does — and on the wasm backend fail sets a flag and returns a sentinel instead of unwinding, so a caller sitting between the failure and its try_or still runs. A rejection has to be a value.
Two behaviours in here are the kind that become security bugs when an implementation guesses. A host that parses as a number is an address, in any base the Standard allows: http://0x7f.1 is 127.0.0.1, and an allowlist that only understands dotted decimal passes it through as a hostname and then resolves it to localhost. And a domain is lowercased while an opaque host is not — case folding belongs to the special schemes, not to hosts.
The authority gate: 24 inputs against node, compared field by field. One line per input as scheme|user|pass|host|port, so a mismatch names the field rather than just the URL. Inputs node rejects must come back None from us too, because a parser that accepts more than the oracle is the failure mode that matters for anything that then makes a request. It found three:
`http://256.1.1.1` and `http://1.2.3.4.5` were accepted as hostnames. Conflating "the IPv4 parse failed" with "so it must be a domain" is exactly the hole above, from the other side. Fixed by asking first whether the host ends in a number: if it does the host is an address and a bad one is invalid, and if it does not IPv4 never applies. http://a@b@h/ produced an unencoded username. The userinfo set includes @, so it would re-split differently on the way out. Now Percent.encode Percent.userinfo.
`str_replace` was returning a buffer with no length header on the C backend. A Mere str in C carries its length in a size_t at [-1], written by __lang_str_alloc. __lang_str_replace allocated with the raw __lang_region_alloc instead, so __lang_str_size of its result read whatever bytes happened to precede the buffer in the region. Those bytes were zero often enough that every replacement came back empty — which is how this surfaced: Url.clean removes tabs with str_replace, so on the C backend it returned "" and every valid URL was rejected, while node and the interpreter agreed with each other the whole time. The cap is a worst case, so the header is also corrected down to what was actually written. The empty-needle guard now tests the length rather than old[0], so a needle that is a NUL byte is replaceable like any other.
A sweep for the same shape found no other instance: __lang_str_alloc is the only other function that region-allocates directly, which is its job.
str_replace had no cross-backend coverage at all — only an interpreter test — which is why this survived. test/parity/string_ops.mere now exercises eleven cases (shorter, longer, equal, absent, whole-string, empty subject, empty needle, overlapping, UTF-8, grow-from-one, then-concatenated) and prints the length of every result as well as the text. The length is the point: the text alone would not have caught a garbage header on every input. Reverting the fix turns that test red.
v0.1.232 — 2026-08-13
_A percent-encode set is a list of bytes somebody transcribed, so it was checked against somebody else's implementation instead._
contrib/url/percent.mere 7 sets, encode / decode
scripts/url_parity.sh derives each set from node's URL, byte by byte
`contrib/url/percent.mere`. The URL Standard has no single escaping rule; it has a stack of percent-encode sets, and which one applies depends on the component being written. Getting the component wrong is not cosmetic — a # left unencoded in a path ends the path. So encode takes the set as a predicate on a byte and the named sets are supersets of one another: c0_control ⊂ fragment, and c0_control ⊂ query ⊂ special_query, and query ⊂ path ⊂ userinfo ⊂ component.
Encoding walks bytes, not codepoints, which is what the Standard says and is the reason a str being a byte buffer is the right shape here. Decoding leaves a % that is not followed by two hex digits exactly as it found it: refusing would make a literal percent sign unrepresentable, and decoding it as zero would invent a NUL.
`scripts/url_parity.sh` derives the sets rather than asserting them. For each byte in 0x20..0x7E it puts that byte alone into a component of an http URL, reads node's serialisation back, and records whether it came out as %XX. That is node's set for the component; ours is diffed against it. A fixture file would only have covered the bytes somebody thought to write down — and this found two wrong sets on the first run: fragment was missing and path was missing ^.
^ in the path set is in there because the oracle encodes it, which is not the same thing as the prose naming it. That is recorded at the definition, so the next person sees a citation rather than a magic number.
Four bytes per component cannot be probed this way and the harness prints them as SKIP rather than passing them silently: a byte that delimits the component under test ends it instead of being escaped in it, a lone space is stripped by URL parsing before escaping happens, . in a path is resolved away, and \ is normalised to / for the special schemes.
test/parity/url_percent.mere holds all four backends to the same output. scripts/url_parity.sh skips when node is absent, like qemu_virt.sh does.
`contrib/http/query.mere`'s `url_encode` / `url_decode` are deliberately untouched. They implement the query-string convention a server wants — an allowlist of alnum -_.~ — which is not any of the Standard's sets. Changing them would change every existing caller's behaviour, so the two live side by side with the difference written down.
Not here yet: the parser itself, punycode, and decode into bytes. That last one is blocked on the llvm str being strlen-based, so Percent.decode "%00" differs by backend; until that changes, callers that can receive %00 should treat this as ASCII-safe.
v0.1.231 — 2026-08-13
_The codepoint pair, written in the prelude rather than five times._
str_of_codepoint : int -> str codepoint_of : str -> int
codepoint_at : str -> int -> int
Q-014's last deferred piece. The Unicode question was settled in v0.1.38/45 — a str is a UTF-8 byte buffer, with a codepoint view composed above it — and one sliver was left open on purpose: the integer form of a codepoint, to be added "when an external program asks for it." Percent-encoding and punycode ask for it.
It went in as prelude source, not builtins. The first attempt added them to the typer, the interpreter and the C backend, and stalled at the point of hand-writing the decoder in LLVM IR and WAT — at which point it was obvious the whole thing composes out of chr, ord and char_at, which every backend already has. The codepoint layer has been prelude-composed since v0.1.38 for the same reason. Six declarations, no codegen touched, and all four backends agree because there is only one definition.
chr and ord stay byte-shaped: the redis, pg and http drivers build 0..255 bytes with them. And the divergence chr carries — out-of-range raises on the interpreter and masks on wasm and llvm — is not repeated: the new pair fails on all four. A scalar value is 0..0x10FFFF less the surrogates, and overlong encodings are refused, because two spellings of one character is how a filter gets walked past.
test/parity/codepoint_pair.mere round-trips both sides of every length boundary (127/128, 2047/2048, 65535/65536, 1114111) and refuses ten ill-formed inputs: negative, above max, both surrogate ends, empty, two codepoints, truncated, C0 80, a bare continuation byte, and a lead byte with nothing after it.
Getting that file green on four backends turned up three older holes, none of them this change's:
- The llvm `str` is not byte-safe. The v0.1.129 work reached C and wasm; llvm still
implements str_len as call @strlen and has no __lang_str_size anywhere, so a str holding a NUL cannot exist there. str_len (chr 0) is 1 on interp, C and wasm, and 0 on llvm, whose chr returns a raw pointer into a 256-entry table. U+0000 is left out of the parity file for this reason and checked in the unit tests instead.
- `fail` does not unwind on wasm. It sets a global and returns a sentinel, so whatever
sits between the fail and the try_or still runs. 1 + (fail "b") survives, because an integer tolerates the sentinel; str_len (fail "b") traps the module and the program dies without even the try_or default. Which failure you get depends on what the consumer does with the value.
- `print` stops at an interior NUL.
str_lensays 2 andprintemits one character,
on the same value. print_bytes (v0.1.219) is the byte-safe writer, but print is silently lossy rather than either correct or refusing.
Also: codepoint_at is let rec for a reason that has nothing to do with recursion. The decl loop is duplicated in test helpers that bind Top_let_rec and Top_let separately, so a plain let here reaching back to let rec utf8_at is unbound in those copies. Same duplication that made the quadratic-inference fix land in six places.
v0.1.230 — 2026-08-13
_A loop was a function calling itself, and whether it survived was up to clang._
while, 200 000 iterations before after
clang -O0 SIGSEGV ok (10 000 000 ok)
clang -O1 and up ok ok
`while` never reached a backend as a loop. The parser desugars while cond do body into a tail-recursive let rec, so by the time codegen sees it there is nothing left that says "iteration" — the C backend emitted a self-recursive function and left the tail call to the optimizer. That made the iteration bound a property of the build: -O0 died with no output at somewhere between 100 000 and 200 000, -O1 did not.
The build that dies is the one scripts/debug_info.sh prescribes: clang -g -w -O0. So the debugger work of v0.1.212-215 — a compiled program you can step through against its Mere source — and a program with a loop in it were mutually exclusive, and nothing said so. A 100 KB input walked one byte at a time is 100 000 iterations.
for i in a..b do body was worse, not better: range materialises the whole range as a list and list_iter then recurses down it. There was no safe way to iterate.
A self tail call now lowers to a goto. Tail position is tracked the way codegen_wasm already tracked it — taken on entry to emit_expr, handed on only by the cases where a subexpression really is in tail position (Annot, both arms of If, the body of Let in each of its four pattern forms, the body of Let_rec, and match arms; deliberately not Region_block or With, which have reclamation left to do after the value). The call becomes argument temporaries, assignments to the parameters, and a jump:
int f(vec* i, long long n, int u) {
int __mere_ret;
__mere_tail: ;
return (cond ? ({ __auto_type __mt0 = i; ...; i = __mt0; ...;
goto __mere_tail; __mere_ret; })
: 0);
}
__mere_ret is never evaluated; it is there to give the statement expression the function's return type, which is what lets a goto sit in expression position at all — the C backend is expression-oriented and cannot restructure a body into a statement loop. Verified by hand first, with a struct return, before any of this was written.
Two things the parity suite caught, both of which had built and linked cleanly.
schema_reflect printed nothing. Rewriting If to name its operands had reversed the order they were emitted in: OCaml evaluates the right argument of ^ first, so the original chain emitted else, then, cond — and emit_expr interns strings and numbers closures as it goes, so the order is part of the output. Restored, with a comment, since nothing in the expression makes it look load-bearing.
Then it looped forever. *(&p) rather than p is why the assignments go through aliases taken at function entry: Json.parse_object binds let j = skip_ws s j five times, so by the tail call the name j is a local shadowing the parameter. Assigning by name wrote the shadow and jumped with the parameter unchanged. The address is taken before the body runs, so it names the parameter however the body rebinds that name.
Neither showed up in dune runtest or ctest.sh — the first needs a program whose output depends on interning order, the second a function that shadows its own parameter and tail-calls itself. MERE_NO_TAIL_LOOP=1 disables the rewrite, which is what made "is it this change?" a one-variable question; MERE_TAIL_ONLY=<substring> narrows it to matching callees.
The prologue is only emitted for functions that actually produced a jump, so a function with no self tail call is emitted exactly as before. That is per function, not per program: the prelude has tail-recursive functions of its own, so all 84 parity programs contain a rewritten function, which is what makes 85/0 there worth something rather than an accident of coverage. examples/bst.mere has one of its own (lookup), and its output is byte-identical before and after, at -O0 and -O2, and against the interpreter.
__attribute__((noinline)) on the show_* helpers has a comment from v0.1.31 saying it exists because an inlined helper's escaped local defeated clang's sibling-call optimisation and deep loops overflowed. That was this bug, seen once from the other end and worked around locally.
The LLVM backend has the same defect — -O0 dies at 10 000 000, and the emitted IR carries no tail marker anywhere. The fix there is musttail, which LLVM honours regardless of optimisation level, but it requires the call to be immediately followed by the ret of its result, and this backend has no tail-position notion and returns from many places. Measured and left open rather than guessed at.
`show` on a float gave four different answers. The interpreter formatted it, C printed <unsupported>, LLVM printed () — a number rendered as unit — and Wasm printed <?show_float?>. Each backend's show generator simply had no float case and fell through to a different placeholder. parity.sh never caught it because no program in the suite showed a float; the probe that found it could not print the time it had just measured. The formatter was already there and already shared, so this only wires it up in the three generators. All four now agree, including 0.1 + 0.2 at 17 digits and the trailing .0 on whole values.
v0.1.229 — 2026-08-13
_Formatting a long function was quadratic in two places, neither of them arithmetic._
16 000 nested lets 22 810-line file
fmt 1.30s -> 0.05s 1.23s -> 0.67s
A run of `let ... in` is now written out as a run. The formatter recursed into the body and concatenated what came back, so every level copied the whole remainder of the function into a new string. All 87 profile samples were inside Stdlib.(^). Each binding in a run sits at the same indent as the one before it, which is what makes the run flattenable at all.
`rename_free_vars` carried its shadow set as a list, extended with @. That copies the whole list at every binding, and each Var scanned it linearly — so a pass whose entire job is renaming names cost O(depth²). It is a set now, which shares structure. This one is not the formatter's: mere -t on the same input went 2.60s to 1.34s, because every path that parses pays for it.
The formatter's output is unchanged: 193 files — every example, every contrib source, every parity program, and a 22 810-line one — format byte-identically before and after. That is the only acceptable evidence for a change to a formatter.
What is left, measured rather than assumed: mere -t on that 16 000-deep chain is still quadratic (0.10 / 0.36 / 1.34 at 4k / 8k / 16k), because the type environment is an association list and looking up a variable bound at the top of a function walks every binding since. The deepest run of consecutive lets in this repository's own sources is 281, and those files type-check in 0.02s. So it is real, it is not biting, and rewriting the environment on the strength of a synthetic chain is the kind of thing Q-023 exists to say no to. Recorded as Q-026.
v0.1.228 — 2026-08-13
_The support matrix is asked for rather than remembered, and it found three holes._
There have been three hand-written versions of "which backend has which host builtin": a table in the design notes, and a list in each of codegen_llvm.ml and codegen_wasm.ml naming the builtins with no lowering. All three had gone stale in the same direction — print_int gained real lowerings in v0.1.190 and file_pwrite_bytes in v0.1.222, and both were still listed as missing, in code that is inert rather than wrong. A table nobody can trust is worse than no table.
scripts/host_matrix.sh produces one instead: fifty one-line programs, each using a host builtin, emitted for each backend, and the outcome recorded as yes, refused (the backend says so itself), MISSING (unbound variable — the compiler blaming the user for a backend hole), or error. The result is docs/host-matrix.md, checked in and diffed on every run.
The first run found three MISSING, all of them the failure this project's loud-failure rule exists to prevent:
- `read_bytes` and `write_bytes` on LLVM and Wasm. They arrived with the
bytes
type in v0.1.216, after the per-backend lists were written, and fell straight through to unbound variable. The hole those lists exist to close, reopened by a later feature — which is exactly what a generated matrix is for.
- `par_map` on Wasm, which said
unbound variable: __pm_f1: a name the user never
wrote, about a function they do not know exists. par_map is desugared at parse time into spawn + channel + list_map, and Wasm cannot resolve the captured function from inside the nesting that produces. The limitation stands; it now names par_map.
The matrix is now 50 builtins, 0 MISSING.
What it cannot see is a builtin that compiles and then does nothing — tcp_set_timeout on Wasm in v0.1.227 returned success and hung. Only running a program catches that, which is scripts/parity.sh and scripts/socket_parity.sh. The two kinds of check answer different questions and neither replaces the other.
v0.1.227 — 2026-08-13
_A capability that quietly did nothing now refuses._
tcp_set_timeout had a Wasm helper that ignored both arguments and returned 0 — the same value the C version returns on success. So a program set a deadline, was told it had one, and blocked forever on the next read. Measured at ten minutes before it was killed.
That is worse than not having the capability at all: a missing feature is a compile error, a silent no-op is a hang. It is refused at the call site now, with the reason, until WASI's poll is wired up. A socket program that never sets a deadline is unaffected.
`scripts/socket_parity.sh` is what found it, and is why this was possible to find at all. parity.sh runs eighty-odd programs across four backends and none of them opens a socket, because a socket program needs two endpoints and a host willing to grant a network. So the whole socket family — tcp_listen through tcp_close, all of it implemented against p2 wasi:sockets — had never been run. It works: the same round trip prints the same four lines natively and under wasmtime -S inherit-network=y.
Two things stay unequal and are recorded rather than fixed:
- A failed read is `0` on Wasm, the same value C uses for a clean end of stream.
Distinguishing them means decoding WASI's stream-error variant instead of its is-error bit. The check above deliberately does not assert on it.
- `sleep_ms` is interp + C only, which is why the parity program contains no
clock at all.
The general shape is the one the host-builtin registry work named: a capability with a per-backend implementation and no single place that says which backends really have it. This is the second silent gap found by running one, after print_int in v0.1.190.
v0.1.226 — 2026-08-13
_A failed read says which failure it was._
> 0 bytes read
0 the peer closed cleanly — end of stream, not an error
-1 nothing arrived before the deadline
-2 the connection is gone
-3 anything else
tcp_read returned read(2)'s -1 for everything. With SO_RCVTIMEO set that covers two opposite events: the deadline passed with nothing arriving, and the connection broke. One means wait again, the other means reconnect. The mraft dogfood told them apart by timing the call and asking whether it had failed slowly enough to have been a timeout — inferring a cause from a duration, and recorded as that repository's P4.
The codes stay negative, so every existing < 0 check is unaffected.
scripts/tcp_read_codes.sh produces all three rather than describing them: a socket nobody writes to, a peer that closes cleanly, and a peer that aborts with data still unread in its own receive queue — which is what makes close() send RST instead of FIN, and the only reliable way to get ECONNRESET without setsockopt(SO_LINGER).
Two things this turned up:
- The socket externs were documented nowhere. Same as the positioned-IO family in
v0.1.222: real, on every network program, and absent from the stdlib reference. Both are now in it.
- Wasm disagrees about `0`. Its
tcp_readgoes through WASIsock_sreadand
returns 0 on error, where C returns 0 only at end of stream. So a program that treats 0 as "the peer closed" is wrong there. No parity test covers sockets, which is why nobody had noticed; recorded rather than fixed, because fixing it means mapping WASI's error set and there is no Wasm program that opens a socket yet.
This answers the concrete case. The general question mraft's P4 asks — how a capability reports why it failed, rather than only that it did, across an FFI boundary that is C-shaped — is not answered by a convention about negative integers, and is still open.
v0.1.225 — 2026-08-13
_The editor stopped accepting names that no longer exist._
Rename a constructor and keep using the old name:
type t = Gamma | Delta;
let x = Alpha in print "b" // Alpha was renamed away
The compiler says unknown constructor: Alpha. The language server said nothing — clean file, no diagnostic, until the build failed. Three versions of one document, driven through mere lsp:
v1 type t = Alpha | Beta; clean
v2 type t = Gamma | Delta; ... Alpha clean <- wrong
v3 type t = Gamma | Delta; ... Gamma clean
The parser has had reset_decl_state since a type in one program could shadow one in the next. The typer's registries — constructors, types, records, views, drop / sync / local types, record aliases — were never cleared. A compiler process checks one program, so nothing ever noticed; a language server checks one document per keystroke. The same leak made a view Pair and a later type Pair collide, which is how this was found: a test declaring a record hit "view Pair must be constructed inside a region block".
Two changes, and the second is the one worth reading:
Typer.reset_type_registries ()runs next to the parser's reset. It restores a
snapshot rather than emptying the tables, because the built-in capability records (Logger, Metrics) are registered at module load and a plain reset deletes them permanently. The snapshot is taken on the first call — which comes before any program's declarations are processed — so it does not depend on where in the file it is written.
- What types exist is now established by `parse_program`, not only by whichever
later walk happens to visit the declarations. Typer.infer on a desugared program never registered anything, so a caller that skipped process_decls type-checked against whatever the previous program in this process had declared. Dozens of tests did exactly that and passed, because nothing was ever cleared: resetting without this made Nil unknown. Registering is idempotent, so the later walks are unaffected.
That second point is the real defect. The first is what exposed it.
v0.1.224 — 2026-08-13
_A name bound by a constructor pattern can cross a thread boundary._
match ports with
| Cons (p, rest) -> let _ = spawn (fn () -> sender p out) in ...
That was refused with cannot capture \p\ of unknown type across a thread boundary, with ports : int list written down two lines above. The capture check's pattern binder handled P_var and a tuple pattern over a tuple type, and bound everything else — constructor patterns, record patterns — with an unknown type, which makes those names unusable across spawn. The mraft dogfood hit it on every peer it spawned a thread for, and the workaround (let q = (p : int) in, capture q) reads like superstition because it is: an ascription that tells the program nothing it did not already know.
The declared payload type is enough to do better, with no unification and nothing this pass may mutate: the constructor registry knows the type parameters and the payload type, and the scrutinee's own arguments say what to substitute for them. Record patterns get the same treatment, field by field.
The diagnosis in the dogfood's PAIN.md was wrong, and worth recording as such. It said the Send check ran before inference had propagated the annotation — a plausible story about pass ordering, told without looking. The types were fully resolved; the binder simply never looked at that shape of pattern. A payload whose type genuinely is a variable is still refused, and now says so accurately (of polymorphic type rather than of unknown type).
v0.1.223 — 2026-08-12
_The reserved-name warning now looks at type names, which is where it was needed._
line 1, col 6: warning: type name `wait` collides with a C type, keyword or libc
symbol — this will be a compile error at codegen, from the C compiler rather than
from here.
A Mere type lowers to typedef struct <name> <name>;, which claims both C's tag namespace and its ordinary one. type wait = ... therefore collides with union wait in <sys/wait.h> and with the wait() declared beside it, and the failure arrived from clang:
error: use of 'wait' with tag type that does not match previous declaration
The compiler has had a list of libc and C-keyword names since v0.1.55 and warns when a top-level let collides with one. It never ran for type declarations. The mraft dogfood named a type wait on its first day and got the collision from the C compiler — the same "documented thing failing at the wrong layer" shape as the print_int bug in v0.1.190.
The function list applies to type names unchanged (a typedef is an ordinary identifier), plus a new list of the struct tags the emitted headers bring in: wait, tm, timeval, timespec, stat, dirent, sockaddr, addrinfo, hostent, termios, winsize, sigaction, iovec, fd_set, div_t and the rest of that family.
Two things about the implementation are worth recording:
- `Top_type` carries no position, and adding one touches ten files. So the
parser records (name, loc) for each type it declares — the same shape as the per-program tables it already keeps for constructors and records — and Pipeline reads that. A warning an editor cannot place is a warning nobody sees.
- The warning is raised on both paths:
process_decls, which the compiler
takes, and infer_program, which the editor takes. It went in on the first one only, and the test that checks it through Pipeline.diagnostics failed — which is exactly the test being worth writing.
Nothing in examples/ or contrib/ trips it.
v0.1.222 — 2026-08-12
_A positioned write that takes the byte type the language grew afterwards._
let n = file_pwrite_bytes h off (bytes_of_str line)
file_pwrite was added for the mbtree dogfood in v0.1.115 and takes Vec[int]. The bytes type arrived in v0.1.216 and got its I/O boundary in v0.1.219 — so the language had a byte string, and the one API that writes at an offset could not take it. The mraft dogfood's write-ahead log had to explode every record into a Vec with one boxed int per byte before writing it.
file_pwrite_bytes : File -> int -> bytes -> int is the same operation over bytes. Both remain: mbtree builds its pages as Vecs and has no reason to change.
All four backends, test/parity/file_pwrite_bytes.mere. Two of them cost almost nothing:
- Wasm: the host import already took a bytes pointer — the Vec version converts
first and then calls it. The new path is the same call with nothing in between.
- LLVM: one
fwriteinstead of a loop, sincebytesis{ i64 len, i8 data[] }
and bytes_len already loads from that layout.
- C: the runtime function had to be defined next to the bytes runtime rather
than with the other file_* ones, because those are emitted before struct mere_bytes has a body — the same ordering the ByteBuf freeze hit in v0.1.218.
The LLVM path also registers the Vec[int] instance even though it uses no Vec: the positioned-IO runtime is emitted as one block and file_pread's body calls the Vec accessors regardless. A program that only writes bytes carries a few unused functions, which is cheaper than splitting the block into per-function flags.
v0.1.221 — 2026-08-12
_A program's output no longer depends on whether someone redirected it._
print lowered to puts, and nothing in the runtime ever called fflush. C line-buffers a terminal and fully buffers a pipe, so a program whose output was redirected to a file printed nothing until it exited or accumulated 4KB.
Every dogfood until now was a batch program — it printed and exited, and exiting flushed. The mraft dogfood is the first Mere program meant to be watched while running, and it logged nothing at all: a server started with > log 2>&1 looked identical to one that had hung.
The C backend now emits setvbuf(stdout, NULL, _IOLBF, 0) in main; the LLVM backend flushes after each print (fflush(NULL), which needs no platform-specific stdout global — glibc and macOS name it differently). Both give piped output the same behaviour a terminal already had.
The LLVM change also removed a duplicate declare i32 @fflush(ptr): the file positioned-IO runtime had its own, and the second declaration is an error rather than a redefinition LLVM tolerates. scripts/parity.sh caught it — file_pio was the only program that emitted both.
v0.1.220 — 2026-08-12
_Type inference was quadratic in the number of bindings. It is linear now._
inference (mere -t) LSP, per keystroke
4 000 bindings 0.16s -> 0.02s 524ms -> 32ms
8 000 bindings 0.50s -> 0.04s 1834ms -> 65ms
16 000 bindings 1.72s -> 0.09s 7741ms -> 135ms
22 466 lines (real) 5261ms -> 1251ms
generalize decided which type variables to quantify by collecting the free variables of every scheme in the environment and quantifying what was not among them. That is the textbook definition, and it costs O(environment) per binding — so checking N bindings cost O(N²).
Nobody noticed while the compiler only ran once per file. The LSP shipped in v0.1.207 re-checks the whole document on every keystroke, which turned a cost nobody paid into 5.3 seconds of latency per character on a 22k-line file. The profile named one function.
The fix is levels (Rémy's ranks): each type variable records how many generalizable bindings it was created inside, unification lowers that number when a variable escapes into an outer type, and generalization quantifies exactly the variables still deeper than the binding — no environment scan at all. The level field on Ast.tyvar is the whole representational cost.
The old definition is kept as an oracle. MERE_LEVEL_CHECK=1 computes both answers at every generalization and reports any disagreement (add MERE_LEVEL_CHECK_TRACE=1 for a call stack). They agree on all 2421 tests and on 390 further files — contrib, examples and the dogfood repositories. Since the quantified set is the only thing this change can affect, that comparison is the correctness argument, and the switch stays so the next change to the level discipline is checked against the definition it replaced.
It found three real defects while being written, none of which any test caught:
- `check_pattern` ran outside the binding's level, so the fresh variables it
makes for a tuple pattern's components dragged the value's variables out with them: let (f, g) = (fn x -> x, fn x -> x + 1) in f stopped being polymorphic. pp_ty prints ('a -> 'a) either way, which is why the existing test passed.
- `trait_elab` has its own copy of the declaration loop and did not get the
level discipline, which cost polymorphism for every binding in a program using traits.
- The value restriction's monomorphic path left variables looking local to
the binding they had just escaped, so the next binding out would quantify them — the one direction of this change that would have been unsound.
The declaration loop now exists in six copies (three in pipeline, one in trait_elab, two in the tests). Each had to be found and fixed by hand here.
v0.1.219 — 2026-08-12
_print_bytes on Wasm, which makes it all four backends._
interp: 41004228290a
C: 41004228290a
LLVM: 41004228290a
Wasm: 41004228290a
Wasm was the one backend that had to refuse this, and the reason was the host boundary rather than codegen: its printing goes through an env.print_no_nl(ptr) import that reads a NUL-terminated string out of linear memory, which is exactly what a byte sequence cannot be. So it needed a new import taking a pointer and a length — env.print_bytes(ptr, len) — and the host side in scripts/run_wasm.js to write that many bytes from memory.
Gated on use, like the imports around it: a program that does not call print_bytes declares nothing new and runs on an older host unchanged. That mattered here, since the playground ships prebuilt .wasm files.
_The test names both halves: the import must appear when the builtin is used, and must not appear when it is not._
v0.1.218 — 2026-08-12
_ByteBuf[R]: the mutable byte buffer that was missing._
_v0.1.216 gave bytes a way out of the program. What it still had no answer for was building or editing one: bytes is immutable, StrBuf appends only and is text, and Vec[R, int] does the job at eight bytes per byte. The thing that asked for it was reconstructing a PNG scanline, which reads the row above it — already reconstructed — and writes the row it is on. Random access both ways, and bytes._
bytebuf_new : int -> ByteBuf[R] n zeroed bytes
bytebuf_len : ByteBuf[R] -> int
bytebuf_get : ByteBuf[R] -> int -> int
bytebuf_set : ByteBuf[R] -> int -> int -> unit
bytebuf_push : ByteBuf[R] -> int -> unit appends, growing
bytes_of_bytebuf : ByteBuf[R] -> bytes freeze a copy
bytebuf_of_bytes : bytes -> ByteBuf[R]
Region-bound like StrBuf, and for the same reason: the bytes live in a region and the region is tracked by a pointer inside the struct rather than by the marker. Freezing copies into the current region, so a bytes frozen inside region R { } can be returned out of it — the mistake strbuf_to_str had to fix once already.
interp + C, which is where byte I/O lives.
_Measured on the dogfood, decoding a 736×724 RGBA PNG: peak RSS 164MB → 117MB, with byte-identical output. The reconstructed image is 2.1MB of bytes, which was 17MB of ints._
_Adding the type re-found v0.1.217's P5 immediately, and twice — in both directions. Once because ByteBuf was missing from the list of region-parameterised constructors, so the marker erased to int. Then again with the names the other way round, which produced the better fix: a type whose C representation does not depend on its region should not carry the region in its tag at all. StrBuf and ByteBuf both lower to one C type each, with the region tracked by a pointer inside the struct, so StrBuf___heap was never carrying information — only an opportunity to disagree. Both now tag as their bare name, which removes the class rather than the instance, and gives StrBuf the fix for a bug it had never happened to trip._
v0.1.217 — 2026-08-12
_Two C-backend bugs the mpng dogfood found, both of which emitted C that a C compiler rejects._
A `let rec`'s names leaked into every later function. The inner-fn lifting pass keeps a set of names that are not captures — top-level names, builtins, externs, and the siblings of a let rec. The sibling names were added to that set and never removed, so a later top-level function whose parameter had the same name had it taken for one of them: not recorded as a capture, and the lifted body then referred to an identifier nothing declared.
The line that was supposed to put the set back read let _ = known_before in, which does nothing. known_before was already being taken two lines above.
_What it took to see it: png.mere has an inner let rec row, encode.mere has a parameter called row, and the second one lost it. Neither file alone reproduces, which is why this survived — the failure needs two files and a name in common._
An unresolved region marker tagged as `int`. ty_tag erases a type variable that survives to codegen (a dead result, an unconstrained value) to int, on the reasoning that no operation ever inspects such a value. That is true of values and false of a region marker: the marker sits in a type's first slot, the typedef for the same type was emitted from a copy where it had resolved to __heap, and the two spellings — Vec_int_int and Vec___heap_int — are a prototype for a type that does not exist. An unresolved marker now tags as the default region.
_Both were found by compiling the program and running the result, which is what CC_CHECK=1 does in the dogfood and what scripts/ctest.sh does here. Neither is visible on the interpreter; neither is visible from reading the emitted C without a compiler. Both have twelve-line repros in the dogfood's PAIN.md, and regression tests here that name the wrong spelling as well as the right one._
v0.1.216 — 2026-08-12
_bytes gets its I/O boundary: read_bytes, write_bytes, print_bytes._
_The bytes type has existed for a while — bytes_len / get / slice / concat / of_hex / of_str / of_vec and their inverses, across interp, C, LLVM and Wasm. What it had no way to do was leave the program. Reading, writing and printing a byte sequence all went through str or Vec[int]._
_Through str does not work, and the mpng dogfood found out the hard way:_
let _ = print_no_nl (chr 65);
let _ = print_no_nl (chr 0);
let _ = print_no_nl (chr 66);
_The interpreter writes 41 00 42. Compiled with -c, the same program writes 41 42 — a str is a NUL-terminated C string there, so "\0" and "" are the same value and nothing downstream can tell them apart. A PNG decoder writing a PPM produced a file seven bytes short of correct, and only on some backends._
The fix is not in printing. A bytes carries its length in every backend ({ len; data[] } in C and LLVM, an OCaml string in the interpreter), which is what makes these three correct where print_no_nl cannot be:
read_bytes : str -> bytes | interp + C |
write_bytes : str -> bytes -> unit | interp + C |
print_bytes : bytes -> unit | interp + C + LLVM |
_LLVM's writes through write(1, …) rather than fwrite to stdout: reaching stdout from IR means naming a symbol that differs between platforms (__stdoutp, stdout), and a file descriptor is the same everywhere — and unbuffered, which is what the name promises._
_Wasm refuses print_bytes at compile time: its printing goes through a host import that takes a NUL-terminated pointer, so this needs a host-side change rather than a codegen one. Loud, not silent._
_The dogfood now reads and writes through bytes end to end, which is the check that the design is the right one — 28 cases, on the interpreter and compiled, producing identical files._
_Also noticed while testing, unrelated and unfixed: exit 0 as a program's trailing expression breaks the LLVM backend (unsupported LLVM codegen type element: 'a, since exit : int -> 'a)._
v0.1.215 — 2026-08-12
_mere -ll -g: the LLVM backend, too. Every backend can now be debugged as Mere._
mere -ll -g app.mere > app.ll && clang -g app.ll -o app
lldb app -o "b twice"
# Breakpoint 2: where = app`twice at app.mere:2:1
The same destination as the C backend's #line — a DWARF line table naming the .mere — reached the most directly of the three, because LLVM IR carries debug information itself. There is nobody to divide the work with: a DISubprogram per function, a DILocation, and a !dbg on the instructions.
On every instruction, which is the constraint that shapes this. A function with a subprogram whose calls have no location is something the verifier objects to, so the location is attached in emit_instr — the one choke point every instruction already goes through — rather than at chosen points. All of a function's instructions share one location, the line its body began on, which is the same granularity the C backend arrives at for an entirely different reason.
Debug Info Version in the module flags is not optional: without it the metadata is stripped as being from an older LLVM and the debugger shows nothing, with no error anywhere to explain why. It has a test of its own for that reason.
`sh scripts/debug_info.sh` compiles a program through both backends and asks lldb where each function is, because a breakpoint resolving to app.mere:8 is evidence and emitted text looking right is not:
ok C both resolves to app.mere:8
ok LLVM both resolves to app.mere:8
_That closes Q-021, and with it every backend: C #line, LLVM !dbg, Wasm a source map, RV32I its own debug map — each reaching the same place by whatever route its output allows. The interpreter needs none, being where the source already is._
v0.1.214 — 2026-08-12
_mere -wg, and a source map: the browser's debugger shows Mere source._
mere -w app.mere > app.wat
mere -wg app.mere > app.map.txt
wat2wasm --enable-tail-call --debug-names app.wat -o app.wasm
node scripts/wasm_sourcemap.js app.wasm app.map.txt app.mere
Writes app.wasm.map and appends a sourceMappingURL custom section, which is what Chrome and Firefox look for. The playground runs on this backend, so this is the debugger a Mere program in a browser has been missing.
The compiler cannot produce the map, and that is not a limitation but a fact about the format. A Wasm source map addresses byte offsets in the assembled binary; this backend emits text for wat2wasm to assemble. So the work splits the way the RV32I debug map splits: mere -wg says which function came from which line, the binary says where each function ended up (its name section), and a script joins them by name. Whoever knows the addresses is not whoever knows the source.
The check is the interesting part. A source map is easy to produce and hard to trust — the segments are VLQ deltas, so an error in one shifts every mapping after it, and the result still looks like a source map. So scripts/wasm_sourcemap.sh decodes the map back and compares it against wasm-objdump:
ok 0x001550 is <both>, and the map says line 8
ok 0x00155c is <thrice>, and the map says line 5
ok 0x001564 is <twice>, and the map says line 2
_Prelude functions are absent from the table, by the rule v0.1.212 established: a position that names a file did not come from the source being compiled. And the binary still validates after the section is appended, which the check also confirms — appending to a Wasm file is only harmless when it is done right._
v0.1.213 — 2026-08-12
_Find references, and rename._
_The same question — where else is this binding — and the difficulty in both is shadowing: two xes in one file may be two different things, and treating them as one is a rename that breaks the program._
let x = 1; // renaming this one touches
let f = fn (n: int) ->
let x = n + 1 in // ... not this one, nor
x + x; // ... these
let _ = print_int (f x + x); // ... but these two
So the walk resolves every occurrence to the binding it refers to, and the answer is the occurrences that resolved to the same one — the reverse of what go-to-definition does, and the one shape Query was missing. Binder positions are included, so the cursor may be on the definition rather than on a use.
Rename refuses what the file does not own. A prelude name or a builtin has its definition somewhere the edit cannot reach, and renaming the uses while leaving the definition is worse than refusing. The refusal is returned from prepareRename, which is where an editor asks before offering a box to type in — so it arrives as a message rather than as a broken file.
_That is the LSP's list done: diagnostics, hover, definition, completion, outline, formatting, semantic tokens, references, rename. What is left is deliberate — incremental sync (nothing to gain yet), and the twenty typer raise sites whose worst case is one error per declaration rather than all of them._
v0.1.212 — 2026-08-12
_mere -c -g: a debugger on the compiled program shows the Mere source._
mere -c -g app.mere > app.c && clang -g app.c -o app
lldb app -o "b mu_twice"
# Breakpoint 1: where = app`mu_twice + 8 at app.mere:2:40
_Verified with lldb and dwarfdump, not by reading the emitted text: the line table names app.mere, and a breakpoint on a function resolves to the line it was written on._
This was reported as "not mechanical after all" in the notes for v0.1.202, and the reason given was that codegen_c is expression-oriented and does not know which output line it is on, so #line — which applies to the next line — cannot be placed. That turned out to be looking at the wrong thing. A function's whole body is emitted as one C line, so a directive per function is not a coarse approximation but the finest granularity the output has; put it inside the braces and the body lands on exactly the line the programmer wrote. No line-tracking writer, no second pass.
The other half of the problem is what to say about the code that has no Mere source — the runtime, and the prelude. Each user function is followed by a directive naming a file that does not exist (<mere runtime>), so a debugger shows no source for those frames, which is the truth, rather than an arbitrary line of the user's file.
The rule for "is this the user's code" is one line, and it is the one v0.1.210 made possible: a position that names a file did not come from the source being compiled. Imports were already stamped; the prelude is now tokenised as <prelude>, so neither can be claimed. An earlier attempt counted prelude declarations instead and was wrong — Trait_elab reorders the list, so the prelude's are not the first N by the time codegen sees them.
_Off by default: without -g the emitted C is byte-identical to what it always was, which the suite checks. LLVM (!dbg) and Wasm (source maps) remain unanswered; the RV32I backend has had its own since v0.1.200._
v0.1.211 — 2026-08-12
_Formatting, an outline, and colour that is not guessing._
`textDocument/formatting` runs the function mere fmt runs. That is the whole point of it living in Pipeline rather than in the CLI: format-on-save and the command line cannot come to different conclusions about what formatted means. It declines twice, deliberately — a file that does not parse is left alone, because replacing a buffer with the best guess of a parser that failed is how somebody loses work, and an already-formatted file produces no edit rather than an edit that changes nothing. It also re-adds the trailing newline the CLI's print_endline supplies, without which format-on-save would strip it from every file, every time.
`textDocument/documentSymbol` lists the file's value declarations for the outline, telling a function from a value by its type. type declarations are absent and honestly so: Top_type carries a name and its variants and no position, so it cannot be pointed at without guessing.
`textDocument/semanticTokens/full` is the compiler saying which names are parameters, which are functions, which are constructors. Syntax highlighting is normally regular expressions guessing at a language; this one does not have to guess. The editor's grammar keeps what it is good at — keywords, strings, numbers — and the distinction it cannot make, a parameter from a global, comes from the tree.
_The encoding is five integers per token and every one is relative to the token before it, which is compact and unforgiving: wrong deltas paint the file at an offset. The test decodes the stream back into positions and names rather than asserting on the numbers._
_The VS Code extension needed no change for any of this — it asks the server what it can do during initialize, so three new capabilities simply started working. That is the argument for keeping the two apart, arriving on schedule._
v0.1.210 — 2026-08-12
_Positions know which file they came from, and the typer reports more than one problem per declaration._
A position carries its file. Loc.t gained file : string option, and since the lexer is the only thing in the compiler that builds a position, stamping the tokens of an imported file was a one-line change that everything downstream inherits: an error raised deep in the typer, about a node that came from another file, now knows which file it is about without anybody having threaded that through. The CLI renders the snippet from that file; the language server publishes against that file's URI. Before this, only syntax errors could say where they came from — by the time the typer runs, imported declarations have been merged into one program.
The typer collects. The two sites that account for nearly every real type error — a mismatch in unify, and an unknown name — report and carry on instead of raising, when a sink is installed. So one declaration can report four problems rather than the first one.
What it carries on with is the interesting choice: a fresh type variable, which unifies with anything, so it neither invents a second error nor silences a real one further along. A distinguished error type would be the textbook answer and would have to be taught to every match on ty in five backends. On a mismatch the two types are left unlinked, since neither is more right than the other and forcing one on the other is how one mistake becomes five.
The other twenty raise sites are unchanged: they end that declaration's check, and v0.1.209's recovery picks up at the next one. The compiler's path is untouched — no sink, same first-error-raises behaviour — and the sink is installed under Fun.protect, because one left behind would make the compiler collect errors instead of stopping. There is a cap of a hundred, because a pathological file can produce errors without end once inference is allowed past them and nobody is reading the hundred and first.
v0.1.209 — 2026-08-12
_More than one type error at a time._
_A file with three broken functions reported one of them: fix, recheck, learn about the next. The check now recovers at declaration boundaries — the same boundary the parser recovers at — so it reports one error per broken declaration, in the editor and in the terminal._
type error: expected `int`, got `str`
--> app.mere:1:28
type error: expected `int`, got `str`
--> app.mere:2:24
2 errors
A declaration that failed still binds its names, to a fresh type variable that unifies with anything. Otherwise every later use of the name is a second error about the same mistake and the real ones are buried — the test for that uses a broken function twice and expects exactly one error.
The compiler's path is unchanged. Recovery is opt-in (infer_program ?on_error): without it the first error is raised exactly as before, because a compiler that carries on past a type error has nothing useful to emit. One code path, two behaviours, rather than a second implementation to keep in step.
_The typer still stops at the first problem within a declaration. Making that collect means teaching every one of its 22 raise sites to produce a value and carry on — a different and much larger change, and one that needs an error type that unifies silently, or every recovery invents cascades of its own._
_One wrinkle worth recording: the pass over the desugared program re-visits every declaration's body, so it re-raises the error the declaration loop already reported. Diagnostics are de-duplicated, which is what makes that harmless — and is the same fix the duplicated exhaustiveness warnings needed in v0.1.208._
v0.1.208 — 2026-08-12
_Diagnostics become data: which file a position belongs to, and warnings too._
A syntax error inside an `import` is reported against the file it is in. Its line numbers describe that file, so reporting it against the importing one was underlining an innocent line. Parse_error_in_file carries the path from the import that raised it, the CLI renders the snippet from that file, and the language server publishes against that file's URI — remembering which other files it has spoken about so it can clear them when the import is fixed. A diagnostic stays on an editor's screen until the server says otherwise, and "never mind" is exactly the message nobody thinks to send.
Warnings are diagnostics now (severity 2 in the protocol): a non-exhaustive match, a top-level name that collides with a C keyword. They were printed to stderr from inside the pipeline, which is fine for a terminal and useless to anything else — an editor cannot underline a line written to a stream it is not reading. The pipeline collects them; the CLI prints them, which is where the decision about how a warning looks belongs.
Two things that fixing this exposed. The non-exhaustive-match warnings arrived twice, because type inference visits a declaration's body once as a declaration and again as part of the desugared program — de-duplicated now. And the messages carried their own line L, col C: prefix, which read as warning: line 3, col 29: warning: … once a caller with the position added its own; the position is data and the text no longer repeats it.
_Still not carried: type errors from an imported file. By the time the typer runs, the imported declarations have been merged into one program and nothing records which file each came from._
v0.1.207 — 2026-08-12
_Completion — the third of the three questions that are really one question._
_Fifth slice of the language-server arc, and the one that needed no new machinery: Query.scope_at already knew what is visible at a position, so this is that list, de-duplicated by name and dressed for the protocol._
Every name visible at the cursor, innermost first, one entry per name — an inner binding shadows an outer one, and offering both would offer a name that cannot be reached. Each carries its inferred type as the detail line and a kind, so an editor draws a function icon for a function.
Two judgements about what not to offer: the prelude's internal helpers (the ones it names with a leading underscore) are left out, and _ is not a name anybody wants back. Prelude names themselves are offered — str_len is exactly what you want in the list — with sortText putting them after the file's own names.
_That completes hover / definition / completion. All three are Query.node_at and Query.scope_at with a different answer attached, which is what moving the check into the library bought: the editor's three questions turned out to be one question the compiler could already answer._
v0.1.206 — 2026-08-12
_Go to definition._
_Fourth slice of the language-server arc, and the first that needs scope: not just what is under the cursor, but what is bound there and where each name came from._
Scope is recomputed, not indexed. Query.scope_at walks down to the position and collects the binders on the way. The walk descends one path rather than the whole tree, it cannot go stale, and there is no invalidation to get wrong — the same reason hover reads the typer's annotations instead of building a table beside them.
A binder covers the parts of itself where it is really visible: a let binds its body but not its own value expression, a fn binds its body, a let rec binds both, a match arm's pattern binds that arm. Each of those is a test, because getting one wrong is how a server sends you to the wrong x.
Two answers it declines to give, both because the honest answer is nothing:
- A prelude name (
print_int) is genuinely in scope, but its position is a
line in the prelude's own text — jumping there would send the editor to an arbitrary line of the user's file. Prelude bindings are therefore marked rather than dropped (completion will want them), which needed the pipeline to record how many declarations the prelude contributed.
- A parameter resolves to the
fnthat introduced it rather than to the
parameter name, since Fun carries the name but not the name's own position.
v0.1.205 — 2026-08-12
_Hover: the type inference gave whatever is under the cursor._
_Third slice of the language-server arc. Point at a name and the editor shows twice : (int -> int); point at a literal and it shows int._
There is no second inference pass and no index. The typer already writes the type it found onto every node it visits (e.ty <- Some t), so the check that produced the diagnostics leaves behind a tree that knows the answer. Pipeline.check now hands that tree back instead of dropping it, and the server keeps it per open document.
What "the node at this position" means here, since it is not obvious: a Loc.t in this compiler is a line, a column and a width — the token a node was built from, not a span over its subtree. So the node at a position is the narrowest node whose own token contains the cursor. That is what makes hovering inside a call answer about the piece under the cursor rather than about the whole application. Query.node_at is that search, and Ast.children is the one generic child walk it needed — written once, because every position question (this one, go-to-definition, completion) needs it and three hand-written 26-case matches would drift apart.
The last tree that type-checked is kept. While a line is half typed the file does not check, and an answer from a moment ago beats no answer at all — so hover keeps working through an edit and catches up when the file is valid again. The test for this edits a good file into a broken one and asserts the tree survived.
v0.1.204 — 2026-08-12
_mere lsp — a language server. Diagnostics in the editor, from the check the compiler runs._
_The second slice of the language-server arc (the first was recovering from syntax errors, so there is more than one to show). What it does is diagnostics: every syntax error in the buffer, republished on each keystroke, and the first type error once the file parses. Hover and completion want a position resolved against a typed tree, which is the next slice._
mere lsp # LSP over stdin/stdout; see docs/lsp.md for editor setup
The check is the compiler's check. infer_program — parse, elaborate, type, plus the borrow/move/Send analyses — moved out of the CLI into Pipeline, where the server calls the same function the four backends start from. A language server that agrees with the compiler on good days is worse than none: it teaches you to distrust the underline. Pipeline.diagnostics is the one entry point that answers "what is wrong with this text", as data rather than as an exception.
Everything the server decides is a function. Lsp.handle : state -> message -> state * message list * bool — so the protocol is tested in the suite without a socket, a subprocess or an editor, and the only untested part is three lines of IO in Lsp.serve. scripts/lsp_smoke.sh covers the process end to end by piping a canned editor session through the real wire format.
Also new: `Json` — a JSON value, parser and writer (~230 lines), because this project has two dependencies and reading a protocol this small is not worth a third. It decodes \uXXXX into UTF-8 including surrogate pairs, and prints integers without a decimal point, since an editor reading "line": 3.0 strictly is entitled to object.
_Known gaps, all written down in docs/lsp.md: one type error at a time (the typer still raises on the first), positions inside imported files are reported against the importing file, and sync is full-text rather than incremental._
v0.1.203 — 2026-08-12
_The parser no longer stops at the first syntax error._
_A file with three broken functions told you about one of them, three times in a row: fix, recompile, learn about the next one. mere <file> now reports all of them, in source order, with a count at the end._
parse error: expected literal, identifier, or '('
--> app.mere:3:20
parse error: expected type
--> app.mere:7:16
parse error: expected 'ident = expr' after 'with'
--> app.mere:12:15
3 syntax errors
_This is the first slice of a language-server arc, and it is the one that pays off on its own: an editor cannot underline three mistakes if the compiler only knows about one, and neither can a person._
How. Parser.parse_program_recover parses, and on an error deletes the declaration that contains it and parses the whole file again, collecting errors until it succeeds (or hits 20). Re-parsing rather than resuming is deliberate: the parser is functional over an immutable token list, so there is no cursor to reset and no half-built state to unwind — deleting a span and starting over is exact, and it needs no changes to the 130-odd places that raise. It costs one pass per error, which for an editor re-parsing on every keystroke is not the expensive part. parse_program itself is untouched, so nothing on the good path changed.
Where a declaration ends is the interesting part. ; at bracket depth zero is the language's real boundary — but a declaration with an unbalanced ( never returns to depth zero, so a depth-only rule deletes the rest of the file and hides every later error, which is the exact failure being fixed. So a declaration keyword in column 1 is accepted as a boundary too: every top-level declaration in this language's sources starts flush left (the formatter emits nothing else), so an indented let is a local binding and one in column 1 is a new declaration. It is a heuristic, and it is consulted only about where to resume after an error.
_Errors from an imported file are not recovered from — their positions belong to another file's token list, so there is nothing in this one to delete. The first is reported and the walk stops._
v0.1.202 — 2026-08-12
_Two loose ends from the bare-metal work: diagnostics that report the line you wrote, and the differential test the QEMU port was aiming at._
The line you wrote. The -rv path compiles a concatenation — a Mere-source runtime prelude, then the user's file — so every position it produced was counted from the top of that text, and a type error in a three-line file was reported at "line 133", against a snippet from an unrelated line or none at all. The debug map already subtracted the prelude; the diagnostics did not.
One function now answers "where is this really?" for both, so they cannot drift apart. A position that lands inside the prelude is deliberately not remapped into the user's file — there is no honest line there to point at — and is shown against the prelude's own text under the name <rv-prelude>, which also makes a prelude bug legible as one.
Two independent machines, same bytes. v0.1.201 booted a bare program and a trap handler on QEMU's virt board. Two additions finish the thought:
examples/riscv_virt_sched.mere— the context switch on virt. This is the
case most worth an outsider's opinion: the trampoline saves 31 registers to a known place and the emulator restores them, so if the two agreed on a wrong order, order-dependent corruption would be invisible to every test we own. It also checks the rule that was hardest to arrive at (gp switches with the rest, because each task has a heap of its own) against a machine with no stake in it.
scripts/qemu_virt.shnow takesMEMU=<memu checkout>and runs each image on
both machines — QEMU and the Mere-written emulator — diffing the two. All three examples are byte-identical on both.
For that diff to mean anything the output has to be a function of the program rather than of the clock, so the scheduler prints one letter per switch rather than one per N iterations: virt gives each task 20ms of real time, our emulator counts instructions, and both print ABABABA.
_The emulator side of this is in the memu project: ./rvrun 8 virt places RAM at 2GB and moves the CLINT to virt's addresses. What had to change there was the decode order — it asked "is this a device?" first, which is only right while the devices are above RAM._
_Still not on virt: the shell (its input would have to be piped in to be diffable) and the user-process pair (a second image needs -device loader rather than -kernel)._
v0.1.201 — 2026-08-11
_The backend's output boots on a machine nobody here wrote: QEMU's virt board._
_Every layer of the bare-metal work is self-written — the compiler, the fifth backend, the kernel, and the RV32I emulator it runs on. So when something misbehaves, "is the binary wrong or is the emulator wrong?" has no answer inside the stack; agreeing with yourself is not evidence. QEMU is an independent implementation of the same specification, which is what the Klaus and Blargg suites are for the 6502 and Game Boy emulators in the sibling project._
mere -rv --bare --load-base 0x80000000 --ram 8 examples/riscv_virt_hello.mere > virt.bin
qemu-system-riscv32 -M virt -bios none -nographic -kernel virt.bin
_Two programs boot: riscv_virt_hello.mere (UART, a run-time-allocated string, recursion, the CLINT read back) and riscv_virt_timer.mere (a registered Mere closure servicing a real timer interrupt). sh scripts/qemu_virt.sh builds both, runs them and diffs the output; it skips cleanly when QEMU is absent, so this is an optional check rather than a dependency._
_What QEMU checks that our own emulator cannot: instruction encodings against a decoder nobody here wrote, the layout _start builds at a load base above 2GB, the 16550 protocol against a real device model, and — the one most worth an outside opinion — the trap contract: mtvec, mstatus.MIE, mie.MTIE, the CLINT's compare register, and the PC a handler returns for mepc._
_The one thing that had to change in codegen: the machine window's length is now a max rather than a sum. It was mmio_base + mmio_len, which is right only while the devices are above RAM — the arrangement the default base 0 forces. virt inverts it: DRAM at 0x80000000 with every device beneath it, so a program handed [0, 0x10010000) could not name its own RAM. The bounds checks were already unsigned, so a length past 2GB is not a negative number to them._
_Also: the -rv family's flags (--bare, --ram, --load-base) are now parsed in any order rather than matched as literal argument lists, which is why the combination this needed — all three at once — did not exist before. That removes eight arms whose only distinction was which combinations somebody had happened to want._
_Not yet on virt: the scheduler, shell and user-process examples (they name our CLINT addresses; the shell wants a receive side; a second image needs -device loader), and running a virt image on our own emulator, which would make the same bytes runnable on both. Both are address swaps rather than redesigns — see docs/bare-metal.md._
v0.1.200 — 2026-08-11
_mere -rvg: a debug map, so a program compiled to machine code can be debugged at the source lines it was written as._
_Nothing in any backend emitted debug information — no DWARF, no source maps, no line directives — so "which line is this?" was unanswerable everywhere. On the RV32I backend that showed up as a working method: every hard bug in the bare-metal arc was found by instrumenting the emulator by hand, a ring buffer of program counters here, a store watchpoint on a save area there, a register dump at trap entry. Ten instruments, written and thrown away, and once by patching codegen to print two registers from inside __oom._
_The map is a text sidecar, because the binary has no header to hold anything — this backend emits code and nothing else. One record per line, addresses ascending:_
S <addr> <name> every label
F <addr> <name> fsz= ra= fp= params= line= a function and its frame
L <addr> <line> <col> the statement starting here
_Two properties it was worth designing for. It is emitted from the same item list the assembler consumes, via a zero-width Meta item that the assembler and the listing both ignore — so -rv and -rvg agree by construction, there is no separate debug build, and the map describes the bytes that actually ran. And the line numbers are the ones the programmer wrote: source positions arrive counted from the top of the prelude-plus-source text the driver builds, and the map subtracts the prelude, so an address whose line lands inside the prelude gets no record — the honest answer for code nobody wrote. (That offset is the same one that makes -rv diagnostics report line 133 for a three-line file; fixing the diagnostics is a separate change.)_
_Frame layout is uniform on this backend, so fsz / ra / fp describe it completely and a backtrace is two loads per frame with no guessing._
_The reader lives in the memu project as riscv-dbg: breakpoints on source lines, a backtrace, and reverse stepping through an undo log — one fixed-size record per instruction, so going back applies the inverse rather than replaying from a snapshot. It is exact, and tested as such (N instructions forward and N back restore every register and a checksum of the heap), and it crosses traps: S walks out of an interrupt handler and onto the line the timer interrupted. Which is the shape of the thing this arc kept needing and building by hand._
_Three tests on the map. dune runtest 2339/0, ctest 13/13._
v0.1.199 — 2026-08-11
_Documentation for the bare-metal work, which existed only as twelve changelog entries and seven example headers._
_The arc built a fifth backend, an operating system on it and a user process on that, and none of it was discoverable: the README did not mention RV32I at all, codegen.md documented three backends, and the nine new examples were missing from the category index. Someone arriving at the repository could not find the tower, let alone the rules for using it._
_docs/bare-metal.md is now the one place: flags, the memory map, raw memory as a window capability and the three ways out that are closed, CSRs and why they are deliberately not a capability, the trap trampoline and its two non-obvious properties, tasks, and when to switch `gp` — the rule that took the arc's hardest bug to find. It ends with what is deferred on purpose (the fantasy console's ambient framebuffer, a QEMU boot for external verification, nested traps, and the absence of an MMU) so the gaps are recorded rather than implied._
_Also stated plainly there, because it would otherwise be easy to overclaim: what isolates the user process is the type system, not the hardware. Everything runs in machine mode; the process is contained because without --bare it cannot obtain a Raw at all, not because an MMU would stop it._
_The examples index gains a section in the same shape as the browser apps — each example beside the thing it forced — and a broken link found on the way (contrib/json/writer.mere, merged into json.mere in 31b4c45) is fixed where this file referenced it._
v0.1.198 — 2026-08-11
_The allocating-handler corruption, solved. The mechanism was none of the three suspects — it was region, and the "fixed" shell had been quietly broken all along._
_The tell was in the emulator's register log: task0's gp moved backwards while it ran — the shell's per-command region R { ... } rollback. The rest follows. The region parked the bump pointer, a timer switch let the background task allocate its loop closure above the mark, and the rollback freed it — live — for the next command to overwrite. The task then resumed with a0 pointing into reused memory and jumped through whatever now sat at its closure's first word._
_Sharing the bump pointer between tasks was the bug. v0.1.192's rule ("gp is machine state; never switch it") missed the other direction: with a shared heap, anything that rolls the pointer back frees what the other context allocated meanwhile. The rule that survives contact is the user-process one, applied inside a single program: contexts share `gp` only if they genuinely share a heap, and a context that uses regions must not. The scheduler and shell examples now give each task an arena carved from machine_scratch — heap up from the bottom, stack down from the top, so the out-of-memory check guards each task for free — and switch every register. raw_len (the partner of raw_base) went in so a kernel can partition a window it was handed without hardcoding the runtime's geometry._
_And v0.1.193's fix had fixed nothing. Making the handler allocation-free moved the corruption out of sight, not out of existence: in the shipped shell the background task's counter froze a few commands in and never advanced again — the task was dead, resuming into reused memory every slice, and the fault-stepping handler swallowed the evidence. The falsification test was the counter, probed between heavy commands: 864, 864, 864. With per-task heaps it climbs monotonically, and the original repro passes with the handler allocating — the rule "a trap handler must not allocate" is back to being good practice rather than load-bearing._
_Two adjacent holes found on the way, both real:_
- _The dedicated trap stack (v0.1.197) sat below
machine_scratch, so a
handler allocating while an arena task was interrupted compared a high gp against a low sp and declared the heap exhausted, spuriously. The stack now sits above the arenas, where the same check instead protects it: an arena that grows into the trap stack is refused._
- _The runtime's abort paths (
__oom,__raw_fault,__pat_fail) report and
exit via ecall — which, with a kernel installed, vectored to the program's own handler; a handler that steps over faults swallowed both ecalls and execution fell off the end of the helper into whatever was emitted next. They now take the machine back (csrrw x0, mtvec, x0) before reporting: a dying runtime owes the program nothing, but it owes the person at the terminal a message._
_All six bare-metal examples verified, the background counter climbing, the self-hosted compiler still byte-identical under its kernel. dune runtest 2336/0, ctest 13/13._
v0.1.197 — 2026-08-11
_A third hypothesis for the allocating-handler corruption, also disproved. The change it prompted is worth keeping anyway: the trap handler gets a stack of its own._
_Until now the handler ran on whichever task's stack it interrupted. That is a design smell independent of any bug — it makes the handler's frame size a constraint on every task's stack, and it means a task with a nearly-full stack turns any trap into a memory-corrupting event. The trampoline now switches sp to a dedicated 8KB stack in the reserved region once every register is safely saved, which is what a kernel does and for these reasons. machine_scratch starts above it, so task stacks are unaffected except for being 8KB smaller._
_It does not fix the corruption. Three mechanisms are now ruled out — the header-before-bump window (v0.1.193), trampoline reentrancy (v0.1.195), and the handler's stack placement (here) — against a signature that is precise:_
- _a task resumes with
a0holding a pointer to the trap save area's window
block — a two-word `Raw` value that only the handler ever constructs;_
- _the next closure tail call reads word 0 of it as a code pointer, which is the
save area's base, and jumps there;_
- _from then on the machine executes its own saved registers as instructions._
_Every write to that a0 slot comes from the trampoline's own save instruction, so the value was in a0 at trap entry, meaning the interrupted code held it — and the only code that holds it is the handler. Which would be reentrancy, which the depth check says is not happening. One of those two statements is wrong and finding out which is the next probe: log the first forty writes to the slot rather than the last, and see the value's first appearance instead of its aftermath._
_Recorded rather than guessed at. The rule stands and every example keeps it: a trap handler must not allocate. All six bare-metal examples verified (uart, timer, sched, shell, user, selfhost — the last still byte-identical to the interpreter), dune runtest 2336/0, ctest 13/13._
v0.1.196 — 2026-08-11
_The self-hosted Mere compiler, running as a user process on a Mere kernel, on a CPU written in Mere._
kernel: running the self-hosted Mere compiler as a user process
(module
... 5,224 more lines of WAT ...
kernel: user process exited after 3 syscalls and 3727 ticks
_The WAT is byte-identical to what the native interpreter emits for the same input. It reached the UART through kernel write syscalls, from a process the timer preempted 3,727 times along the way._
_Nothing new was needed. The compiler image is examples/riscv_user_selfhost.mere — the contrib self-hosted compiler asked to compile let x = 10 in x * x + 1 and print the result, with no idea it is a user process. The kernel is examples/riscv_bare_selfhost.mere, which is v0.1.194's kernel with a bigger tenant: 24MB for the image, because the compiler's heap peaks between 14 and 18MB and this backend's allocator never frees. That measurement, made back in v0.1.186, is why --ram exists._
_The tower, bottom to top: a language; a backend of its own that emits RV32IM; a CPU written in that language to run it; a kernel written in it too, with traps, a timer, a scheduler and a syscall boundary; and the language's own compiler running as a process on that kernel._
_v0.1.147 reached the third floor of that and called it the north star. This is the fifth._
v0.1.195 — 2026-08-11
_Chasing the allocating-handler corruption from v0.1.193. Two hypotheses tested and disproved, one narrowed to a single instruction, and a permanent diagnostic for the class._
_The repro is deterministic: the shell with its register-copy loops moved back inside the handler (where they are closures, and a closure is an allocation), driven by a 33-command session. It stops at exactly the same byte every time._
_Working backwards with the emulator, which is where this backend's debugging lives:_
- _the guest ends up executing the trap save area as if it were code —
mcause
2 (illegal instruction), the PC marching forward four bytes per trap because the fault path returns mepc + 4;_
- _it got there from
jalr zero, 0(t1)— the tail call this backend emits for a
closure — with `t1` holding the save area's base address;_
- _
t1was loaded two instructions earlier bylw t1, 0(a0), the code-pointer
fetch. So a0 was pointing at the two-word block a Raw window is, not at a closure: word 0 of that block is the window's base, and the window in question is the save area._
_A pointer to the save-area window is a value only the handler ever holds. So some register belonging to the handler ends up restored into a task. That reads like a reentrancy failure — the save area is one global buffer, so a trap taken while the handler runs would overwrite the interrupted context with the handler's own. The trampoline now counts trap depth and refuses to nest, printing what happened and stopping, because with one save area there is nothing left to resume. It is placed after the register-save loop, since checking any earlier would clobber a register before saving it — which is the exact bug it exists to catch._
_And it does not fire on the repro. So the corruption is not a nested trap either. That is worth as much as a positive result: two mechanisms are now ruled out (the header-before-bump window, fixed in v0.1.193 and not the cause; and reentrancy, ruled out here), and the failure is pinned to one instruction with a known-wrong register. What remains unexplained is how a handler-local pointer reaches a task's register file at all._
_The rule stands and is now enforced by construction in every example: a trap handler must not allocate. The nested-trap check stays regardless — it turns a whole class of silent corruption into a sentence, which is this project's usual trade._
_All five bare-metal examples verified after the trampoline change (uart, timer, sched, shell, user), dune runtest 2336/0, ctest 13/13._
v0.1.194 — 2026-08-11
_A user process. A separately compiled, ordinary Mere program running under a Mere kernel, printing through kernel syscalls, preempted by the timer, and unaware that any of that is happening._
kernel: starting a user process at 8MB
user: hello from a user process
user: I do not know a kernel exists
user: fib 20 = 6765
user: exiting
kernel: user process exited after 9 syscalls and 21 ticks
_The user program in examples/riscv_user_prog.mere is not --bare, holds no capability, names no device and touches no CSR. It calls print. That lowers to the same ecall every hosted Mere program on this emulator has always used — what changed is who answers: with mtvec set, an environment call traps (cause 11) and the kernel in examples/riscv_bare_user.mere reads fd, buffer and length out of the register save area and writes the bytes to the UART. Neither the emulator nor the user program is in that conversation. That is the whole idea of a syscall boundary, and it is why the user program needs no cooperation: it asks the machine, and the machine now has a kernel._
_`--load-base <addr>` is what makes two programs fit in one address space. Everything PC-relative in the emitted binary never cared where it lived; the absolute parts did, and they all went through three places — the globals region, the stack top, and the assembler's resolution of la for string literals and lambda entries. They now shift together. The kernel says where its user process goes; the process does not know._
_One thing this got wrong first, and the mistake is the interesting part. v0.1.192 found that gp — the heap's bump pointer — must not be switched between tasks, because two tasks in one program share a heap and switching it makes them allocate over each other. Carrying that rule over to a user process broke immediately: the kernel resumed with the user's gp, its next allocation landed at the user's heap top, and the out-of-memory check compared that against the kernel's much lower stack and correctly declared the heap exhausted._
_So the rule is not about tasks at all. Switch `gp` exactly when the two contexts do not share a heap. Two tasks in one program share one; two programs built with different --load-base do not. Both examples now say so, and say why._
_On the emulator side (memu): ecall traps to mtvec when a kernel is installed and falls back to the host ABI when none is, and a second image user.bin loads at 8MB if the file is there — a bootloader's job, done by the bootloader, since a kernel with no filesystem has to get its first process from somewhere. The instruction fetch is guarded now too, which is what made the previous slice's debugging possible at all._
_A consequence worth noting: any --bare program that installs a trap vector must clear it before returning, or _start's exit ecall vectors into its own handler instead of halting. The timer and shell examples now do._
_Two new tests. dune runtest 2336/0, ctest 13/13, parity 84/84._
v0.1.193 — 2026-08-11
_A shell, on a machine with no operating system. And a rule the previous slice got wrong._
_examples/riscv_bare_shell.mere reads a line from the UART, dispatches a handful of commands, and reports on the machine it is running on: bg twice shows a counter that climbed while the shell sat waiting for a keystroke — nothing yielded to it, the timer took the CPU away and the handler gave it to the other task. fault reads an address the machine does not have, and the same handler that schedules fields the access fault, counts it, and steps over the faulting instruction, so a bad command does not take the machine down. Each command runs inside region R { ... }, which is what keeps the heap flat across a session on a backend whose allocator never frees._
_It needed no new language features. That is the point of the slice, and it is also the signal that this dogfood has stopped generating pressure._
_The correction. v0.1.191 said the trampoline's save-and-restore of gp gives "a region per trap, for free". That is wrong, and this shell is what proved it: a trap handler that allocates corrupts the program it interrupted. Reproducibly — three sessions, three corruptions — and cleanly fixed by moving the handler's two register-copy loops to top level, where they are functions rather than closures and so allocate nothing._
_Two candidate mechanisms were investigated and neither fully explains it. The runtime's variable-size allocators did claim a block and write its header before advancing gp, which leaves a window where an allocating trap lands on a block that already has contents; those are reordered here to bump first and fill after (__str_concat, __str_of_int, __substring, __strbuf_new — the vec allocators were already in the right order). That is a real latent bug and worth fixing on its own, but it was not the whole story: with it fixed the corruption still reproduced. Removing the handler's allocation fixes it; so does not rolling gp back. The honest state is that the interaction between a bump allocator and an involuntary interruption is not yet understood, and the working rule — do not allocate in an interrupt handler — is one real kernels keep anyway. It is now what the example does and why._
_dune runtest 2334/0, ctest 13/13, parity 84/84; the interpreter-vs-emulator sweep is unchanged at 39 / 38 / 8._
v0.1.192 — 2026-08-11
_Preemptive multitasking. Two Mere tasks, neither yielding, on a machine with no operating system._
_A context switch turned out to need no new mechanism. The trampoline already saves the interrupted register set to a known place and restores from that same place before mret, so switching tasks is a memory copy in the middle of a handler: the save area into the outgoing task's TCB, the incoming task's TCB back into the save area, and its PC as the handler's result. The scheduler in examples/riscv_bare_sched.mere is thirty lines of ordinary Mere._
_What a bare program could not do was name any of it, so five primitives went in — all narrowing from the machine capability, so none of them hands out authority the program did not already have. Only coordinates:_
- _
trap_save mach— the save area, so a handler can reach both sides of a switch_ - _
machine_scratch mach— reserved RAM the runtime is not using, which is where
a task stack comes from. A bare program owns no fixed address of its own: the heap grows up from 2MB and the stack down from the top, so any address it picks is one the compiler is already using. The timer example's first draft learned this by sharing a word with a top-level binding._
- _
raw_base w— a window's base as a number, because a stack pointer is an
address and the hardware wants the number. Not authority: touching anything still needs a window._
- _
closure_code f/closure_env f— a task IS a closure, so starting one means
building a context whose PC is its code and whose a0 is its environment. ABI knowledge, which a kernel legitimately has._
_One word must not be switched, and finding out why is the interesting part: gp, the heap's bump pointer. Save and restore it per task and two tasks allocate over each other, each rolling the pointer back to where it stood when it last ran. The heap is machine state, not task state. That is a real collision between this language's memory model and concurrency, and the answer here — share the bump pointer, leave the word alone across a switch — is the simplest one that is correct. A per-task heap would be the other, and it is not needed yet._
_Also: the RV32I backend had print_int but not print_bool, which the parity file added in v0.1.190 immediately caught by being the one file -rv refused. It lowers the way LLVM and Wasm do it, as print of a literal._
_Two new tests. parity 84/84, ctest 13/13, dune runtest 2334/0. The interpreter-vs-emulator sweep of test/parity is 39 identical / 38 refused / 8 mismatching — one better than last slice, since print_int_bool now runs there too._
v0.1.191 — 2026-08-11
_The machine takes traps, and a Mere closure services them. A timer interrupt arrives while the program is doing something else._
_A trap handler cannot be an ordinary function: it is entered with every register live and it leaves with mret, not ret. The tempting move is to give the language a naked fn or an interrupt attribute. That is not needed — codegen already emits _start, so it can emit the trampoline too, and the user writes plain Mere._
_The harder question was how a handler gets the machine capability. It needs one for anything useful (a context switch is a memory copy), and an interrupt has no caller to hand it anything. So the handler is registered rather than named: set_trap_handler (fn cause -> ...) takes a closure, which captures whatever it needs. The trampoline stores it, points mtvec at itself, and calls it with mcause; the result is the PC to resume at. Everything else the handler might want is a csr_read away — mepc, mtval — so nothing has to be packed into a tuple, which would mean allocating inside a trap._
_mscratch holds the save area's address, because at trap entry there is no free register to build one in — which is what that CSR exists for. gp, the bump pointer, is saved and restored with the rest, so whatever the handler allocated is reclaimed when it returns: a region per trap, for free. The corollary, worth stating, is that a handler must not leave an allocated value somewhere that outlives it._
_(v0.1.193: that last paragraph is wrong. An allocating handler corrupts the program it interrupted, reproducibly. The rule is that a trap handler must not allocate at all — see v0.1.193.)_
_On the emulator side traps vector to mtvec with mepc / mcause / mtval set and MIE moved into MPIE. Three causes so far: an unimplemented instruction (2), a load or store past the end of RAM (5 / 7), and the timer (0x80000007). The access faults are an improvement in their own right — an address past RAM used to take the emulator down with a Mere "index out of bounds", reporting the host's problem instead of the guest's. A guest with mtvec still zero halts, as before, rather than jumping to address 0._
_The timer is a CLINT with mtime and mtimecmp in the MMIO region. mtime advances once per instruction: a clock in units of work done, which is what a deterministic emulator can honestly offer and enough for a scheduler tick. The interrupt is taken between instructions._
_examples/riscv_bare_timer.mere arms it and spins. The ticks arrive anyway, which is the mechanism preemption is made of: once a handler runs without the interrupted code's cooperation, a scheduler is a matter of what that handler chooses to return._
_The example's first draft kept its tick counter at 0x200000 and quietly shared a word with a top-level binding — that address is where globals live. A bare program owns no fixed RAM: the heap grows up from 2MB and the stack down from the top. The handler's state is an ordinary Mere cell it captures instead, which is both correct and the better demonstration._
_Three new tests (twenty-one on this backend). parity 84/84, ctest 13/13, dune runtest 2332/0._
v0.1.190 — 2026-08-11
_print_int and print_bool, which the reference has claimed for every backend since Phase 22 and which only the interpreter had._
_Found while smoke-testing the RV32I work: let f = fn n -> let _ = print_int n in n; would not compile on the C backend. Not a shadowing or value-position subtlety — codegen_c had no arm for the name at all, so it fell through to the closure path and emitted a call to an undefined mu_print_int. The refusal then arrived from clang, as "use of undeclared identifier", about a symbol the user never wrote. LLVM and Wasm at least said what they meant ("no lowering yet"), though LLVM's message named its scope as "interp + C", which was not true either._
_So a documented builtin worked in one of four places, and the one that failed worst failed at the wrong layer. All four now agree. C prints through printf (%lld, matching its int width) and puts for the bool. LLVM and Wasm route through the str_of_int they already had, so neither needed a new runtime call or host import — print_int x becomes print (str_of_int x) at emit time. That also needed print_int added to each one's use-scan, since the helper @show_int / $show_int is only defined when something registers it: exactly the v0.1.42 gap, one level further in._
_test/parity/print_int_bool.mere keeps the four in step over zero, negatives and a bool from a comparison. parity 84/84, ctest 13/13, dune runtest 2329/0._
_The reference line that claimed all this now says what actually happened._
v0.1.189 — 2026-08-11
_Machine CSRs, and a non-blocking byte of stdin so a UART can have a receive side._
_CSRs. csr_read / csr_write lower to CSRRS-from-x0 and CSRRW-to-x0, and the register number has to be a literal — it is a 12-bit field of the instruction, so a computed one has nowhere to go, and saying so beats emitting something that reads a register nobody asked for. They are --bare only: a trap vector means nothing under a host. Unlike raw memory they are not behind a capability, and that is a decision rather than an oversight — a CSR has no base and length to narrow, and the hardware already has machine / supervisor / user mode to separate a kernel from a user process. Duplicating that in the type system before there is a user mode to protect would be speculative._
_The disassembler had been rendering every CSR instruction as ebreak, because it only knew inst = 0x73. It now decodes the six CSR forms plus mret and wfi, which matters more than it sounds: this backend's debugging story is reading its own listings._
_A receive side. The plan for the UART's input half was "the emulator reads host stdin with read_key", and that does not work: read_key blocks. A CPU polling a line-status register cannot stop the machine to find out whether a byte is waiting — and once there are timer interrupts it will have other things to do while nothing is arriving. So the dogfood forced a new builtin instead: `stdin_byte : unit -> int`, one byte or -1 when nothing is ready, on the interpreter and the C native backend (select with a zero timeout, so it leaves stdin's flags alone and composes with tty_raw and with a plain pipe alike)._
_With that, examples/riscv_bare_echo.mere echoes what you type through the UART and stops at q — polling the status register, reading the data register, all through the window capability main was handed, with no host syscall anywhere in it. That is the last piece the shell needs._
_Four new tests (eighteen on this backend). parity 83/83, ctest 13/13, dune runtest 2329/0._
_Recorded while here, unrelated and pre-existing: on the C backend let f = fn n -> let _ = print_int n in n; does not compile — the builtin ends up in value position and the emitted C calls an undefined mu_print_int. It reproduces on the pre-session compiler, so it is not from this arc; print on a str refuses at codegen instead. Worth its own slice._
v0.1.188 — 2026-08-11
_Bitwise operators on the RV32I backend, and the INT_MIN bug they found._
_The UART example from the previous slice had to read a status bit with / 32 % 2, because this backend had no bit_and. A device driver is mostly masks and shifts, so that was the next thing in the way. All six lower now: AND / OR / XOR as R-type, or the I-type form when the operand is a small literal, bit_not as xori -1, and bit_shl / bit_shr as SLL / SRA — arithmetic, because bit_shr is documented as floor division by 2^n on every backend._
_Shift counts of 32 or more needed a decision. RV32 shifts use only the low five bits of the count, so a bare SLL would quietly make bit_shl x 33 mean x << 1. What the other backends produce, read back as 32 bits, is zero for a left shift and the sign bit for a right shift, so that is what this emits: folded when the count is constant, and three extra instructions (or a branch) when it is not. A silently different answer was not on the table._
_Then bit_shl 1 31 printed -./,),(-*,(._
_Not a shift bug — str_of_int and print_int both began by negating a negative value to make it positive, and 0x80000000 negated is still 0x80000000. Every remainder after that came out negative and '0' + negative is punctuation. Both helpers now go the other way: make the value negative (negating a positive is always safe) and take each digit as -(x % 10). So INT_MIN prints. This had been latent since the backend's first slice; nothing before now had produced 0x80000000, because nothing before now could shift._
_The bitwise builtins take three files off the interpreter-vs-emulator sweep's "refused" list. One joins the matching set; the other two ask for integers wider than 32 bits (int64_bitwise shifts 1 by 40, and riscv_core is the RV32I emulator itself, which needs headroom above 32 bits before masking), so they cannot match on a 32-bit target and are recorded as such. The sweep now reads 38 identical / 38 refused / 8 mismatching._
_Three new tests. parity 83/83, ctest 13/13, dune runtest 2325/0. The UART example reads its status bit with bit_and now._
v0.1.187 — 2026-08-11
_Raw memory, as a capability rather than an ambient builtin. mere -rv --bare hands a program the machine and it writes to a UART._
_A kernel needs to reach hardware, and this backend's answer so far was one builtin per device with the address baked into codegen: fb_set stores to the framebuffer, key loads from the key register, present ends a frame. Adding a UART, a timer and an interrupt controller that way means a compiler change per device, which is the opposite of what a dogfood is for — the kernel would teach the language nothing. The alternative that fits what this language already says about itself (README: effects are capability values you pass) is to make the address space a value._
_A `Raw` is a window onto physical memory: a base and a length. It is opaque, nothing constructs one, and there is no function that mints one — the only source is the argument --bare hands to the program's top-level main, and raw_window can only narrow. Offsets are relative to the window, so a driver holding a UART window cannot express an address outside it. All three ways out are closed and each fails differently: forging one from ints is a type error, widening one faults at construction, and an offset past the end faults at the access._
_So this reads as a promise rather than a hope:_
let putc = fn (uart: Raw) -> fn (c: int) -> raw_poke8 uart 0 c;
_putc can touch the UART and nothing else — not the heap, not the stack, not another device — and that is visible in its signature instead of being a claim about its body and every body it calls. That is the whole argument for a value over an ambient builtin._
_The bounds check is not free and not optional: a window's length is a runtime field, so there is nothing to fold at compile time even when the offset is a literal. Three instructions on an MMIO poke buys a guarantee that holds, which is the better trade than a nominal one._
_Device MMIO now sits above any RAM — the UART's data register is at 0x10000000, the address QEMU's virt machine uses, so a driver written today is not inventing a private convention. --ram is capped so RAM can never reach it, which means a device address does not move when the RAM size does. (The fantasy console's framebuffer and keys predate this and still live in the reserved top of RAM; they move with --ram and will migrate when that demo is next touched.)_
_--bare also refuses print / print_int / print_no_nl / print_err: those lower to the emulator's write syscall, which no real machine answers, and a bare program that depends on the courtesy would break the moment it left the emulator. A UART window is three lines away — see examples/riscv_bare_uart.mere, which prints through one and reads the 16550 line-status register back._
_Every other backend refuses these names, because there is no honest physical address in a hosted process. The C backend had to be taught to: without an arm of its own the raw_\ names fell through to the closure path and emitted a call to an undefined `mu_raw_poke8` plus an unknown `Raw` C type, so the refusal arrived from clang as "type specifier missing" — loud, but about the wrong thing. LLVM and Wasm already refused._
_Five new tests, fourteen on this backend now. parity 83/83, ctest 13/13, dune runtest 2322/0; the interpreter-vs-emulator sweep of test/parity is unchanged at 37 / 41 / 6._
v0.1.186 — 2026-08-11
_A RAM size instead of three hardcoded addresses, and the self-hosted compiler runs on the RV32I emulator again._
_v0.1.185 ended by reporting that the self-hosted compiler no longer fits on this backend. The reason it did not fit was not really its appetite: the stack top, the print scratch buffer and the fantasy-console framebuffer were three immediates baked into codegen, which pinned the heap's ceiling at 0x7E0000 and gave every program on this backend the same 5.86MB no matter what it was doing._
_They now derive from one number. The top 128KB of RAM holds the scratch buffer and the MMIO; the stack starts just below that and grows down; the heap grows up from 2MB; everything between the two growing ends belongs to them. At the default 8MB every address comes out exactly where it has always been, so an emulator sized for the old layout needs no change. mere -rv --ram <MB> (and -rvs --ram) raises it._
_With that, the measurement the previous slice could not make: the self-hosted compiler's heap peaks somewhere between 14 and 18MB — it fails at --ram 16 and completes at --ram 20. At 20MB it compiles 1+2 on the Mere-written RV32I emulator and emits WAT byte-identical to the native interpreter. The tower from v0.1.147 stands again, and this time the requirement is written down rather than implied: a binary says how much RAM it wants, an emulator is told the same number, and a mismatch fails loudly at the first allocation past the limit instead of corrupting a frame._
_The emulator side of that (a RAM size argument, and dropping the 100M instruction budget the compiler ran past) lives in the memu project, not here. Two tests lock the layout: the default still puts the stack at 0x7E0000, and --ram 32 moves it to 0x1FE0000. dune runtest 2317/0; the interpreter-vs-emulator sweep of test/parity at the default size is unchanged at 37 identical / 41 refused / 6 mismatching._
v0.1.185 — 2026-08-11
_Three holes in the RV32I backend, found by asking what a program that never returns would need._
_The next dogfood for this backend is a bare-metal kernel: a scheduler, a trap handler, a shell. Every one of those is a loop that never ends, and iteration here is recursion — explicitly, and also under while, which the parser desugars to a tail-recursive local closure. So the first question was how long a recursion this backend can actually sustain, and the answer was: not long, and it does not say so. A zero-allocation tail-recursive counter completed at 500,000 and died silently at 1,000,000; raising the emulator's instruction budget twentyfold did not change that, so it was the stack, not the clock. A while loop, which pays for a closure frame per iteration, died at 300,000._
_Tail calls. A saturated call in tail position now tears the frame down first and jumps, so the callee returns straight to our caller and the stack stays flat. The tail-position bookkeeping mirrors codegen_wasm's wasm_tail_pos (which lowers to return_call): compile_expr clears the flag for every subexpression and the cases whose value IS the enclosing value — if branches, let bodies, match arms, annotations — put it back. Direct calls become j u_f instead of jal ra, u_f; the closure form, which is the shape a local let rec loop = fn ... and every desugared while actually take, becomes jalr x0 on the code pointer. Calls with nine or more arguments keep the old path: args 9+ travel on the caller's stack, which a teardown would drop. The counter now runs 10,000,000 iterations in constant stack._
_Regions. region R { ... } compiled to its body and nothing else — the comment said "no reclamation" and meant it. gp is the only allocation state on this backend, so the fix is the one Wasm already uses on __lang_bump: park it on entry, roll it back at the closing brace. Eight rounds of 100,000 allocations inside a region now complete with a flat heap; the same program without the region reports exhaustion. The body is deliberately not in tail position — a tail call out of a region would skip the rollback._
_A region also could not call anything. vars_in, which decides which top-level functions are reachable and therefore emitted, had no Region_block case, so a function called only from inside a region was never emitted and assembly died with undefined label u_f. Two lines reproduce it: let f = fn x -> x + 1; and region R { f 41 }. It stayed hidden because a region body of nothing but builtins resolves fine. free_vars_of had the same gap, which would have dropped a capture._
_Heap exhaustion. The heap grows up from 2MB and the stack grows down from 0x7E0000 with nothing between them, and when they met the bump pointer overwrote a live frame's return address with whatever it was allocating — the program then jumped into the middle of a string. No message, no exit code, just an emulator reporting a wild load. Every bump now checks bgeu gp, sp and lands on __oom, which prints and exits 3._
_That check immediately reported something. The self-hosted compiler, compiled with -rv and run on the Mere-written RV32I emulator, now says it is out of memory — and with the check patched out it dies the old way, on a wild load, so this is not a regression from this slice. The self-hosted codegen's output has grown about a hundredfold since that demo was first verified (a five-line input now emits ~5,200 lines of WAT, nearly all of it fixed prelude), and building that with ++ on a bump allocator that never reclaims needs far more than the 5.86MB this memory map leaves. Recorded, not fixed: it wants a configurable RAM size and a memory map that separates the MMIO region from the heap's path, which is the same work the kernel needs._
_This backend had zero automated coverage — the region bug was a hard crash that nothing in the suite would have caught. Seven tests now assert on the emitted listing, so they lock the instruction actually chosen (j vs jal, the bump park/restore, the heap check) rather than just that codegen ran. parity 83/83, ctest 13/13, dune runtest 2315/0. A differential sweep of test/parity through the interpreter and through -rv + the emulator: 37 identical, 41 rejected by the backend as unsupported, 6 mismatching — byte-for-byte the same three numbers and the same six files before and after this slice._
v0.1.184 — 2026-08-11
_to_json on LLVM, and a parity harness that counts what it did not check._
_This one is not a dogfood, and saying so is the point. Nothing forces to_json on LLVM: everything native in this repo goes through C, and no app names LLVM. Inventing an app to justify the capability would invert the discipline the browser dogfoods have been run on all week — apps force capabilities, not the reverse._
_What justifies it is coverage, and coverage is a number. scripts/parity.sh treated a clean refusal as a passing row, so a backend that checked nothing said so in a word buried in a line. It now tallies per backend and names the cases:_
unchecked on llvm: 7 of 83 (refused at emit time)
nested_tuple
of_json_composite
...
_Five of LLVM's seven were the JSON family, and that number grows with every test that touches to_json — contrib/schema being a library means future apps using it would go unchecked there too._
_to_json turned out to be show with different literals: the same structural walk, the same per-type functions, the same recursion. So it is not a second emitter but one with a mode — emit_struct_fn ~json:true — which is also the arrangement that keeps them from drifting. The differences are all shape: [..] for a tuple, {"f":..} for a record, "," rather than ", " in a list, a quoted name for a nullary constructor and a single-key object for a carrying one, and null for unit. Option is the one real special case — JSON spells it as the value or null, not as {"Some":4} and "None" — and getting that wrong was the last diff before the backends agreed._
_LLVM's blind spot is 7 → 6, and every remaining case is of_json. That stays deferred: decoding needs a JSON parser written in the target language, C and Wasm each have their own hand-written one, and nothing rides on a third. The number is in the harness now, so the cost of leaving it is visible rather than argued about._
_parity 83/83, ctest 13/13, dune runtest 2308/0._
v0.1.183 — 2026-08-11
_of_json_like: the target type from a witness value, so a decoder can live inside a polymorphic function — and contrib/schema, which is what that was for._
_of_json reads its target type off the call node. At a use site with an annotation that is exactly right; inside a generic helper it is a variable, and the interpreter has no runtime types to resolve it with. The monomorphizing backends managed — with v0.1.182's fix a generic setter compiled and ran correctly on C — so a program could be compiled and not interpreted, which is the wrong kind of split for a language whose parity harness treats the interpreter as the reference._
_A witness closes it. of_json_like : 'a -> str -> 'a takes a value of the target type; the interpreter reads the type off its runtime shape (a record carries its type's name, a constructor gives one through Typer.constructors) and the compiled backends read it off its static type, which is the same variable the result unifies with. of_json_opt_like is the non-crashing form, and the one a generic setter actually needs: it tries a shape and finds out whether it decoded. The witness is never an imposition — replacing one field of a record means holding the record._
_contrib/schema/reflect is the payoff: schema_fields / schema_text / schema_with over any record, from the name-keyed view the compiler already synthesises for to_json. examples/claims had carried a per-record copy of exactly this since v0.1.176 and now imports it, so a form generated from a record declaration is a library rather than a trick in one app._
_Two collection gaps surfaced while wiring it. collect_mono_variant_instances walked fn signatures and not fn bodies, so person option — produced by of_json_opt_like inside a generic setter whose own signature never mentions an option — was never registered and the emitted C named an undeclared struct. And the new collector arms called ty_tag before checking the type was concrete, which inside the generic skeleton it is not._
_A polymorphic record still needs the annotation: a value carries its type's name but not its type arguments, so a witness cannot describe Box[int]. LLVM is unchanged — it has no to_json at all, which is a separate gap._
_test/parity/schema_reflect.mere runs the one implementation over two record types. parity 83/83, ctest 13/13, dune runtest 2308/0, claims browser check 17/17 against the native server._
v0.1.182 — 2026-08-11
_The monomorphization bug, found and fixed. specialize_single_use_local_fns read "one concrete use" as "used at one type"._
_Three slices ago this was a silent miscompile; two ago it became a named refusal; here it is a working program. The cause is one line of judgement in a codegen pass._
_specialize_single_use_local_fns fixes a local polymorphic fn's type in place when its body contains exactly one concrete use of it — a good optimisation, since a fn used at one type needs no multi-instantiation machinery. But find_all_concrete_arrows_in only reports the uses it can already read, and a use sitting inside a polymorphic function is not concrete yet:_
let hold = fn (v) -> (v, v);
let mk = fn (v) -> (hold v, hold 1);
_hold 1 is concrete. hold v is not, and only becomes so once mk is instantiated — at which point it may be a different type. The pass counted one arrow, unified hold's definition with int, and the str instantiation had nothing left to unify with, so it was emitted with the int body. It now also requires that the body contain no unresolved use of the name. False only when a use genuinely cannot be read yet, so a fn that really is used at one type still specializes._
_How it was found is worth recording, because reasoning about it was wrong three times. Instrumenting generalize showed hold correctly generic with one quantified variable. Instrumenting the skeleton collection showed it arriving at codegen as int -> (int * int). Removing each post-typing pass in turn changed nothing. Watching the specific type variable's binding site printed Loc.dummy — which is not a source location at all, and which only codegen uses._
_The payoff: contrib/state/store uses contrib/state/cell again. It had been writing its three one-slot vecs out longhand since v0.1.178 for exactly this reason — a store instantiated at two state types calls cell_new at a type derived from S and at plain int for its token counter — so the module that names the trick can now use it._
_test/parity/poly_helper_fixed_and_free.mere has outlived two expectations (wrong, then refused, now 1s2 on all four backends) and is kept because the shape is easy to break again. parity 82/82, ctest 13/13, dune runtest 2308/0, claims browser check 17/17 against the native server, and all five browser clients rebuild._
v0.1.181 — 2026-08-11
_Two corrections and a better repro. No new capability._
_The monomorphization repro from v0.1.179 used vec_new and read as if mutable containers, the narrow value restriction or the region model were involved. None of them are. Reduced further, hold allocates nothing:_
let hold = fn (v) -> (v, v);
let mk = fn (v) -> (hold v, hold 1); // parameter-derived AND fixed
let (a1, a2) = mk 1 in
let (b1, b2) = mk "s" in ...
_Three neighbouring shapes compile and run — one instantiation of mk, hold called only at the parameter-derived type, and hold used directly at two types — so it is the combination and nothing smaller. test/parity/poly_helper_fixed_and_free.mere now holds that version._
_The search for where the skeleton gets fixed narrowed without landing. generalize is called on hold and instantiate substitutes without mutating, so use sites cannot be reaching back to the definition — and yet by the time codegen takes its pristine clone the skeleton reads int -> (int * int) while mk beside it is still 'a -> (('a * 'a) * (int * int)). Whatever fixes it runs before codegen. Recorded in the test rather than guessed at._
_The second correction is to v0.1.180's memory measurement, and is written into that entry: two samples read as a leak in contrib/store/kvlog, and nine thousand requests show RSS oscillating (4528, 3568, 5840, 7536, 6656, 8272 KB) rather than climbing. That is malloc churn as the region takes and returns blocks. The native HTTP arena is doing its job; there is no kvlog leak to fix, and the item is withdrawn rather than carried._
_parity 82/82, ctest 13/13, dune runtest 2308/0._
v0.1.180 — 2026-08-11
_The native HTTP runtime gets the four externs a middleware stack needs, so examples/claims runs on both hosts from the one source._
_It implemented six. contrib/http's middlewares need four more: http_current_status and unix_time for access_log, http_arena_mark for the per-request arena, http_send_file for static. A server with a middleware stack could therefore be built for the Node host and not the native one — which is half of "the same source runs on both", and the half nobody had checked because the browser dogfoods only ever ran the Wasm host._
_All four are in. The arena checkpoint is the interesting one: the C region is a chain of malloc'd blocks with a bump pointer, so a mark records the block chain and the pointer, and the release frees every block newer than the mark and winds the pointer back. It runs after the response has gone out on the socket, not before, because the body it just wrote lives in that arena. The checkpoint is taken where the request began rather than wherever the handler calls http_arena_mark, so the request line and the body copy — made by the accept loop before the handler is entered — are inside it too. Still opt-in: nothing is released unless a handler asks._
_http_send_file writes the file straight to the socket without it passing through a Mere str, which is the point of that binding — a str is NUL-terminated and a file is not — and tells the accept loop not to write a second response after it._
_Measured on the native build: 116 KB reclaimed per request, with no new blocks allocated on the paths that fit, so the region stays in its first 4 MB. Two thousand 404s in a row moved RSS 3920 → 1920 KB, i.e. down._
_(Correction, made while writing v0.1.181: this entry first read a second measurement — 1920 → 6336 KB over two thousand requests that read the kvlog — as locating a leak in contrib/store/kvlog. Sampling nine thousand requests instead of two thousand shows RSS oscillating rather than climbing: 4528, 3568, 5840, 7536, 6656, 8272 KB. That is malloc churn as the region takes and returns blocks, not a per-request leak, and two samples were not enough to say which.)_
_examples/claims/browser_check.mjs passes 17/17 against the native server, unchanged from the Wasm host. parity 82/82, ctest 13/13, dune runtest 2308/0._
v0.1.179 — 2026-08-11
_The monomorphization gap v0.1.178 recorded without a repro, isolated to seven lines — and made loud, which is as far as this slice goes._
_The shape, self-contained and with no imports:_
let hold = fn (v) -> let c = vec_new () in let _ = vec_push c v in c;
let mk = fn (v) -> (hold v, hold 1); // parameter-derived AND fixed
let (a1, a2) = mk 1 in
let (b1, b2) = mk "s" in ...
_A polymorphic helper called at both a parameter-derived type and a fixed one, inside a function that is itself used at two types. Written (hold v, hold v) it is fine, and used directly at two types it is fine; it is the combination._
_What it produced: hold monomorphized into an int copy and a str copy, both declarations with the right signature and both bodies with the `int` one's operations — a mere_vec_str* function calling mere_vec_int_push. Not a wrong answer, no answer, and only when the C compiler saw it, with a message about pointer conversions naming nothing in the source._
_The mechanism, traced: codegen keeps a pristine clone of each polymorphic skeleton and unifies a fresh copy with every instantiation's arrow. That unify was wrapped in try ... with _ -> (), and when it fails the spec's body belongs to another type. It cannot fail while the skeleton is genuinely polymorphic — and here it is not. Instrumenting the clone showed the typer hands codegen hold already fixed at int -> Vec[__heap, int], while mk beside it is still 'a -> .... Written (hold v, hold v) the skeleton arrives polymorphic._
_All three compiled backends now refuse instead of emitting, with a message that names the function, the arrow it cannot take, and the type it is stuck at. That turns a wrong program into a named one; it is not the fix. The fix is upstream of codegen, in whatever fixes the skeleton before it is cloned, and test/parity/poly_helper_fixed_and_free.mere will change from UNSUP to a real answer when that lands._
_This is why contrib/state/store writes its three one-slot vecs out longhand instead of using contrib/state/cell: a store instantiated at two state types calls cell_new at a type derived from S and at plain int for its token counter. The module that names the trick does not get to use it, and now says why. parity 82/82, ctest 13/13, dune runtest 2308/0._
v0.1.178 — 2026-08-11
_A page with two views, so that a watcher's lifetime becomes a question — and the answer that drop type cannot give._
_examples/claims grew a tab bar over one store: a Form view that builds the generated controls, the line rows and the actions, and a read-only Summary view. Switching drops the old view's DOM wholesale. Dropping the nodes did not drop the watchers, because contrib/state had no way to take one back, so three round trips between the tabs left eleven watchers running, nine of them painting into nodes no longer in the document. The counter in the tab bar is how that became visible at all._
_store_watch now returns a token and store_unwatch takes it; the app collects a view's tokens and releases them on close. 11 → 2, and constant however far the user navigates._
_The language question this was chosen to ask has a negative answer, which is worth having. Mere's one mechanism for enforced release is a drop type bound by with, whose close runs at scope end — and a subscription's lifetime is not the scope that created it. The view is built in one event and torn down in another, so by the time with would close the handle the view has not even been shown. A drop type also cannot be placed in a region, which is where the watcher list lives. So releasing stays a convention, recorded in contrib/state's README rather than papered over._
_Two compiler findings on the way. A binding called entry did not compile on LLVM at all: values and basic-block labels share one namespace there and every emitted function opens with a block called entry, so a parameter of that name claimed the slot first — "unable to create block named 'entry'". Not a wrong answer, no answer, from a name nothing warns about; entry is the obvious name for an element of an association list, which is how it turned up. llvm_safe_local renames the parameter, since the entry: label is written into every hand-authored runtime blob in that file. test/parity/reserved_local_entry.mere holds it across the shapes that emit a parameter separately — top-level fn, lifted inner fn, closure adapter._
_And the store could not use contrib/state's own cell: a store instantiated at two different state types puts cell_new at three types derived from S, and that shape does not survive monomorphization on C or LLVM. The builtin vec does, so the module that names the trick writes its three slots out longhand. Recorded in the source; a minimal repro is not yet isolated._
_parity 81/81, ctest 13/13, dune runtest 2308/0, claims browser check 17/17 — including that three round trips between views leave no watchers behind._
v0.1.177 — 2026-08-11
_Stage 4 of the shared-schema dogfood: change the schema in ways other than adding a str, and see what the derived machinery does. It found a hole in the mechanism and a soundness hole underneath it._
_The mechanism first. claim_with took the JSON shape to write back from the shape already there, which is exact except for an absent optional — null cannot say what it would have held. The first version guessed "string" and named the two fields that needed clearing, and adding a seat: int option to the record broke setting it: the guess produced a string and the decode failed. Where the value cannot say, ask the decoder instead — try the shapes in order and keep the first that survives of_json_opt. Blank text tries null first, which is how an optional is cleared without anything knowing which fields are optional. All four cases now behave: an optional str and an optional int clear to None, a required str blanks to "", a required int to 0, and the hardcoded field names are gone._
_Underneath it: the probe set every field from text, including the variant-typed status, and "7" was accepted. The interpreter's decoder built a nullary constructor from whatever string arrived without asking whether the variant had a case by that name, so of_json_opt returned Some holding a value outside its own type, which to_json then printed straight back out. Both compiled backends answered None — the reference the parity harness compares against was the one in the wrong, which is why nothing had caught it. The object form was checked for existence but not ownership, letting a payload case of an unrelated variant through; constr_info.type_name closes both, and a case of this variant arriving in the wrong shape ("Rejected" as a bare string when it carries a payload) is refused too._
_test/parity/of_json_variant_tag.mere holds it and fails without the fix. It took a dogfood that round-trips a variant field through a form to surface, which is the same shape as v0.1.175's finding: the boundary bugs live where a value is only ever decoded. parity 80/80, ctest 13/13, dune runtest 2308/0, claims browser check 14/14._
v0.1.176 — 2026-08-11
_Stage 2 of the shared-schema dogfood: the form is the record._
_The plan was to let the app pick the mechanism rather than choose one in advance, and the app picked one that needs no language change at all. A record already has a total, name-keyed view of itself — the one the compiler synthesises for to_json. Reading it back gives the field names in declaration order, and replacing one key and decoding gives a setter. claim_fields / claim_text / claim_with are that, in about forty lines of schema.mere, and the controls on the page are built from the list at startup. index.html now has a <div id="form"> and no field markup._
_Measured the same way as stage 0, by adding a cost_center: str field:_
| stage 0 | stage 2 | |
|---|---|---|
type claim | edit | edit |
blank_claim | compile error if omitted | compile error |
a rule + claim_problems | silent | derived |
| the field table in app.mere | silent | gone |
| index.html markup | silent | generated |
_Two edits, in one file, and the compiler forces the second. Silent sites 3 → 0. Checked rather than assumed: adding the field and changing nothing else puts a control labelled "Cost center" on the page and round-trips its value to storage. claim_problems is derived too — it walks the same field list and applies rule_for, so a rule is registered in one place instead of two._
_Two mechanisms were considered and rejected on evidence. A build-time generator over contrib/parser foundered on the parser itself: the self-hosted parser has no expression-level record literal, field access or record update (ERecordLit and friends are declared in ast.mere and never constructed), so it cannot parse a schema file that also contains rules — the prerequisite is larger than the generator. Compiler-synthesised derive was not needed once the JSON view turned out to be enough._
_Two things this does not fix, both named in the source. of_json is not polymorphic — to_json is, but of_json is directed by an annotation at the call site, so claim_with has to name claim and there is one copy per record type. And rule_for still ties a field name to its rule by hand; nothing in a declaration can say which rule judges which field, but nothing checks the names either._
_examples/claims/browser_check.mjs covers what a dump cannot: that a generated <input> behaves like one, that leaving it runs the rule the schema associates with its name, and that the server's own refusals land in the generated error slots. 14/14 in Chrome. parity 79/79, ctest 13/13, dune runtest 2308/0._
v0.1.175 — 2026-08-11
_Stage 1 of the shared-schema dogfood: a type reachable only as another type's field is still a type the C backend has to declare._
_The C backend decides which structs to emit by walking what the program mentions — the expression tree and the function signatures. Neither walk descended into a record declaration's field types, so:_
type item = { name: str };
type f = { title: str, items: item list, note: str option };
_left item, list_item and option_str undeclared while the emitted code referred to all three. unknown type name 'list_item'._
_Two walks were short, and both are fixed the same way. collect_record_names now follows a record's field types when it registers it, so a record reachable only as another record's field element gets declared; and collect_mono_variant_instances walks monomorphic record declarations the way it already walked monomorphic variant declarations, so container specializations reachable only as field types get registered. seen is set before recursing, so a self-referential record terminates._
_This had been latent since records and of_json first coexisted, because a program that builds one of those values anywhere registers the type on the way past — and, as the test found the hard way, so does taking one apart: mentioning Some in a pattern is enough. The failure needs container-typed fields that are only ever decoded, which is exactly what a program looks like once the schema is shared and the codec is synthesised, since of_json becomes the only producer. examples/claims was written that way and found it; test/parity/of_json_field_only.mere holds it, and reads only the scalar fields on purpose — an earlier draft called list_len and matched Some, and passed on the broken compiler._
_examples/claims now emits and compiles as C; what stops a native build is only the two HTTP externs the native runtime does not implement, which is the gap recorded in v0.1.174 and still its own slice. parity 79/79, ctest 13/13, dune runtest 2308/0._
v0.1.174 — 2026-08-11
_examples/claims: the shared-schema dogfood, stage 0 — build it the way you would build it today, and count what is written by hand._
_An expense claim: title, date, purpose, a list of lines each with a category that may carry its own text, an optional note and approver, a status. Enough type variety that a naive answer to "generate the boilerplate" would not survive it. schema.mere declares the record and the rules, and both replicas import it._
_Half the boundary already costs nothing, which is worth stating before the complaint. to_json / of_json are synthesised per type, so neither side writes down a field name for the network — examples/profile has a hand-written serialize and parse_pairs, and this one has neither. The validation functions are the same functions on both sides, so a rule and its message exist once. problem is a declared type, so the server's refusals arrive as records rather than as text to parse, and a server-only rule (over ¥100,000 needs an approver, and the approver must exist) lands in the same error slot a local complaint would._
_The measured half: adding a cost_center: str field to claim takes four edits in Mere and one in HTML, and the compiler catches exactly one of them — the record initializer. The rule, the field table and the markup are all silent. Stopping after the record and its initializer leaves a program where both replicas compile, the server stores the field and the browser round-trips it, and it never appears on screen; the rendered page mentions it zero times. That is the number stage 2 exists to move._
_Two compiler findings fell out on the way. An extern whose result is unit could not be used as a value at all: unit is Mere's int 0 and C's void, and the closure adapter returned the call directly, so the emitted C returned void from a function declared to return int. That is the shape of middleware — contrib/http's with_arena wraps http_arena_mark — so no native build of a server using it could compile. Fixed, with test/ctests/extern_unit_as_value.mere holding it. And of_json on the C backend does not register container instances that are reachable only through a record field type: a record with both a nested record and a list/option field emits unknown type name 'list_item'. It is latent because a program that constructs one of those values by hand anywhere registers it — which the claims app happens to do, and a generated codec would not. That one is stage 1._
_The native HTTP runtime implements six externs, and with_arena / with_access_log need two it does not have (http_arena_mark, http_current_status). So a server that releases its per-request allocations exists on the Wasm host and not the native one. Recorded, not fixed here. parity 78/78, ctest 13/13, dune runtest 2308/0._
v0.1.173 — 2026-08-10
_Two bugs behind one build failure: a regression from v0.1.172, and the latent one it exposed._
_v0.1.172 left mere-ruby unable to compile — 19 use of undeclared identifier 'mu_..._as_value' — while parity, ctest and the OCaml suite all stayed green. Both harnesses that compare stdout are blind to a program that never links, and scripts/ctest.sh, which does invoke the C compiler, had not been run._
_The regression: in codegen_c the two arms that made a direct call to a top-level or inner-lifted fn were folded into the new shadowing guard. That guard is position-aware — a builtin used above a later same-named binding is still the builtin — so any call it declined fell through to the closure path instead of the direct one. Which of the three call shapes to use is a question about what the name is, not about where in the file the caller sits, so the fallthrough arm is back. (LLVM and Wasm were never affected: their equivalent arms stayed in the match as a catch-all after the builtin arms.)_
_The latent bug that made it fatal: a polymorphic fn used at more than one type is emitted once per instantiation under a mangled name, and _as_value closure wrappers are only defined for those. The unmangled name stays registered in toplevel_fn_names so source-level call sites still dispatch — so value position asked for <base>_as_value, which is never defined. let ident = fn (x) -> x used at two types and then passed as a value has never compiled on the C backend; it now picks the instance from the reference's own type, the same way the direct-call path does. Wasm refuses this case, which is honest — it has no single function to hand out either — but said "unbound variable" for a variable that is plainly bound, and now names the actual limitation._
_test/parity/multi_inst_as_value.mere covers it. A parity case rather than a ctest because Wasm's refusal is a documented UNSUP there, and because parity reports a backend whose output will not compile as MISCOMPILE — which is exactly the signal that was missing. mere-ruby builds again and its corpus is 51/51 against ruby. parity 78/78, ctest 12/12, dune runtest 2308/0._
v0.1.172 — 2026-08-10
_Shadowing a builtin, in general — the family the join fix in v0.1.169 was one member of._
_Each backend dispatches on builtin names in about a hundred match arms, and each arm decided on its own whether to ask if the program had bound that name itself. Those questions had been added one incident at a time, after someone was bitten: C had guarded 39 of 95, Wasm 14 of 139, LLVM 6 of 70. The rest were silent. let str_len = fn (s: str) -> 999 returned 999 on the interpreter and 5 on all three compiled backends — no error, no warning, a different answer. Both a top-level and a local binding were affected._
_Each backend now asks once, before any builtin arm can match: if the head of the application spine is a name the program bound, the call goes to the ordinary call paths — inner-lifted fn, top-level fn, or closure value — and never meets the builtin arms. Those three paths were already contiguous at the end of each backend's App group, so they lifted out into one function (emit_user_app, and a trio of local helpers in codegen_c) with nothing duplicated. Safe by construction: the guard is false unless the program actually bound the name, so a program that shadows nothing reaches exactly the arms it reached before._
_The first version of the guard was wrong in a way worth recording. It asked whether a name was bound anywhere in the program, but top-level bindings are sequential — the typer rejects a forward reference — so a builtin used above a later same-named binding is still the builtin. Under the first version a helper written before let show = ... silently started calling the user's show. Each backend now records the declaration position of every top-level fn and the guard compares against the body being emitted, with <= so a recursive fn still counts as binding its own name. Wasm needed the position question asked ahead of its name-only tables (fn_closure_table_idx, top_globals_wasm also hold top-level fn names), and LLVM needed the cursor reset for main's body, which it emits without a host-scope switch._
_test/parity/shadow_builtin.mere covers the shapes the guard has to recognise separately: bound at top level, locally, or as a capturing inner fn that gets lifted; called with one argument or three (the name three levels down a spine); used as a value rather than called; and the ordering rule. parity 77/77, dune runtest 2308/0._
_Swept up while here: file_size / file_pread / file_pwrite left the LLVM "no lowering yet" list they had been on since v0.1.163, the stale claim in test/parity/file_pio.mere that LLVM refuses positioned I/O, and the example count in mere --help (118 → 282)._
v0.1.171 — 2026-08-10
_contrib/state: the one-slot-vec trick, named — and the thing naming it does not fix._
_Mere has no mutable cell, and every browser client built one the same way: a vec allocated once, pushed once, then read and written at index 0. It works — the bump arena is page-lifetime, so the slot survives across event firings — and by v0.1.170 there were ten of them across four apps, each with a comment explaining the trick. contrib/state/cell names it: cell_new / cell_get / cell_set, three lines over the same vec._
_That is the smaller half, because the noise was never the real problem. examples/profile shipped with three of those cells and seven hand-written calls to a recompute function, one after each mutation site; the seventh was added during debugging, and the eighth would have been a screen quietly disagreeing with the state behind it. contrib/state/store holds one value and a list of watchers, and store_update writes and then tells them. The seven calls became zero, and the app's three cells became one store._
_Making the screen derived forced the app to be honest about something the first draft had fudged. Focusing a field takes its complaint off the screen but does not make the value right — Save has to stay withheld until the field is left and re-judged. The ad-hoc version got that by writing to the DOM behind the state's back, which worked only because nothing else ever repainted. The model now carries "what is wrong" and "what we are currently saying out loud" as two separate facts, which is what they always were._
_examples/tasks keeps cells and no store, deliberately: its screen is not derived from its state but reconciled against it — rows that stopped matching are removed one at a time so the row being typed into is never rebuilt — and a watcher that redrew on every write would destroy the one thing that app exists to protect. Two of its four cells are a timer handle and a retry delay that nothing renders at all._
_chat, tally and tasks moved to cell; profile moved to store. All four still pass their harnesses (chat and tally and tasks headless, profile 15/15 in Chrome). test/parity/store_watch.mere locks watcher order, the immediate first run, read-modify-write, and two stores at different types across all four backends. parity 77/77, dune runtest 2308/0._
v0.1.170 — 2026-08-10
_A form, as opposed to a list — and the six bindings that separate the two._
_The three browser clients before this one were lists. A list is edited one item at a time, every keystroke is worth acting on, and the only question the UI ever asks is "what is on screen". examples/profile is a settings form, which is a different animal: it is a set of fields with a shape. It is valid or it is not. It differs from what was loaded or it does not. Answering either question means holding two versions of the same record at once — what the server confirmed and what the user has done since — and comparing them. The change count, whether Save is offered, what Revert restores: all three are derived from that pair, none of them is a flag anyone sets._
_Six bindings fall out, and each is here because the form cannot be written without it. dom_on_blur and dom_on_focus are what make validation feel like a form rather than a nag — judge the field when the user leaves it, take the complaint back down when they return to fix it; checking on every keystroke tells someone their email is invalid while they are still typing the part before the @. dom_on_change is the only event a <select> produces, so without it the theme picker is inert. dom_checked / dom_set_checked exist because a checkbox has no useful value — its meaning is el.checked, a property that is not the attribute of the same name. And dom_remove_attr closes a hole dom_set_attr left in v0.1.152: a form could disable its save button while the input was invalid and then never enable it again._
_The fields are described once — id, label, a getter, a setter, a check — and everything else loops over that list: painting, comparing, validating, reverting. Record fields cannot be named at runtime, so without the get/set pair each of those loops would have been written out once per field._
_The server rejects things on purpose. The client checks what it can see; the server also enforces a rule the client cannot know (a reserved display name) and answers 422 with per-field messages, which land in the same error state a local complaint does. run_dom_headless.mjs gained --pick and --check for the two controls that are not text, removeAttribute and checked in its DOM stub, and value / checked in its dump — a filled-in form used to render identically to an empty one. examples/profile/browser_check.mjs covers what the stub cannot: it can fire a blur, but only a browser can cause one. 15/15 in Chrome, parity 76/76, dune runtest 2308/0._
_Correction: the version bump in v0.1.169 left test/test_basic.ml's version assertion pinned at 0.1.151, so that slice's suite was red as committed. Fixed here._
v0.1.169 — 2026-08-10
_args () on LLVM — the last reason a Mere CLI ran on three backends out of four — and the shadowing hole finding it exposed._
_The B+-tree in mbtree has run on all four backends since v0.1.163, when positioned file I/O landed on LLVM. Its command line had not: args was still interp + C, so mere -ll refused the program that wrapped the store. LLVM now stores main's argc/argv into globals and folds them into a str list, dropping the program name, which is what interp and C hand back. main keeps its no-argument signature unless the program actually asks for argv. The strings are argv's own rather than copies into the current region: a str is a plain NUL-terminated pointer on this backend and argv outlives the program, so there is nothing to allocate and nothing that can outlive what it points at._
_Writing the parity case for it turned up something worse. The test defined a join helper — the obvious name for joining strings — and LLVM lowered the call to pthread_join(i64) and handed it a str list. join is also the thread builtin, and LLVM guarded str_eq / is_digit / is_alpha / is_space against a user's same-named binding but not join, which C has guarded since Phase 30.0 with a comment predicting exactly this. The guards had been added one incident at a time; they now go through one user_shadows_llvm that asks about locals, lifted inner fns and top-level fns alike, the way C's user_shadows does. test/parity/shadow_builtin.mere locks it down, and docs/reserved-names.md gained the section distinguishing this axis — a name Mere owns — from the C-symbol collisions the rest of that document is about._
_mbtree data.db set 42 100 / get / selftest now produce identical output on interp, C, LLVM and Wasm. parity 76/76, dune runtest 2308/0. mere -v also reports the truth again: lib/version.ml had been left at 0.1.151 while the changelog ran to 0.1.168._
v0.1.168 — 2026-08-10
_A list you can filter and edit, and the three bindings it forced._
_Every browser client so far only appended, and a list that only grows can be redrawn wholesale: dom_set_text el "" drops every child and the rest is rebuilt from state. examples/tasks cannot. Each row holds an <input> you type into, so rebuilding the list while you are editing takes the field out from under the caret — which makes three things necessary rather than convenient. dom_remove detaches one node and leaves its siblings, and their focus and half-typed text, alone. dom_on_input fires on every keystroke, so the filter can narrow as you type instead of on submit. dom_set_timeout / dom_clear_timeout do two jobs: a keystroke cancels the pending filter and queues a new one, so a burst of typing costs one re-render rather than one per character; and a save that fails comes back on a doubling delay instead of being dropped._
_Filtering reconciles rather than redraws: rows that stopped matching are removed one at a time, rows that started matching are built and appended, and a row that stays is never touched. The browser check asserts exactly that — after typing in the search box, the row being edited is the same DOM node with its caret still at offset 3 — along with delete removing exactly one row, and a save failing and landing on the retry. All nine pass in Chrome._
_examples/tasks/server.mere is deliberately thin: contrib/store/kvlog behind one tab-separated mutation endpoint, plus POST /api/flaky to fail the next N saves on purpose, because a retry path that is never taken is a path that is never tested. scripts/run_dom_headless.mjs gained --type <id>=<text>, which sets a field and fires input the way a keystroke does, and its DOM stub now implements remove(). parity 74/74, dune runtest 2308/0._
v0.1.167 — 2026-08-10
_An open region is closed to the heap before codegen, which unblocks the last LLVM gap the dogfoods hit._
_vec_new and file_pread hand back a Vec[R, T] whose region no region block ever constrained, and the typer is content to leave R open. That is harmless on the C backend, which erases types, and fatal on LLVM, which does not — and the failure arrived three removes from its cause. An open R makes the Vec non-concrete; a helper taking one is therefore not resolvable and gets dropped from the backend's function list; the inner fn that uses it treats it as a capture instead of a known top-level name; and the emitted call site refers to an SSA value that does not exist. What clang reported was a type mismatch on a register several functions away._
_A pass before lifting closes any open region variable to the default heap region. That is completing the inference rather than working around it, which is what keeps every downstream name stable — an earlier attempt tagged open regions leniently instead, and made a struct's name depend on how far unification had progressed, so the same type was registered under one name and referred to under another._
_The B+-tree from the mbtree dogfood now compiles and runs on LLVM, and a tree it writes there reads back correctly under the C backend. test/parity/lifted_vec_capture.mere locks the shape. mbtree's CLI still needs args(), which has no LLVM lowering. parity 74/74, dune runtest 2308/0._
v0.1.166 — 2026-08-10
_Two corrections to the capture typing added in v0.1.165, and a precise account of what still blocks mbtree on LLVM._
_ty_is_concrete is the wrong question to ask about a capture. A Vec[R, int] whose region is still a variable is not concrete, yet llvm_ty_of lowers every Vec to ptr and never looks at the region — so rejecting it left the capture untyped. The lookup now asks whether the backend can represent the type at all._
_v0.1.165 also fell back to searching every top-level fn body for a concrete type of the captured NAME. Names are not unique across functions, so a capture of b : Vec[R, int] picked up an unrelated b : int from elsewhere in the program and was declared i64. Removed — a wrong type is worse than no type, and the loud error is what should happen._
_What still blocks mbtree, stated exactly: get_i8, a top-level curried helper, is not among the backend's fn_decls, so the free-variable scan treats it as a capture of the inner fn that uses it. Its type carries an unresolved region, and the emitted call site refers to %get_i8 — an SSA value that does not exist, because the value is a top-level function. Two things would have to change: the name would have to be known as a top-level binding so it is never captured, or a captured top-level function would have to be materialised at the call site._
_A fix was attempted and backed out: tagging the region position of Vec / Map / StrBuf leniently makes a struct's name depend on how far unification has progressed, so the same type got registered under one name and referred to under another, and test/parity/bytes_typed_fn_unused.mere regressed to llvm:MISCOMPILE. Any real fix has to keep the tag stable. parity 73/73, dune runtest 2308/0._
v0.1.165 — 2026-08-10
_A capture the LLVM backend cannot type is refused instead of guessed._
_When an inner fn is lifted, each captured variable becomes a parameter, and the capture's type came from the first Var occurrence inside the fn body whose recorded type was fully concrete. When none was, the search fell through to its initial value — TyUnit, which lowers to i64. So a capture the backend could not type was silently declared i64 while the call site passed a pointer, and the failure surfaced as a clang type error about an SSA register, several steps from the cause. That is what stopped the mbtree dogfood from building here._
_The lookup now prefers a concrete type, falls back to any recorded one, and takes the binding site's type when a later use has been generalized; failing all of that it raises a Mere codegen error naming the variable. mbtree's real blocker is now stated plainly: unsupported LLVM codegen type element: 'a — capturing a polymorphic value, which this backend cannot represent and the lifting code already says so in a comment. The monomorphisation there covers a local let rec applied at one type; a polymorphic top-level helper captured by an inner fn is still open._
_parity 73/73, dune runtest 2308/0._
v0.1.164 — 2026-08-10
_A partially applied extern is a value, not a call._
_The Wasm extern path collapses a curried App chain into one call $name, taking the argument count from the call site rather than the declaration. That is right when the application is saturated and emits a call with the wrong arity when it is not, so worker_call req on a two-argument extern produced WAT that wat2wasm rejected outright. Found while writing contrib/async, where the whole premise is that an extern with its request applied is already a Task — the tally client had to wrap every one in fn (cb) -> … cb._
_Under-application now eta-wraps: the missing arguments become a lambda and the expression routes through the ordinary anonymous-closure path, so the partial application is a closure like any other value. The wrapper is gone from examples/tally/app.mere. Two regression tests pin both halves — a partial application reaches its callee through call_indirect, a saturated one still emits a direct call. parity 73/73, dune runtest 2308/0._
v0.1.163 — 2026-08-10
_Positioned file I/O on LLVM, closing the last backend gap in that group._
_file_openrw / file_size / file_pread / file_pwrite / file_fsync / file_close were interp + C + Wasm; parity reported llvm:UNSUP for test/parity/file_pio.mere. The LLVM runtime now implements them over libc with the same contract as everywhere else — a handle, bytes crossing as a Vec[int], a short read past EOF rather than a padded one — and that case is llvm:MATCH. All four backends agree on writing past the end, overwriting a window in the middle, and reading a window back after a reopen._
_A File travels as an i64 here rather than a raw ptr: lifted inner functions type every parameter uniformly, so a pointer could not be passed through one. The runtime converts at its own boundary._
_Not fixed, and worth naming because it is what stops the mbtree dogfood from building on LLVM: a lifted inner function declares all its parameters i64, so passing a closure to one fails to typecheck in the emitted IR. args() also has no LLVM lowering. Neither is about files — mbtree hits both — but they are the two things between this backend and running that dogfood. parity 73/73, dune runtest 2306/0._
v0.1.162 — 2026-08-10
_mere install fetches before checking out._
_A dependency's repo is cloned once and cached, keyed by repo-and-rev, but an existing cache was never refreshed. So the first install after a new commit was pushed failed with fatal: reference is not a tree: <sha>, and the only remedy was deleting ~/.mere/cache by hand — a git error with no mention of a cache, a long way from the cause. It now fetches before checkout._
_Found while repointing the mere-blog dogfood at a current revision, which also repaired its packaged path: its host was pinned to a release whose JS glue predates the current value representation, so a build made with a current compiler linked and then failed on its first request. With everything pinned to one commit, mere serve runs that app on the vendored host end to end._
v0.1.161 — 2026-08-10
_Sequencing for callback-shaped work, and what it does not fix._
_Everything that crosses a thread or a network hands its result to a closure. One such call costs an indentation level, which is nothing. The cost shows up when steps depend on each other: an ordinary fold over a list where each element is a round trip cannot be written as a fold, and becomes a recursion carrying its own continuation. The tally client had exactly that recursion written out by hand, and so would every app that sequenced two requests._
_contrib/async/async.mere carries it once. A Task is (str -> unit) -> unit; async_each is the fold, async_map collects results in list order rather than completion order, async_then chains, async_all fans out and joins. It is ordinary Mere — no language support, just the callback shape given names — and test/parity/async_combinators.mere pins the ordering across all four backends. examples/tally/app.mere now uses it and the hand-rolled recursion is gone._
_What it does not fix is the nesting: async_then still puts each step inside the previous one's closure, so the four-step sync in tally is four levels deep. Removing that needs syntax, not a library — some form of do-notation or await — and shipping the combinators first is how to find out how much of the pain was the fold and how much is the nesting. On the evidence of one app: the fold was the part worth removing, and the nesting is legible enough to live with for now._
_Found while writing it: an extern cannot be partially applied on the Wasm backend. worker_call req emits a direct one-argument call rather than a closure, so a Task built from an extern still needs an explicit fn (cb) -> extern req cb. Worth eta-wrapping the way nullary builtins already are._
v0.1.160 — 2026-08-10
_A request's allocations can be released when the request ends._
_The default region is a bump arena with no free, so a long-lived server keeps every request's working memory forever — the parsed body, the strings a handler concatenated on the way to a response, and every string the host wrote in. A handler that builds a 200-piece string cost 162,840 bytes per request, none of it reclaimed. The dogfoods noted this as a constraint twice without measuring it; the number above is a 300-request run against a Node host, reading the module's own __lang_bump._
_contrib/http/arena.mere marks the arena on the way into a request, and the glue rewinds it once the response has been copied out of linear memory. Same handler, same run: 24 bytes per request, and 200 consecutive responses are byte-identical, so nothing is reclaimed early. A plain region block around the handler gets most of the way there on its own (162,840 → 722, the copied-out response), which is worth knowing when no host cooperation is available._
_It is opt-in, and the reason is the rule it depends on: nothing a handler allocates may outlive its response. That holds for ordinary request/response work and fails for a handler that stashes a value in something longer-lived — an in-memory session store, a cache, a subscriber list. mere-blog is exactly that case and deliberately does not use it; examples/tally/server.mere keeps all its state in a log file and does._
_parity 72/72, dune runtest 2306/0; the tally client still syncs against the wrapped server._
v0.1.159 — 2026-08-10
_args() returns the arguments on a plain Wasm host._
_The runners have supplied arg_count / arg_get for a long time; the builtin was never wired to them and returned Nil unconditionally, so a CLI compiled to Wasm silently saw no arguments while the same source on C saw them all — a disagreement with no error anywhere. It now builds the list from the host, and a host that reports 0 yields Nil, which is what the hardcoded empty list got right for a browser. contrib/dom answers 0 so a page keeps the behaviour it had._
_Verified by running the same program on both backends with arguments: foo bar gives n=2 [foo] [bar] on C and on Wasm. parity 72/72, dune runtest 2306/0._
v0.1.158 — 2026-08-10
_One definition of the host boundary, instead of five._
_v0.1.157 added a number so a host and a module could refuse each other. This removes the reason they drifted in the first place: readCStr, writeStr, bumpAlloc and the closure caller existed in five separate copies, and a fix to one never reached the others. run_wasm.js grew the string length header and open-coded it at four call sites; contrib/dom learned the i64 closure convention and contrib/http did not; a fixed 4KB scratch window outlived the move to the shared heap in exactly one glue. Each was a wrong answer at runtime, never a link error._
_scripts/mere_host.js now holds all of it — makeMarshal for the four value operations plus writeBytes / readBytes / copyToStr for the byte buffers a driver moves, and makeClosureCaller for dispatch — with the layout written down once at the top. run_wasm.js, run_http_server.js, pg_env.js and contrib/http/http.glue.js all use it: 251 lines deleted, 48 added._
_contrib/dom/dom.glue.js keeps its copy, because it is an ES module a browser fetches and cannot require the shared one. That is now the only place the layout appears twice, so scripts/check_host_abi.js checks it: the ABI constants must agree, and the copy must still write the length header, return a pointer past it, and pass closure arguments as BigInt. Each of those three was wrong in some host at some point._
_parity 72/72, dune runtest 2306/0; the chat, tally and mere-blog apps all rebuild and pass their checks. The ABI guard earned itself immediately — a stale mbtree build from before v0.1.157 was refused by name instead of quietly returning empty strings._
v0.1.157 — 2026-08-10
_An ABI number, so a host and a module can refuse each other._
_Nearly every bug the three dogfoods turned up was one shape: a boundary written when a Mere value was 4 bytes and a str was a plain C string, left behind when both changed. A host and a compiled module agree on far more than the import list — the value representation, the string layout, how a closure is called — and none of it is checked, so an older host links cleanly and then returns empty strings or crashes on the first request. That is how a request line arrived as "", how a password hash hashed to nothing, and how a result set came back with every column empty._
_Every module now exports __mere_abi, and every host checks it before touching the instance: scripts/mere_abi.js for the Node runners and contrib/http, inlined in contrib/dom since it loads in a browser. A module without the global is refused as pre-ABI-1 with a note to recompile; a newer one tells the host to update. The self-host emitter in contrib/codegen/codegen_wasm.mere emits the same global, so both codegens stay in step and the bootstrap fixpoint still holds._
_ABI 1 names what was implicit: i64 values with addresses in the low word, [i32 len][bytes][NUL] strings whose value points at byte0, closure records of { i32 env, i32 fn_idx } called as (i64, i64) -> i64, 8-byte compound fields and 16-byte { tag, payload } variant cells._
_An audit of the remaining host boundaries found three more of the same shape, all in run_http_server.js and all fixed by routing through the shared writeStr rather than open-coding the layout a fourth time: getenv, sha256_hex and __lang_str_of_float. Open-coding is precisely how these drifted — run_wasm.js grew the length header at each of its own call sites and none of the copies elsewhere followed. parity 72/72, dune runtest 2306/0._
v0.1.156 — 2026-08-10
_of_json gets a working Wasm backend, and the last of the raw-C-string handoffs go._
_The Wasm JSON decoder was written for the 4-byte value model. The parser runtime carried i64 signatures over i32 bodies, and the generated __ojnode_* decoders built records with 4-byte fields and variant cells of 8 bytes, so typed request decoding had no Wasm backend at all. The runtime's JSON tree is private to the parser and now says so — i32 throughout — while every decoder produces a real Mere value: 8-byte record and tuple slots, 16-byte { tag, payload } cells, and __mj_atoi accumulating in i64 so a number past 2^53 survives. test/parity/of_json_composite.mere covers records, lists, options (including None-on-error), nested lists of records, variants both nullary and payload-carrying, and the wide integer._
_`show` and `to_json` on C returned bare string literals for bool, unit, closures, the empty list, nullary constructors and every null — and a bare literal is not a Mere str, which carries its length in the word before byte0. print (show true) looked fine because print formats with %s and stops at the NUL; print ("x" ++ show true) segfaulted, and str_len (show true) returned whatever preceded the literal in rodata. Same for str_repeat at n≤0, str_join of an empty list, and __lang_fail_str._
_Two Wasm hosts still handed over raw bytes: mem_to_str in scripts/pg_env.js, which is where every column value a database driver reads off the wire becomes a str — an entire result set came back empty — and read_file in scripts/run_http_server.js. And contrib/http/http.glue.js read the response body by scanning for a NUL, so a .wasm asset served out of read_file truncated to nothing at its first byte._
_Together these close the gap the previous entry left open: mere-blog now builds and runs on both backends against Postgres — signup, cookie sessions, typed request decoding, authenticated post creation — and serves the same Wasm admin client byte-identically from either. The admin UI's browser checks pass against the native binary and the Wasm host alike. parity 72/72, dune runtest 2306/0._
v0.1.155 — 2026-08-10
_Four boundaries that a database-backed web app walks straight into, found by building an admin UI for the mere-blog dogfood._
_`to_json` on Wasm had the same rot show did at v0.1.153: every case of the emitter built 4-byte cells with i32 fields and kept the accumulator str in an i32 local, so a module that serialized anything was rejected by wat2wasm. to_json_int also read its i64 argument through i32 comparisons, so sign and magnitude were wrong past the low word. mere-blog serializes a typed record on every response, so the whole Wasm backend was unavailable to it. test/parity/to_json_composite.mere covers it._
_The native runtime handed Mere raw C strings. A Mere str is [size_t len][bytes][NUL] with the value at byte0, so __lang_str_size reads the length from s[-1] — and the native HTTP server passed a static char[] request line straight to the handler, which read back as "". Every route on a native build 404'd. Same for http_current_body and http_get_header, and for the crypto/encoding externs: __to_hex / __to_b64 / gen_request_id malloc'd their results, so sha256_hex returned "" and every password hash with it. They now allocate through __lang_str_alloc._
_A `unit` parameter broke extern closure adapters on C. extern fn f: unit -> str lowers to str f(void), but the adapter that lets it be used as a value passed its argument through, calling a 0-arity function with one._
_The native HTTP response measured its body with `strlen`. A Mere str carries its length and may contain NULs, so a .wasm read through read_file was truncated at its first zero byte — a native Mere server could not serve its own compiled client. It now uses __lang_str_size._
_Together those let mere-blog build and run natively against Postgres with a current compiler — signup, cookie sessions, authenticated post creation — and serve a Wasm admin client whose form validates drafts with the same validate.mere the server enforces, compiled to Wasm on one side and C on the other. parity 71/71, dune runtest 2306/0._
_Known and not fixed: of_json on Wasm is still entirely 4-byte-model, runtime and generated decoders both, so typed request decoding has no Wasm backend until that is rebuilt. Native and interpreter are unaffected._
v0.1.154 — 2026-08-10
_A local-first app in Mere, split across three replicas of one store. contrib/store/kvlog.mere is an append-only key/value log over the positioned file I/O from v0.1.153: a write appends a record and fsyncs, a read replays. examples/tally runs it three ways — store.mere owns the browser's copy off the UI thread, server.mere owns the authoritative copy, and both import the same kvlog source, so the two replicas are one store compiled twice rather than two stores that agree on a format. A log written by the C backend reads back under Wasm and interp, and the reverse. file_size joins the positioned group on Wasm, since an append-only store needs the end of the file without reading it._
_The UI half (app.mere) reaches storage through a new worker_call binding in contrib/dom: a request string in, a reply handed to a closure later. In a browser that is postMessage to a Worker, which is where the store has to live because an OPFS access handle — the only synchronous positioned file I/O a browser offers — exists only off the main thread. scripts/run_dom_headless.mjs --worker <store.wasm> runs the same split under Node with replies deferred to a later turn, so the asynchrony is faithful and the whole app is testable without a browser: add a counter, reopen in a fresh process and see it persisted, press +1, sync against a running server, and lose the server and watch it fall back to local. examples/tally/store.worker.js carries the OPFS binding, and scripts/check_browser.mjs drives it in Chrome: write through the store, reload, restart the browser process, and confirm the counters come back off disk. Playwright is not a dependency, so a missing install is a SKIP. The check also confirms createSyncAccessHandle is absent on the main thread, which is the constraint the whole split exists for. A log written by the browser through OPFS parses identically under kvlog.mere compiled to C and run natively, and compiled to Wasm and run under Node — one source, three hosts, one byte format._
_What the split says about the language. Every endpoint is straight-line code — handle in store.mere returns its reply as a value and never mentions asynchrony — and every crossing is a callback. With one request in flight that costs an indentation level, which is what the chat client measured in v0.1.152. The cost shows up when steps depend on each other: sync reads local state, fetches remote, writes the merge back key by key, then publishes it, and the middle step is an ordinary fold over a list where each element is a round trip. Written against callbacks it becomes a recursion that carries its own continuation and calls it when the list runs out. That rewrite — a fold that cannot be a fold — is the clearest argument so far for giving Mere a way to name the result of a call._
_Also found: a let rec that closes over a Vec is rejected by both compiled backends ("captured variable has no recorded type" on C, "inner-lifted capture not in scope" on Wasm), so kvlog threads its byte buffer through as an explicit parameter. And dom_set_text el "" turns out to be the way to clear a container, since setting textContent drops every child — no dom_remove binding needed yet._
v0.1.153 — 2026-08-09
_Positioned file I/O on the Wasm backend, and the three broken builtins that finding it uncovered. file_openrw / file_pread / file_pwrite / file_fsync / file_close were interp + C only, on the reasoning that Wasm has no filesystem. That is true of the browser main thread and false of a Worker, which gets synchronous positioned read / write / flush from an OPFS access handle — the same contract these builtins already describe. They now lower to host imports, with bytes crossing in the mere_bytes layout the read_file_bytes path already uses, so there is no per-byte host crossing and the host never needs to know the Vec layout. scripts/run_wasm.js backs them with positioned fs calls against a handle table._
_The result: mbtree — a persistent B+-tree written against those builtins — compiles to Wasm unchanged and passes its durability selftest (20 keys, node splits, a root split, fsync, close, reopen) on interp, C and Wasm alike. A tree file written by the C backend reads back correctly under Wasm and interp, and the reverse, so the on-disk format is one format across backends rather than three that happen to agree._
_Getting there needed three fixes to builtins that were stale for the i64 value model and that nothing exercised. show over any composite type emitted WAT that wat2wasm rejected — the tuple, record, variant and list emitters all still built 4-byte cells with i32 fields and kept the str accumulator in an i32 local — so print (show [1, 2, 3]) could not be assembled at all. vec_to_list on Wasm had the same rot. On LLVM, vec_to_list stored the payload tuple into the node by value, but the Phase 24 variant layout makes that field a pointer, so 16 bytes went into an 8-byte slot and the helper segfaulted for any input. test/parity/show_composite.mere and test/parity/file_pio.mere cover both gaps; the existing 68 parity inputs only ever showed scalars and never opened a file handle. Full suite: parity 70/70, dune runtest 2306/0._
_Known, not fixed: args() on the plain Wasm backend is hardcoded to the empty list even though run_wasm.js supplies arg_count / arg_get, so a CLI program silently sees no arguments there while the C backend sees them. Correct for a browser host, wrong under Node._
v0.1.152 — 2026-08-09
_The chat client rewritten in Mere, and the four stale host boundaries it exposed. examples/chat/app.mere replaces the 74 lines of hand-written JavaScript that examples/http_chat.mere used to serve, so both halves of the demo are now Mere and both share contrib/http/escape.mere for JSON escaping — one implementation compiled to C on the server and Wasm in the browser. contrib/dom gains 12 externs in the three groups a document-shaped app needs and a game does not: element construction (dom_create / dom_append / dom_set_attr / dom_set_value / dom_scroll_to_end / dom_on_submit), request/response (dom_fetch blocking + dom_fetch_async callback, sharing dom_fetch_status / dom_fetch_header), and server push (dom_sse). Both request shapes ship deliberately: the app performs its bootstrap through each so the cost of a callback continuation is visible in Mere source rather than argued about._
_Four host-side boundaries had drifted from the compiler and only a real app touched them. `str` layout: mere_strbuf_to_str in lib/codegen_wasm.ml hand-rolled its bump allocation and skipped the [i32 len] header that __lang_strlen reads from ptr-4, so every strbuf_to_str result read back as "" on Wasm alone — invisible under print, which exits through the host and scans to NUL. The same header was missing from the writeStr helpers in contrib/http/http.glue.js, contrib/dom/dom.glue.js, scripts/run_wasm.js and scripts/pg_env.js; run_wasm.js had the correct sequence open-coded at four call sites, which is why the fix never reached the shared helpers. Closure ABI: http.glue.js still passed plain numbers to the (param i64 i64) closure type v0.1.127 introduced, so every contrib/http demo threw on its first request when rebuilt. Scratch memory: dom.glue.js still wrote host strings into a fixed 4KB window at 56K that wrapped around, which a multi-KB bootstrap response overruns. Missing import: the runners never supplied time, which the prelude imports unconditionally. test/parity/strbuf.mere closes the coverage gap that let the first of these live — StrBuf had no parity case among the other 68._
_contrib/http/static.mere now serves assets through a new http_send_file host extern instead of read_file. Mere strings are NUL-terminated, so the old path truncated any binary file at its first zero byte, and a .wasm module begins with one: the server could not serve its own compiled client. Bytes now go from disk to socket without entering Mere, which also distinguishes an empty file from ENOENT. scripts/run_dom_headless.mjs runs a mere -w module against contrib/dom under Node with a small DOM, fetch, synchronous XHR and EventSource, so browser-targeted Mere is testable without a browser._
v0.1.151 — 2026-08-08
_RV32I fantasy-console I/O + a browser build. Two new externs the mere -rv backend lowers to memory-mapped I/O and a syscall, turning a mere -rv program into a playable cartridge: key n reads the held state of button n from an input register at 0x7F9000 + n (a lbu), and present () ends a frame and yields to the host via ecall a7=100, resuming on the next instruction next frame — so a program's main loop is a coroutine whose state lives on the RISC-V call stack. Paired with the existing fb_set (framebuffer store), these three are the whole hardware contract. contrib/site/playground/rvconsole.mere is the memu RV32IM emulator compiled to WebAssembly and wired to the DOM (ROM via dom_rom_byte, input via dom_key_held, framebuffer blitted to a <canvas>), and game.mere is an arrow-key-playable cartridge; both ship to the playground. lib/codegen_riscv.ml, contrib/site/build_full.sh, contrib/site/build.mere._
v0.1.150 — 2026-08-08
_Full structural == / != on the RV32I backend. A comparison at a compound type (tuple, record, or payload-carrying variant) now generates a per-type __eq_<tag>(a,b) helper that recurses over the structure — mirroring codegen_c's eq_<tag>. Helpers are deduped by a type tag and emitted from a worklist, so recursive types (e.g. a cons list) terminate; type parameters are substituted with the concrete arguments at the use site, so list int and list str get distinct monomorphic helpers. Verified byte-identical to the interpreter across tuples, records, single- and tuple-payload constructors, Circle 5 vs Dot, a recursive ilist, option, and a tuple with a string field. Only == on functions is rejected. lib/codegen_riscv.ml._
v0.1.149 — 2026-08-08
_A framebuffer primitive for the RV32I backend: fb_set x y v lowers to a byte store into a 64×32 framebuffer at 0x7F8000 (above the stack). Declared in a program as extern fn fb_set: int -> int -> int -> unit;, it lets a Mere program draw pixels; an emulator that renders that region turns it into a tiny "fantasy console" — a Mere program, compiled by mere -rv, drawing graphics on the Mere RISC-V CPU (see the memu project's riscv-console demo). lib/codegen_riscv.ml._
v0.1.148 — 2026-08-08
_Correct == / != on enums (RV32I). A comparison at a non-primitive type was comparing heap pointers; now an all-nullary variant type (an enum) compares its tag word, which is exact. Compound values (tuples, functions, payload-carrying constructors) would need a recursive structural equality — they now raise a clear Codegen_error pointing at pattern matching instead of silently comparing pointers. Ints/bools/type-variables are unaffected. lib/codegen_riscv.ml._
v0.1.147 — 2026-08-08 — the Mere compiler runs on the Mere CPU
_The self-hosting tower closes: the Mere-written compiler (lexer + parser + typer + Wasm codegen), lowered by mere -rv to a ~380KB RV32IM binary, runs on the Mere-written RV32I emulator and emits WAT byte-identical to the interpreter. Self-language → self-backend → self-CPU._
_The last bug was a memory-map overlap: the globals+heap region sat at 0x10000 (64KB), but the self-hosted compiler's code is ~88KB and extended past it, so a global write corrupted the code and the program jumped into garbage. Small programs (<64KB of code) never hit it. Fix: move globals+heap to 0x200000 (2MB), well above any code — layout is now code [0,2MB) | globals+heap ↑ | stack ↓ from 0x7E0000 | print scratch 0x7F0000 (needs an ≥8MB emulator). Global slot addressing now materialises the full address, so the global count is unbounded. Verified: the self-hosted compiler produces byte-identical WAT on RV32I for arithmetic, let/if, and a recursive factorial (133 lines of WAT); all existing tests still pass. lib/codegen_riscv.ml._
v0.1.146 — 2026-08-08
_Long-range conditional branches on the RV32I backend. A bare B-type branch reaches only ±4KB and silently truncated its offset in large functions (the self-hosted compiler has functions well past that). Every conditional branch is now emitted as an inverted branch skipping a J-type jump (±1MB reach), so branch targets are correct at any distance. All existing tests stay byte-identical to the interpreter. lib/codegen_riscv.ml._
v0.1.145 — 2026-08-08
_An injected Mere-source runtime prelude for the RV32I backend (lib/rv_prelude.ml) — the string / char-class / Map tail the self-hosted compiler needs, written on top of the primitives codegen_riscv emits instead of hand-assembled. mere -rv prepends it to the user source, so it goes through the normal typer + desugar; the definitions shadow the builtins of the same name (compile_app resolves user bindings first) and reachability emits only the ones a program uses. Provides is_digit/is_alpha/is_space, not, and str_starts_with / str_ends_with / str_index_of / str_contains / str_repeat / str_rev / to_lower / to_upper / str_trim / str_join / str_split / str_replace / str_unescape / int_of_str. Map is an assoc-list in a one-cell Vec with str_eq keys (mirroring the self-hosted Wasm backend): since the typer forces the Map type on the map_new name, codegen_riscv intercepts the map_* builtins and dispatches to rvmap_* helpers. Also relocated the print scratch buffer out of the heap region (to 0x7F0000, stack to 0x7E0000) so large programs don't clobber it. Verified byte-identical to the interpreter across the whole string/Map surface; existing tests still pass. With this, the Mere-written compiler compiles to a ~94k-instruction RV32I binary and runs on the emulator (reaching its own typer) — the last correctness gaps on real input are being chased. lib/codegen_riscv.ml, lib/rv_prelude.ml._
v0.1.144 — 2026-08-08
_Vec — a mutable, growable array — on the RV32I backend (M3). A Vec is a [len][cap][dataptr] cell over a cap-word buffer; vec_push doubles the buffer when full (allocating a new one and copying, since the bump heap can't realloc). vec_new / vec_push are runtime helpers; vec_get / vec_set / vec_len are inlined. Verified byte-identical to the interpreter across push/get/set/len, growth well past the initial capacity, and an iterating sum. (ref turned out to be unused in the self-hosted compiler — Mere's mutability flows through Vec/Map, so no separate reference cell is needed.) Remaining for self-host: the Map collection and a tail of string builtins (str_replace / str_join / str_split / …). lib/codegen_riscv.ml._
v0.1.143 — 2026-08-08
_A big step toward self-hosting on RV32I — driven by feeding the Mere-written compiler (lexer + parser + typer + codegen) through mere -rv and closing each gap it hit:_
- _Top-level value bindings (globals). Non-function top-level
lets now
live in a fixed memory region (below the heap), initialised in order at the start of __main; any top-level function can read them. The peeler no longer stops at the first non-function binding, so functions defined after a value binding are still lifted._
- _Recursive local closures (
let rec f = fn ... in ...): the closure is
allocated first and f bound to it before the captures are filled, so the body's self-reference resolves._
- _Fully-recursive pattern binding (arbitrarily nested tuples / records /
constructors; as-patterns; string patterns) via a container-pointer parked on the stack. Record update { r | f = e }. region { } is a no-op (the bump heap doesn't reclaim)._
- _String / char builtins:
str_of_int,ord,chr,char_at,
substring, print_no_nl, fail, and int-only show; plus StrBuf (strbuf_new / strbuf_push / strbuf_to_str / strbuf_len)._
- _> 8-argument calls: args beyond a0–a7 are passed on the stack with
caller cleanup._
- _Fix: a user binding (local / global / top-level) now shadows a
same-named builtin, matching the interpreter._
_All existing RV32I tests remain byte-identical to the interpreter. The self-hosted compiler now gets much deeper before hitting the remaining gaps (more string builtins like str_replace, and the Vec/Map collections). lib/codegen_riscv.ml._
v0.1.142 — 2026-08-08
_Records on the RV32I backend (M3, second slice). A record is a heap block whose fields are laid out in declaration order (from the Top_record decl); a Record_lit reorders its fields to that order, evaluates, and fills the block, a p.field reads the field's slot (the field's index is resolved from p's type via the typer's .ty), and a record pattern T { f = a, .. } binds each field by its offset — in both let and match. Verified byte-identical to the interpreter across construction, field access, out-of-order literals, record pattern destructuring in let, a string field, nested records (s.a.x), and field patterns with guards in match. Next: attempt the Mere-written selfhost-compile. lib/codegen_riscv.ml._
v0.1.141 — 2026-08-08
_String builtins + content comparison on the RV32I backend (M3, first slice toward self-hosting). == / != / < / <= / > / >= on str-typed operands now compare content, not pointers (dispatched on the typer's .ty), via new __str_eq / __str_cmp runtime helpers (the latter normalised to -1/0/1, matching the interpreter). New builtins: str_of_int (itoa into a heap string), str_eq, str_compare, ord, chr, char_at, substring (end-exclusive, matching the interpreter's String.sub s start (end-start)), and print_no_nl. Verified byte-identical to the interpreter across equality/ordering, signed str_of_int, char access, and substring. Next M3 steps: records, then attempting the Mere-written selfhost-compile. lib/codegen_riscv.ml._
v0.1.140 — 2026-08-08
_Closures on the RV32I backend (M2, final slice) — the last piece before self-hosting. A closure is a heap block [code_ptr][captured...]; fn x -> body captures the locals its body uses, lifts the body to a top-level lambda (code(env in a0, arg in a1) — captures loaded from the env at entry, param from a1), and evaluates to the block. Application splits two ways: a saturated direct call to a known top-level function keeps the fast register-allocated jal path, while everything else (lambdas, higher-order params, curried/partial application through values) evaluates the head to a closure and applies arguments one at a time via an indirect jalr. That unlocks first-class and higher-order functions: verified byte-identical to the interpreter for apply/twice/compose, free-variable capture, and — the milestone — the prelude's own list_map / list_fold / list_filter / range / list_product driven by lambda arguments over a Cons/Nil list, all running on the Mere-written CPU. (Partial application of a bare top-level function still wants an explicit lambda; a follow-up.) lib/codegen_riscv.ml._
v0.1.139 — 2026-08-08
_Strings on the RV32I backend (M2, third slice). A string is a pointer to [len:4][bytes][pad to 4]. Literals become rodata blocks emitted after the code, loaded by a new LoadAddr item (lui+addi of the label's absolute address — the binary loads at 0, so absolute = offset); a new Bytes item carries the raw data. print writes the bytes then a newline (print_endline semantics, matching the interpreter), ++ calls a new __str_concat runtime helper that bump-allocates and byte-copies both operands, and str_len reads the length header. Verified byte-identical to the interpreter for literals, concat chains, str_len, and strings flowing through functions, an ADT payload, and tuple destructuring. mere -rvs now also lists the rodata. Next: closures (the last M2 piece before selfhost). lib/codegen_riscv.ml._
v0.1.138 — 2026-08-08
_ADTs and pattern matching on the RV32I backend (M2, second slice). A constructor is a heap block [tag][payload] — the tag is the variant's index within its type (from Top_type decls), the payload is one word (an int, or a pointer; a tuple pointer when the constructor has several fields). match stashes the scrutinee in a binding slot, then for each arm tests the pattern (constructor tag compare, int/bool literal, or an irrefutable tuple/var bind) — branching to the next arm on mismatch — and binds its variables before running the body; guards are supported. Covers the top level plus one level of sub-structure, enough for Option/Result and typical enums (deeper nesting raises a clear Codegen_error). Verified byte-identical to the interpreter across a nullary enum, single- and tuple-payload constructors, a recursive ilist (sum/len/max over a hand-rolled cons list), and the built-in option. Next: strings, then closures. lib/codegen_riscv.ml._
v0.1.137 — 2026-08-08
_The RV32I backend grows a heap (M2, first slice) — tuples. _start now sets gp as a bump-heap top pointer (heap at 0x10000, below the print buffer and stack); a tuple literal evaluates its elements onto the memory stack, bump-allocates an n-word block, and fills it (no call between the bump and the stores, so the block pointer stays put), leaving the pointer as its value. A tuple-pattern let (a, b, ...) = e loads each field into its binding. This is the first non-integer value representation — values are now "a word that is either an int or a heap pointer". Verified byte-identical to the interpreter across tuple construction, 3-field tuples, tuple-returning functions, elements that are themselves calls, tuple-in/tuple-out (swap), and nested-tuple dot products. Next slices: ADTs + Match, strings, closures. lib/codegen_riscv.ml._
v0.1.136 — 2026-08-08
_A disassembler for the RV32I backend — the debugging surface the direct byte-emitter skipped. lib/riscv_disasm.ml decodes one RV32IM word to a readable mnemonic (mirroring the emulator's imm_ decoders, inverse of the enc_ encoders), recognising the mv / li / ret / j / nop / beqz pseudo-ops. Two new modes use it: mere -rvs file.mere prints an assembly listing of the compiler's own output (address, hex, mnemonic, with real label names on jumps/branches), and mere -rvd file.bin disassembles a flat binary. This makes the register-allocated code inspectable — e.g. factorial shows the param pinned in s1, n <= 1 folded to slti a0, a0, 2, and n * fact(n-1) as mul a0, s1, a0 — and sets up debugging for the heap/closure work ahead._
v0.1.135 — 2026-08-08
_Register allocation for the RV32I backend (M1). The M0 stack machine kept every named binding in a memory frame slot and every intermediate on the memory stack; M1 puts a function's params and lets in the callee-saved registers s1..s11 (spilling only the 12th-plus binding to memory), folds the hot n - 1 / n < 2 of recursion into a single addi / slti, and reads binop/comparison operands straight out of their registers when possible — a value in a callee-saved register survives the other operand's evaluation, nested calls included, so no spill is needed. Static instruction count drops 16–32% (~23% average) across the sample programs; still byte-identical to the interpreter across factorial, Fibonacci(25), Ackermann, deep recursion (sumto 1000), gcd, and a 14-local function that exercises the memory-overflow path. lib/codegen_riscv.ml._
v0.1.134 — 2026-08-08
_A fifth backend — Mere lowers to native RV32IM machine code. Where -c / -ll / -w delegate to a C compiler / LLVM / a Wasm runtime, mere -rv file.mere emits a flat little-endian binary directly (no external assembler or linker) that runs on the Mere-written RV32I emulator: the self-made language now runs on the self-made CPU. This is the M0 vertical slice — 32-bit integers, arithmetic, comparisons, short-circuit &&/||, if, let, top-level (mutually) recursive functions, saturated calls, and print_int — lowered by a simple stack machine (a0 accumulator, fp-relative frame slots, no register allocation yet) with a two-pass label assembler and a self-contained _start / print_int (itoa + ecall write) runtime. Only the top-level functions reachable from main are emitted, so the prelude is skipped entirely. Anything outside the slice (closures, strings, ADTs, heap) raises a clear Codegen_error. Verified byte-identical to the interpreter across recursion (factorial, Fibonacci, Ackermann), mutual recursion, gcd, short-circuit logic, and signed div/mod. lib/codegen_riscv.ml._
v0.1.129 — 2026-08-07
_Byte-safe strings on the Wasm backend — the arc that made C strings byte-safe (v0.1.127) now extends to Wasm. A str in linear memory is [i32 len][bytes][NUL]: the length header lives immediately before the pointer, so embedded NULs survive (("a" ++ chr 0 ++ "b") has length 3) while NUL-free strings stay host/C-interop compatible via the preserved terminator. __lang_strlen reads the header; a new __lang_str_alloc centralises header-writing allocation; every string producer (concat, substring, trim, rev, repeat, replace, upper/lower, escape/unescape, split/join, str_of_bytes, hex_of_bytes, char_at, show_int, JSON string cells, read_stdin) and every string literal now carries a header, and == / compare / starts_with compare over the header length instead of scanning to a NUL. Host glue (run_wasm.js, playground) writes the header for read_file / str_of_float / getenv / arg strings. This gives the browser mere-ruby playground binary-safe strings (pack / unpack1 / embedded-NUL length now match ruby byte-for-byte)._
v0.1.128 — 2026-08-06
_First native MIDI input capability — the seed of a MIDI dogfood. Six extern fn entry points (midi_init, midi_default_input, midi_open_input, midi_poll, midi_read, midi_close), backed by a PortMidi runtime in the C backend. The polling model (midi_poll then midi_read) matches the synchronous FFI shape that tcp_*/udp_* already use, and a whole MIDI message packs into one int, so — unlike the socket externs — the read side needs no byte arena._
_The surface is uniformly int -> int: the two conceptually-nullary calls take an ignored dummy 0. That is not cosmetic — the codegen synthesizes a first-class closure adapter (name(__x)) for every concrete A -> B extern, so a unit -> int signature would emit midi_init(__x) against a void-param C function and fail to compile. Keeping everything int -> int sidesteps it and matches the arena-FFI convention._
_The PortMidi runtime (and its #include <portmidi.h>) is emitted only when a program declares a midi_* extern (a new uses_midi gate, mirroring uses_tls), so non-MIDI native builds need no portmidi. Native-C only, like the socket capabilities; linking asks for -lportmidi the same way TLS asks for -lssl. Example: examples/midi_listen.mere echoes Note On/Off as note names (C4, A4, …) with velocity and channel._
_Verified: generated C compiles, links, and runs against a faithful PortMidi stub — the "no input device" exit path is clean, and a scripted event stream decodes correctly (Note On C4 vel80, Note Off C4, Note On A4 vel100). The gate excludes portmidi from non-MIDI programs (tcp_smoke: 0 refs)._
v0.1.127 — 2026-08-06
_The Wasm backend's int is now 64-bit — the receipt from "The Int That Stayed 32 Bits" came due. The trigger was exactly the documented one: a real dogfood (a Date.now()-driven clock, and mere-ruby running Ruby arithmetic) that needs 64-bit integers in the browser. Epoch-milliseconds (~1.75e12) overflowed i32: the clock trapped at int_of_float, and any Ruby snippet above 2^31 either failed to compile (literals, loudly) or couldn't run._
_The design is the uniform one the receipt named for long-running use: every value slot widens from 4 to 8 bytes — ints are true i64, pointers carry a 32-bit address zero-extended, wrapped back to i32 exactly at memory operations, tags/indices/the allocator stay 32-bit internally. The JS boundary keeps its 32-bit ABI (host imports are declared $name_h taking i32 pointers; generated in-module shims adapt), so hosts stay BigInt-free except where true i64 values cross: the closure call type (param i64 i64) (result i64) (JS glue passes BigInt(env)) and channel payloads (the ring is now BigInt64)._
_Verified: the four-backend parity suite is 59/0 including three new int64 cases (big-literal arithmetic, epoch divmod, int_of_float above 2^31, 64-bit bitwise); ctest 12/0; self-host fixpoint all-passed; a live clock prints the correct time on Node; and mere-ruby — a 17k-line Ruby interpreter — compiles to Wasm (500k lines of WAT), validates, and computes 1234567890123 + 1 correctly. Also fixed en route: the C backend's int_of_float truncated through a 32-bit (int) cast (epoch-ms came out as INT_MAX), and time / print_no_nl gained Wasm host wiring._
v0.1.115 — 2026-08-04
_Positioned write — the write half of the file API, and the forcing function for an on-disk store (a paged B-tree dogfood, mbtree, comes next)._
_file_pread (v0.1.83) could read an arbitrary window of a file, but there was no way to write one: write_file / write_file_bytes only replace a whole file. Three new builtins complete random-access file I/O, interp + C only (the LLVM/Wasm MVP backends have no filesystem and cleanly refuse):_
- _
file_openrw : str -> File— open a read/write handle, creating the file if
absent and never truncating an existing one (r+b, falling back to w+b)._
- _
file_pwrite : File -> int -> Vec[int] -> int— seek to the offset and write
the byte vec, extending the file past its end if needed; returns the count written._
- _
file_fsync : File -> unit— flush buffered writes to stable storage
(fflush + fsync), for commit points in a durable store._
_file_pread and file_close now also accept the read/write handle, so a store reads and writes through one file_openrw handle. On the interpreter the handle is a Unix.file_descr (V_rwfile); on C it is a single FILE*. Round-trips are byte-identical across interp and C._
v0.1.114 — 2026-08-04
_contrib/mlint grows from a one-rule demo into a real linter, and forces a C-backend codegen fix._
_mlint now carries three rules, all as dyn Rule trait objects — unused bindings, unused parameters, and shadowed bindings (the last threading its own scope environment) — with a two-method trait (rname + check). It reads a source path from args and lints that file (falling back to a built-in sample), so it runs on the interpreter and the native C backend; the LLVM/Wasm MVP backends cleanly refuse args/read_file (no filesystem)._
_C-backend fix (found by mlint, affects any dyn Trait): the arrow-type collector skipped polymorphic records' field types, but the struct emitter monomorphizes a generic trait dictionary Trait__dict 'a (left generic on Trait__pack's dictionary parameter) at the TyParam-erased default int, emitting Trait__dict_int. When a method's field closure type at 'a = int (int -> R) is instantiated nowhere else, the emitted struct referenced an undefined C type. mlint's check : 'a -> program -> diag list triggered it (int -> program -> diag list appears nowhere else); examples/trait_object.mere compiled only because its int -> int / int -> str closures exist elsewhere. Fixed by walking polymorphic-record field types at their monomorphized instances in collect_arrow_types._
_Ergonomics note recorded in contrib/mlint/README.md: a trait-object consumer must annotate its parameter as dyn Trait (fn (ru : dyn Rule) -> …) to select object dispatch; an unannotated fn ru -> check ru … is inferred with a Rule 'a => dictionary constraint instead._
_Full suite including the bootstrap fixpoint stays green — 2291 checks, 0 failures. Line-numbered diagnostics remain deferred: the self-host AST is position-less, and adding spans would ripple through the whole self-host compiler plus the bootstrap._
v0.1.113 — 2026-08-04
_A linter for Mere, written in Mere (contrib/mlint), plus two contrib/parser fixes it forced. mlint parses source into the shared self-host AST and runs lint rules over it — dogfooding the trait system on an AST-sized program: rules are dyn Rule trait objects, diagnostics derive (Eq, Ord) for dedup + sort, and Ord is a super-trait of Eq. Its one rule so far flags a let binding whose name never occurs in scope; it runs on all four backends._
_Forced upstream in contrib/parser:_
- _
parse_program_ast : str -> program— the parser only exposed
parse_str_program (a debug string); AST consumers (a linter, an analyzer) need the program value._
- _the parser defined its own
list_appendfixed totop_decl list, which
shadowed the prelude's polymorphic list_append for every importer (so list_append on any other element type failed to type). Renamed the private helper to append_top_decls._
_(Both are self-host compiler components; the full suite including the bootstrap fixpoint stays green — 2291 checks, 0 failures.) Recorded pain in contrib/mlint/README.md: the self-host AST is position-less (message-only diagnostics), and type T = Ctor; — a single nullary variant — parses as a type alias unless written type T = | Ctor;._
v0.1.112 — 2026-08-04
_Fix a parser declaration-table leak across programs parsed in one process. Pipeline.parse_program reset only imported_files, so the constructor / record / module / alias tables accumulated: a type Rect = { ... } record in one program left Rect registered, and a later program's constructor pattern Rect r then mis-parsed as a record pattern (expected '{' for record pattern). This bites any host that parses several programs in one process — the CI test binary hit it (a module Shapes { type Rect = {...} } test poisoning a later trait-object test's Rect constructor), aborting the run._
_Pipeline.parse_program now calls Parser.reset_decl_state () once before parsing the prelude (which re-registers its own types / constructors), giving each program a clean parser state. The REPL, which drives Parser.parse_program directly to accumulate definitions across lines, is unaffected. Regression test added; the full test binary now runs to completion (2289 checks, 0 failures)._
v0.1.111 — 2026-08-04
_Fix an inner-function over-capture on the LLVM and Wasm backends. When a lifted inner function A calls another lifted inner function B, A's captures are extended with B's (the transitive-capture closure) so A can forward them. But a capture of B that is bound inside A's own body — a let local, a nested-fn param, or a match-arm binder — is already in scope in A and must not be threaded in; otherwise A over-captures, and when A's host calls it the host is asked to pass a name it never had (use of undefined value %row on LLVM, inner-lifted capture row not in scope on Wasm)._
_The interpreter and C backend already excluded body-bound names (v0.1.48); this ports that exclusion to LLVM and Wasm. Surfaced by examples/sudoku.mere, whose inner cell (which fills a row) captures the match-arm variable row bound in its enclosing load — sudoku now runs on all four backends (previously llvm:MISCOMPILE / wasm:UNSUP). Locked by test/parity/inner_capture_match_binder.mere; parity 47/0, unit suite green._
v0.1.110 — 2026-08-04
_Structural == / != and < <= > >= on compound values (variant / record / tuple, including recursive ones like list) now work on the LLVM backend. Previously the LLVM backend refused structural == (a clean UNSUP) and mis-compiled structural ordering, so a program comparing compounds — e.g. deriving Eq/Ord for a variant key — only ran on interp / C / Wasm._
_Implemented as the LLVM siblings of the C backend's eq_<tag> / cmp_<tag> (and the interpreter's value_eq / value_compare): per-type define i1 @eq_<tag> and define i64 @cmp_<tag> (returning <0/0/>0) that recurse structurally over components — extractvalue for tuples/records, tag + boxed payload for variants (recursive variants via the pointer-to-node layout), and strcmp for strings — matching how show_<tag> already walks these shapes. The Cmp handler lowers ==/!= to a call to @eq_<tag> and ordering to @cmp_<tag> compared against 0; the needed types (and their transitive components) are collected into eq_types / cmp_types and emitted._
_Effect: deriving Eq / Ord for a variant or record key now compiles on all four backends, so ordset over such a key works everywhere. Locked by test/parity/struct_eq_cmp.mere (variant/record/tuple/list eq+cmp), derive_variant.mere, and two LLVM-IR test_basic assertions; parity 47/0, unit suite green. (The interp, C, and Wasm backends already implemented structural comparison — this brings LLVM to parity, so all four backends now agree.)_
v0.1.109 — 2026-08-04
_derive — generate trait instances from a trait's defaults. derive (Eq, Ord) int; (or single derive Eq int;) expands to one empty impl Ti T {} per listed trait; each empty impl inherits the trait's default method bodies. A trait is therefore derivable iff every method has a default — deriving a trait with a method that has no default is the ordinary "missing method" error._
_This makes structural instances a one-liner: a trait Eq 'a { eq : 'a -> 'a -> bool = fn a -> fn b -> a == b; } (default in terms of the builtin structural ==) is derivable for any key type, and derive (Eq, Ord) int; gives working Eq / Ord instances with no hand-written bodies. Pure sugar over the empty impl + default-method machinery (v0.1.101); adds the derive keyword and works inside module bodies too._
_contrib/ordset now carries structural defaults on its Eq / Ord traits and examples/ordset_demo.mere derives the int instances (its color key keeps a custom rank-based ordering). Note: structural == / < on a variant / record is still an LLVM-backend limitation, so deriving for such a type works on interp / C / Wasm but not LLVM (a pre-existing gap, independent of derive). Locked by test/parity/derive.mere and two test_basic assertions; parity 45/0, unit suite green._
v0.1.108 — 2026-08-03
_Trait objects: dyn Trait. A heterogeneous collection of values that all implement a trait, with dynamic dispatch — [dyn Shape (Circ 2), dyn Shape (Rect 3)] is a (dyn Shape) list, and a trait method called on a dyn Shape (area o) dispatches dynamically. A function that consumes objects annotates its parameter fn (o : dyn Shape) -> ... (a function inferred as Shape 'a => would instead expect a concrete dictionary-carrying value)._
_Implemented entirely as elaboration, with no new backend support: for an object-safe trait (every method takes the trait parameter as its single self argument and doesn't otherwise mention it — so area : 'a -> int qualifies, eq : 'a -> 'a -> bool does not), trait_elab auto-generates an object record Trait__obj of self-capturing method thunks plus a constrained packer Trait__pack. dyn Trait e is sugar for Trait__pack e; dyn Trait (type) is sugar for Trait__obj; and a method use on a value of object type lowers to reading and forcing the captured thunk. Because a dyn Trait is just a record of closures, every backend handles it unchanged. Adds the dyn keyword._
_This is ergonomic sugar over a pattern already expressible by hand (a record of self-capturing closures); the sugar removes the per-instance-type boilerplate. Locked by test/parity/trait_object.mere, two test_basic assertions, and examples/trait_object.mere (circles + squares in one list on all four backends); parity 44/0, unit suite green._
v0.1.107 — 2026-08-03
_Traits and impls may now be declared inside a module body. Previously a module body accepted only let / let rec / nested module (types were already allowed and kept global); trait / impl were rejected, so a reusable trait-based library could not be namespaced. They are now accepted and — like types — kept global (the module namespaces only its functions), so a consumer writes bare impl Ord T but calls M.of_list. The top-level trait/impl parsing was factored into shared helpers used by both the top-level and module-body parsers._
_Also fixes a spurious non-exhaustive-match warning for a type declared inside a module: the module qualifies constructor uses (M.Leaf) but the variant registry keys on the bare name, so the exhaustiveness checker now normalizes a qualified constructor to its bare last segment before comparing._
_Surfaced by the contrib/ordset dogfood — a generic sorted set (BST) over an Ord key, with a consumer (examples/ordset_demo.mere) that instantiates it at int and a user-defined color variant. The dogfood also hit a pre-existing limitation (a top-level polymorphic value binding like empty = Leaf : 'a tree has no use site to fix 'a and is rejected by the LLVM backend), worked around in the library by exposing empty as a thunk. Locked by test/parity/trait_in_module.mere and a test_basic assertion; parity 43/0, unit suite green._
v0.1.106 — 2026-08-03
_Extend v0.1.105's local-fn duplication to a single self-RECURSIVE local let rec f = fn ... in body used at several concrete types. Each type gets its own monomorphic copy, and the recursive self-call inside each copy is redirected to that copy, so the C and LLVM backends compile it (LLVM previously refused). The transform requires that f is not shadowed in either its body or its continuation, which keeps the self-call rename unconditional and safe. Mutual (multi-binding) local let rec groups used at several types are still left alone — a rarer remaining case. Locked by test/parity/local_poly_rec_multi_type.mere; parity green, unit suite green._
v0.1.105 — 2026-08-03
_A LOCAL polymorphic function used at several distinct concrete types now compiles correctly on the C and LLVM backends (let id = fn x -> x in (id 1, id 1.5) and friends). Both backends lift a local function to a single top-level function, so a multi-type use previously defaulted it to one type and miscompiled on C, or was refused on LLVM. The interpreter and Wasm already handled it._
_Fix: a pre-pass (duplicate_multi_use_local_fns) that, before lifting, splits such a local binding into one monomorphic copy per distinct concrete use type and rewrites each use to its copy — turning the unsolved "multi-instantiate a lifted local fn" problem into the already-solved monomorphic case for every backend. It is deliberately conservative: it fires only on a non-recursive local let (nested inside a function body, so top-level functions are left to the ordinary multi-instantiation machinery), only when the function is not shadowed and every use is at a concrete type. Implemented once and shared by both native backends._
_Locked by test/parity/local_poly_multi_type.mere (int / float / bool) and local_poly_multi_type_hof.mere (a higher-order local fn at two types); parity 39/0, unit suite green. Still open: the recursive local case and top-level mutually-recursive functions used at multiple types._
v0.1.104 — 2026-08-03
_Constrained recursive functions in a LOCAL let rec ... in — self- and mutually-recursive — now work. Previously only top-level let rec groups were handled; a local one failed with "ambiguous trait constraint". Two changes: the typer now records a local let rec binding whose scheme carries trait constraints into trait_local_constrained (it already did this for a local non-recursive let), and trait_elab threads the group's shared dictionary through every intra-group reference of a local let rec group, mirroring the top-level handling._
_The C and LLVM single-use monomorphization pass is extended to local let rec groups so a local constrained recursive function used at a single non-int type (e.g. float) emits at that type instead of defaulting to int. This is restricted to dictionary-taking (trait-constrained) members and excludes each member's own body from the use scan, so it cannot mis-specialize a genuinely polymorphic recursive function used at several types (e.g. the prelude's list_fold)._
_Works on all four backends at a single instance type (int / float / user variant). Locked by test/parity/trait_local_rec_self.mere, trait_local_rec_mutual.mere and two test_basic assertions; parity 39/0, unit 2283/0. (A local polymorphic recursive group used at several distinct types remains gated by the same pre-existing multi-instantiation limit as top-level and trait-free polymorphic recursion.)_
v0.1.103 — 2026-08-03
_Super-traits: trait Ord 'a : Eq 'a { ... }. A super-trait declares that any instance of the sub-trait must also be an instance of the super-trait. impl Ord T now requires impl Eq T (checked transitively, and for every super in a multiple-super list : Eq 'a, Show 'a); omitting it is a clear error rather than a confusing failure at a later use site._
_Method access needs no special dictionary machinery: Mere's inference records a separate constraint for every trait method actually used, so a generic function that uses both an Ord method and an Eq method on one value already receives both dictionaries (there are no signature-level constraint annotations that could under-specify this). This is why super-traits reduce, for Mere, to the declaration plus the well-formedness guarantee._
_As part of this, impl method bodies are now type-checked at their concrete instance type (param := target), so an impl body that calls another trait's method on the instance value — e.g. a super-trait method — resolves to that trait's concrete dictionary instead of leaving an unresolved dispatch variable. (Same-trait sibling calls are still inlined before type-checking, so no self-referential dictionary arises.) Locked by test/parity/trait_super.mere and three test_basic assertions; parity 37/0, unit suite green._
v0.1.102 — 2026-08-03
_Support top-level mutually-recursive constrained functions (a let rec f = ... and g = ...; group where the members require a trait). This used to be rejected ("mutually-recursive constrained function ... is not yet supported")._
_Intra-group references are typed monomorphically (before generalization), so they carry no constrained-use obligation and must be threaded by hand: a reference to a constrained group member — itself or a sibling — is applied to the dictionary parameter(s) that member expects. Because the group is typed monomorphically, mutually-recursive members share the dispatch variable(s), so those dict parameters have the same names as the current member's own and are in scope. This generalizes the single self-recursive case (which becomes the one-element instance of the same code path)._
_Scope: a single instance type works on all four backends. A polymorphic mutual-rec group used at two distinct types hits a separate, pre-existing backend multi-instantiation limitation (it fails the same way for trait-free mutually-recursive polymorphic code). Local (let rec ... in) constrained recursion — self or mutual — remains a distinct open path. Locked by test/parity/trait_mutual_recursion.mere and two test_basic assertions; parity 36/0, unit suite green._
v0.1.101 — 2026-08-03
_Trait method DEFAULTS, and impl method bodies that reference sibling methods. A trait method may now be written m : ty = expr; an impl that omits m inherits that default. Both this and an impl body that calls a sibling method (e.g. neq = fn a -> fn b -> if eq a b then ...) previously failed — the sibling reference had an unresolved dispatch type ("ambiguous trait constraint"), and resolving it to a dictionary field would have required the dictionary to reference itself, which Mere has no way to express._
_Both are solved the same way: trait_elab completes each impl Trait T before any type-checking — every method gets a source body (the impl's own, else the trait default, else a "missing method" error), and every reference to a sibling method name inside a body is replaced by that sibling's (recursively inlined) source body. Cyclic defaults are rejected. After completion each method is a self-contained body with no trait-method name references, so type inference sees ordinary Mere and the dictionary stays a plain, non-recursive record — every backend is unchanged. An impl may still override a default by providing the method. Locked by parity cases trait_sibling_method.mere / trait_default_method.mere and five test_basic assertions; parity 35/0, unit 2277/0._
v0.1.100 — 2026-08-03
_Fix a dictionary mix-up when a generic function carries two trait constraints on the SAME type variable (e.g. (Num 'a, Sh 'a) => 'a -> str). The elaboration's variable→dict-parameter map was keyed by the type variable's id alone, so the second constraint's dictionary parameter clobbered the first — and a Num method use resolved to the Sh dictionary, failing with "record Sh__dict has no field: add"._
_Fix: key the map by (variable id, trait). Since resolve_dict already knows which trait a method belongs to, each method use now selects the correct dictionary. A function constrained by multiple traits on one variable elaborates and runs correctly on all four backends. Locked by the parity case trait_multi_constraint.mere and a test_basic assertion; parity 34/0, unit 2272/0._
v0.1.99 — 2026-08-03
_Monomorphize single-use local polymorphic functions on the C and LLVM backends. A local let f = fn ... in ... is let-generalized by the typer, so its binding keeps an unresolved scheme while each use site instantiates a fresh concrete copy. Both native backends lift such a local fn to one top-level fn and default its residual type variable to int — so a local fn whose sole use is at, say, float was emitted as an int-typed C function and the float call site mismatched at compile time. (The interpreter and Wasm already handled the general case.) This is the long-standing "local polymorphic fn not multi-instantiated" limitation that the v0.1.98 trait local-let support ran into — reproducible without traits._
_Fix: a whole-program pre-pass (specialize_single_use_local_fns) run before fn-type resolution and inner-fn lifting. When a local fn is used at exactly one concrete type, it unifies the binding type with that use arrow; because the body's type variables are shared mutable union-find cells, this propagates into the body, so the lifted fn AND any generic callee inside it (e.g. list_fold) resolve concretely. Running before resolve_fn_types is essential — the top-level multi-instantiator only sees a generic callee's concrete use once the enclosing local fn's body is concrete (cf. the v0.1.28 poly-through-poly fix)._
_Effect: a constrained generic function defined in a local let now compiles on all four backends at any single instance type — int, float, or a user-defined variant. Multiple distinct use types on one local fn are left to the existing defaulting (a larger multi-instantiation increment). Locked by the trait-free-equivalent parity cases trait_local_let.mere (int) and trait_local_let_variant.mere (user variant); the four-backend differential harness stays 33/0 and the unit suite 2271/0._
v0.1.91 — 2026-07-31
_Real TLS on the native backend: tcp_starttls / tcp_starttls_verified are no longer stubs. Declaring either extern swaps in an OpenSSL-backed runtime — tcp_starttls does the handshake with SNI; tcp_starttls_verified adds peer certificate verification and hostname matching (with an optional CA-bundle path) — and tcp_read / tcp_write / tcp_close route through the per-fd SSL* when a socket has been wrapped, so the rest of a client is unchanged between HTTP and HTTPS. The whole thing is gated on a uses_tls flag set only when a program declares a starttls extern: a plaintext TCP program emits zero OpenSSL references and still links with just -lm, so TLS's external dependency is opt-in. A TLS program compiles with, e.g., -I$(brew --prefix openssl@3)/include -L.../lib -lssl -lcrypto.
Surfaced and verified by the mhttps dogfood — an HTTPS GET client in pure Mere: it fetches https://example.com/ and https://api.github.com/ (the latter requiring a valid verified handshake) and prints HTTP/1.1 200 OK for both. test_basic guards that starttls pulls in the OpenSSL runtime and that a plaintext program does not. TLS on the Wasm host is still separate (browsers do TLS transparently); native LLVM shares the C runtime. suite passed._
v0.1.90 — 2026-07-30
_mere install detects same-major version conflicts instead of silently picking one. A module path may resolve to only one revision in a build; if two packages in the dependency graph demand the same path at different revs, the installer used to fetch both and let the second overwrite the first (last-writer-wins). It now tracks the resolved sha per module path and fails with both revisions named, asking for explicit reconciliation (pin it in the top-level mere.toml). This is the correct answer for Mere's model: it pins exact revisions with no version ranges, so there is nothing to minimize over — the npm/cargo MVS problem does not arise. Incompatible majors sidestep the conflict entirely by living at different module paths (.../v2, SIV, v0.1.89), which the check leaves untouched (distinct paths → no conflict). Together with v0.1.88–89 this closes the version-resolution axis of Q-013 for an exact-pin package manager. Hand-tested against a diamond (two libs pinning the same library at different revs → conflict; the SIV app with distinct paths → installs clean). suite: 2255 passed / 0 failed._
v0.1.89 — 2026-07-30
_Two modules with the same name now coexist correctly on every backend — the language-side half of Go-style Semantic Import Versioning (SIV). A library and its /v2 (an incompatible major) both name their module the same (module Greet { ... }), so their members desugar to identically-qualified top-level names (Greet.hello defined twice, of different types) that shadow by declaration order. The interpreter already honoured that — a closure captures the env at its definition and the env prepends, so a reference binds to the most-recent prior definition — but the native backends resolve a top-level name globally and mis-assigned one version's body to the other (a C build emitted Greet.hello : str -> str with the other version's closure-returning body). New pipeline pass Ast.uniquify_toplevel_module_shadows walks the decls in order and, when a dotted (module-qualified) name is redefined, alpha-renames the redefinition (Greet.hello → Greet.hello__v2) and rewrites later references, so all four backends see distinct symbols. Only dotted redefinitions are touched, so ordinary programs are unaffected.
Verified with the version-resolution dogfood — an app whose two dependencies pull incompatible majors of a shared library via SIV distinct paths (.../mgreet and .../mgreet/v2): it now prints the v1 and v2 results side by side on interp, C, LLVM, and Wasm. This closes the native gap that the dogfood surfaced; combined with the full-path installer (v0.1.88), SIV works end to end. Minimal-version selection (MVS, for compatible same-major demands) remains the one deferred package axis. suite passed._
v0.1.88 — 2026-07-30
_mere install grows up to match the Go-style import model (Q-013): full-path layout, cross-repo transitive resolution, and a verifying lockfile. Driven by the first genuinely multi-repo dogfood — a 3-repo transitive graph mcalc → mbigfmt → mbignum where the app never names the leaf. Three gaps surfaced and were fixed:_
- _Full-path layout. Installs went to
.mere_modules/<bare-name>/, but a
Go-style import resolves to .mere_modules/github.com/owner/repo/, so even a direct dependency failed to resolve. The installer now reads each fetched package's own mere.toml [package] path and installs under it (bare-name fallback for legacy packages)._
- _Cross-repo transitive deps. The installer only followed
../relative
imports (monorepo siblings); a dependency declaring its own [dependencies] in another repo was never followed. It now reads each fetched package's [dependencies] and queues them, so transitive cross-repo deps are pulled in (deduped on the (git, rev, subdir) coordinate for diamonds)._
- _Verifying lock.
mere.lockalready recorded resolved shas + content
hashes; now a re-install parses the existing lock and, if a pinned (git, rev) coordinate produces a different hash, fails loudly (go.sum-style tamper/corruption detection) instead of silently building against changed content._
_The resolved lock records the full transitive graph, so mcalc's lock pins mbignum even though mcalc only depends on mbigfmt. Verified end-to-end: mcalc fact/fib N matches python on interp and C, reading both deps out of the full-path .mere_modules/. Unit tests cover the [package] path parse and the write_lock/read_lock round-trip; the fetch path is git-integration-tested by hand. suite: 2248 passed / 0 failed._
v0.1.87 — 2026-07-29
_A user top-level binding named main now compiles on every backend (finishing the follow-up left open in v0.1.86). Mere has no main convention — the entry point is the file's trailing expression — so a main binding is just an ordinary value that happens to share the synthesized entry's name. The C backend already mangled it (mu_main), but LLVM and Wasm emitted the raw name and hit a duplicate-main link/assemble error. Rather than patch each backend, the fix is one backend-agnostic pass in the pipeline: Ast.reserve_toplevel_main alpha-renames a top-level main to a reserved name (__mere_user_main) via the existing scope-aware rename_free_vars, so an inner main (a local let or a parameter) still shadows and is untouched. Verified: let main = fn () -> 42 in main () prints 42 on interp, C, LLVM, and Wasm. The v0.1.86 note's "not fixed" caveat is superseded. suite passed._
v0.1.86 — 2026-07-29
_str_eq on the LLVM backend. The interpreter and C backend had string equality, but the LLVM backend never defined it, so any LLVM-compiled program using str_eq failed at emit with "unbound variable: str_eq" (surfaced by the bignum and mpath dogfoods, both of which pattern on single characters). The 2-arg call now lowers to a new @__lang_str_eq runtime — a byte compare over two NUL-terminated strings returning i1 — mirroring the C backend's strcmp path. Guarded in test_basic; verified equal/unequal/empty/prefix cases match the interpreter on a compiled LLVM binary.
Not fixed here (documented in the mpath dogfood's PAIN): the Wasm backend emits a user top-level binding named main as $main, colliding with the exported entry $main ("redefinition of function $main"). It is narrow (only a literal main binding, only on Wasm) with a trivial rename workaround; a proper fix mangles or reserves the entry name and is left as a follow-up. suite passed._
v0.1.85 — 2026-07-29
_A module-level value binding compiles on the C backend, and a new contrib/bignum library. Writing bignum surfaced the bug: let base = 1000000000 inside module Bignum { ... } carries a dotted name (Bignum.base), and the C let-emitter used the raw binder name for its internal temporaries, so it emitted __auto_type __let_tmp_Bignum.base = ... — a . in a C identifier, which does not compile. Every earlier contrib module bound only functions (whose names already route through name-mangling), so a module value binding had never been exercised. Fix: flatten the dot for the __let_tmp_ / __let_result_ temp names (Bignum__base); ordinary undotted names are untouched, so only the previously-broken module-value case changes. Guarded in test_basic.
contrib/bignum/bignum.mere is a reusable arbitrary-precision natural-number library — little-endian base-1e9 limbs as a persistent int list (from_int / add / mul_small / mul / cmp / to_str / fact / fib). Base 1e9 keeps limb products under 2^63 on the 64-bit backends. examples/bignum_demo.mere imports it by full path and prints 100!, fib 200, and a product that match python's bignum exactly on interp and C. Backend reach: interp + C exact; Wasm runs add/fib but mul overflows its 32-bit int (a base-1e9 product is ~1e18); LLVM rejects a polymorphic inner-closure capture (a known monomorphization gap) — both documented in the library README. suite: 2244 passed / 0 failed._
v0.1.84 — 2026-07-29
_Fire-and-forget threads: detach : ThreadHandle -> unit. spawn returns a joinable handle, and the only way to reclaim a worker's resources was join, which blocks. A server's accept loop that spawns one handler per connection and never joins therefore leaked a joinable thread per connection, so a long-running server slowly exhausted thread resources. detach h releases the thread without waiting for it (pthread_detach on the C backend; a no-op on the reference interpreter, whose domains are not the server target). Surfaced by the mhttpd dogfood (a concurrent HTTP/1.1 server): with detach plus a fixed pool of arena buffers handed hand-to-hand over a channel, mhttpd serves 400 sustained concurrent requests where the naive version aborted at ~256 (each connection had leaked a fresh 64 KB from the no-free byte arena). No new limitation in the byte arena itself — its loud abort-on-exhaustion already told the server to pool and reuse buffers; detach is the missing concurrency primitive. suite passed / 0 failed._
v0.1.83 — 2026-07-29
_Positioned reads: file_pread : File -> int -> int -> Vec[R, int]. file_pread handle offset len seeks to offset and reads up to len bytes from an open handle, returning them as an int vec (same byte representation as read_file_bytes). Reads fewer than len bytes only at EOF (partial tail). Until now the only binary read was read_file_bytes, which loads the whole file — fine for a WASM module inspector, useless for a random-access format where the point is to touch only the pages you need. This is the capability a B-tree file format wants: read one page at an offset without paying for the whole file. Interp uses seek_in/really_input on the in_channel; the C backend emits __lang_file_pread (fseek + a bounded fgetc loop building a region vec_int). Scope is interp + C, inheriting the file I/O family's boundary — a program that preads first opens the handle with file_open, which already refuses cleanly on LLVM/Wasm ("v0.1.59 scope = interp + C"), so file_pread needs no separate unsupported arm there. Surfaced by the msqlite dogfood (a read-only SQLite reader): it now prints SELECT * FROM t byte-for-byte against sqlite3 on both interp and C, reading the header, sqlite_master, and a table's leaf page by positioned reads. suite: 2242 passed / 0 failed._
v0.1.82 — 2026-07-29
_The LLVM backend's region allocator grows the arena instead of overrunning it. __lang_region_alloc was a pure bump — add the aligned size to the top pointer and return, with no bounds check — so once the default 4 MB arena filled, allocations ran past the malloc'd buffer and corrupted the heap. An allocation-heavy program (a per-pixel renderer that materializes a float triple per pixel) crashed with SIGSEGV at larger image sizes on LLVM while interp, C, and Wasm all produced the same checksum; the C backend had gained bounds-checked, block-chained growth in v0.1.25 (found by a long-running server) but the LLVM runtime never received it. The LLVM %__lang_region struct now carries a 4th blocks field (a chain of malloc'd blocks, each with a 16-byte header holding the prev link so the data that follows stays 16-aligned); __lang_region_alloc compares top + aligned against base + cap and, when it would overrun, chains on a geometrically larger block (doubling until it fits) via a new __lang_region_add_block helper; __lang_region_init seeds the first block and __lang_region_free walks the chain. Blocks never move, so pointers into earlier blocks stay valid across growth — matching the C semantics exactly. New guard test/parity/region_growth.mere folds a checksum over ~6 MB of live region allocations (two shallow loops so it exercises arena growth, not stack depth) and now matches across all four backends. Guard verified by reverting the fix: the LLVM column goes DIFF (empty output from the crash) on the old allocator and returns to MATCH with the growth in place. suite: 2240 passed / 0 failed; parity 28/28 (was 27)._
v0.1.81 — 2026-07-29
_Go-style full-path imports adopted in-repo: the self-host resolver learns the module path, and contrib's cross-package imports migrate. v0.1.80 taught the OCaml resolver to resolve an import under the project's declared module path (mere.toml [package] path) to local files; but the self-host toolchain has its own import inliner (inline_imports_in in contrib/codegen), whose resolve_import_path only knew base-dir-relative resolution — a first migration attempt made the self-host tests fail with a doubled path (contrib/eval/github.com/.../ast.mere). resolve_import_path now mirrors the OCaml resolver with its signature unchanged: an import whose first segment looks like a host name (contains a dot, not dot-relative) walks up from the importing file probing for a mere.toml that declares a matching module path and resolves module-root-relative; anything else keeps the historical behavior, and plain relative imports never probe. A missing mere.toml reads uniformly as "" on every backend (the language-level read_file fails catchably under try_or on interp/C; the Wasm host returns an empty string — its ENOENT log is now silent since a miss is an expected probe result). The helpers follow the file's Phase 54.32 style (inner loops hoisted to top-level rec fns with explicit args, the wasm-codegen capture workaround). With that in place, the repo declares path = "github.com/merelang/mere" in a root mere.toml and contrib's 11 cross-package ../ imports (typer/fmt/codegen/eval -> parser, http -> log, feed -> xml, site -> markdown/path, webhook -> http) migrate to full-path spelling — the same import now works in-repo (module-path-local) and vendored (.mere_modules/<full-path>/). site/playground keeps its own build pipeline untouched. suite: 2238 passed / 0 failed; self-host bootstrap fixpoint all-passed; ctest 12/12; parity 27/27._
v0.1.80 — 2026-07-29
_Module-path-local resolution for Go-style full-path imports (Q-013, compiler side). A project declares its module path in mere.toml ([package] path = "github.com/owner/repo"); the OCaml resolver walks up to the nearest such mere.toml and resolves an import that starts with the declared path to local files relative to the module root. External consumers already resolved full-path imports via .mere_modules/<full-path>/ (the walk-up resolver handled deep paths unchanged) — this adds the in-repo half, so a package's cross-package imports use the same spelling in-repo and when vendored. Resolution order: module-path-local, importer-relative, .mere_modules walk-up, -I/MERE_PATH; a project with no declared path is unaffected. Two unit guards; packages.md documents the convention. suite: 2238 passed / 0 failed._
v0.1.79 — 2026-07-28
_Doc-only: memory-model.md documents the heap-element overwrite leak in copy-on-store containers (a hot loop overwriting the same vec_set/map_set slot with fresh strings grows O(writes) — the container region is bump-allocated, so old copies are unreclaimable; measured ~550 MB for 4M overwrites; scalar elements unaffected). Eliding the copy needs type-level region tracking on str (deferred); prefer scalar slots or StrBuf reuse in hot-overwrite loops._
v0.1.78 — 2026-07-28
_A race/cancellation example, and the finding that the structured-concurrency "select" gap is smaller than assumed. channel_recv_timeout ch 0 turns out to be a general non-blocking try-recv (empty -> None, ready -> Some, verified on interp and C), which makes poll-based select and cooperative cancellation expressible with the primitives already in the language: examples/race.mere spawns N workers, takes the first to finish via a shared results channel, then broadcasts one cancel token per worker that each worker observes with a 0-ms recv on every step and stops early. So the only genuinely-missing piece of E-1 is a blocking multi-channel select (an efficiency win over busy-polling), not a capability gap — it stays deferred. (channel_recv_timeout is interp+C scope, v0.1.48, so the example is interp+C; LLVM/Wasm report it cleanly unsupported.) No compiler change. suite: 2236 passed / 0 failed._
v0.1.77 — 2026-07-28
_Map-accumulator memory fix (C backend), surfaced by a word-frequency dogfood. Counting words into a Map[str, int] over a 37.6 MB file with only ~13 distinct words held 62.5 MB of RSS — O(file), not O(distinct keys). Isolated to map_set: copy-on-store (v0.1.30) deep-copied the key into the map's region UP FRONT, before the hash lookup, so every update to an existing key leaked one key copy into the never-reclaimed bump region. A pure churn of 8M sets over 3 keys reproduced it (62.5 MB). The fix copies only what is actually stored: hash and key comparison use the caller's (content-identical) key, an existing-key update copies just the new value, and only a fresh insert copies the key. Churn and word-count both drop to ~1.3 MB (~45x), same output. This is the ubiquitous counter / histogram / accumulator pattern (a KV server, a frequency table), previously O(total writes). LLVM and Wasm were already flat here (they do not copy-on-store). Added an in-process guard that the key copy sits after the lookup loop. suite: 2236 passed / 0 failed._
v0.1.76 — 2026-07-28
_More parity coverage (test/parity/ 22 -> 26) and a documented known divergence. Added four shapes that agree across all four backends: an or-pattern match arm with a shared binding, functional record update ({ base | f = e }), float builtins with int_of_float, and a variant whose constructors carry different payload shapes. Probing the divergence-prone corners recorded their status: a nested let-record/constructor pattern (let pt { x = a } = p) is a parser limitation (rejected before codegen, not a backend gap); ref/:= mutable cells are not a Mere idiom. One genuine but already-known correctness divergence was reconfirmed and is deliberately NOT in the pass/fail corpus (it would be a permanent red): a 64-bit integer computation (100000 * 100000 = 10^10) is correct on interp and C but silently wrong on LLVM and Wasm, whose integers are i32 — the documented i64-widening limitation (a non-goal per the earlier LLVM assessment). No compiler change. suite: 2235 passed / 0 failed._
v0.1.75 — 2026-07-28
_Port the v0.1.70 referenced-but-unresolved poly-fn recovery to the Wasm backend. A polymorphic helper whose arrow keeps a residual type variable at every use site (e.g. result_and_then applied only to Ok, so the error type never grounds) is dropped by the resolver as unused — but if emitted code still references it, the direct call site emits call $<name> to a function that was never defined, which C fixed in v0.1.70 but Wasm still hit, producing an invalid module (undefined function variable "$result_and_then"). The parity harness (v0.1.73) surfaced it. Wasm's resolve_fn_types now runs the same recovery fixpoint C does: after the normal resolution, scan the emitted spine (Codegen_c.find_live_arrow, which accepts arrows with tyvars and skips dropped fn definitions) for live references to still-unresolved skels, and emit each with Codegen_c.deep_erase_tyvars erasing residual tyvars to int (both helpers are backend-agnostic and reused directly). The unconstrained-error result program now emits a valid module and runs on Wasm (== interp/C == 42); LLVM continues to report it a clean codegen error (documented subset limit). Added test/parity/result_residual.mere and an in-process guard asserting the Wasm definition is emitted, not just called. suite: 2235 passed / 0 failed._
v0.1.74 — 2026-07-28
_Parity corpus expansion (test/parity/ 9 -> 21) plus a harness classification fix. Added twelve diverse self-contained programs — nested tuple pattern, prelude option/result/list helpers, nested-variant match, string/char ops, curry-3, value shadowing, negative div/mod, tuple-capturing closure, boolean short-circuit — each run through all four backends. All 21 now agree with the interpreter. Two backend divergences surfaced along the way: (1) a nested let-tuple pattern (let ((a,b),(c,e)) = t) is rejected by C, LLVM, and Wasm with a clean "not supported in <backend> codegen subset — use match" (a consistent, documented limitation); the harness's emit-classifier now recognizes that phrasing as UNSUP rather than a hard failure, so it is not a false red. (2) A prelude result helper left with a residual (error) tyvar — only Ok used, so the error type never grounds — makes the Wasm backend emit a call to an undefined function (invalid module), where the C backend recovers it (v0.1.70) and LLVM errors cleanly; the corpus uses a fully-concrete (int, str) result instead, and the Wasm gap (it should recover like C or error like LLVM, not emit an invalid module) is recorded for a follow-up. No compiler change. suite: 2234 passed / 0 failed._
v0.1.73 — 2026-07-28
_A four-backend differential (parity) harness — scripts/parity.sh + test/parity/ — plus the resolver bug it immediately caught. The harness runs each program through every backend (interp / C / LLVM / Wasm) and diffs stdout against the interpreter, classifying each backend MATCH / DIFF / MISCOMPILE / UNSUP (clean "unsupported" at emit) / SKIP (toolchain absent). It exists to catch the "interp-accepts / backend-rejects" and "backends-disagree" family before a dogfood stumbles on it. On its first run it found one: a top-level fn named f taking a tuple and matching on it compiled on interp / LLVM / Wasm but MISCOMPILEd on C. Root cause: the concrete-arrow discovery that drives per-instantiation monomorphization (find_concrete_arrow / find_all_concrete_arrows_in / find_live_arrow) walked into the bodies of resolved poly helpers ignoring binder scope — so list_fold / list_map's parameter f was read as a use of the user's top-level f, forcing a bogus int -> int -> int monomorphization that treated the tuple parameter as curried (.f0 / .f1 on a scalar). The fix makes those scans skip a Fun parameter that shadows the searched name; only the Fun binder is treated as shadowing, because a let / let-rec binding the name may itself be the poly fn's definition whose use sites live in its body (narrowing to Fun keeps chained multi-instantiation discovery working). This is the same name-collision family as the earlier index / y0 param cases, but in the monomorphizer rather than name mangling. Added a corpus of nine self-contained parity programs (arithmetic, recursion, let-pattern, tuple/ADT/record match, closures, strings, mutual recursion) and an in-process guard for the fixed mono. suite: 2234 passed / 0 failed._
v0.1.72 — 2026-07-28
_contrib/stream — region-scoped line streaming combinators, closing the ergonomic side of the strings-lifetime gap. A str carries no region tag, so the escape checker must assume any line read in a streaming loop may escape and keeps it in the program-lifetime region; a naive line loop therefore grows to O(file). The reclaim machinery to avoid this has existed since v0.1.31 (a region R {} block redirects the thread-local current region, and an escape-clean block result — an int/bool/unit — is copied out while the block's scratch, including the line string, is freed on release), and the memory-model doc measured it, but there was no reusable combinator, so a streaming tool had to hand-roll the per-line region (mgrep's line loop simply didn't, and grew to ~file size). module Stream { each_line, count_lines } packages the pattern: each_line path cb runs a side-effecting str -> unit callback per line, and count_lines path pred returns a match count, each processing the line inside a per-line region block. Measured on a 57 MB / 1,000,000-line input, native C backend: peak RSS 61.7 MB (naive loop) -> 1.4 MB (combinator), a 44x drop, same output. Added examples/stream_lines.mere and an in-process codegen guard asserting the region redirect -> file read -> restore -> release ordering that makes the reclamation hold. The deeper fix — a type-level region tag on str so region-scoped strings can also be stored into outer containers — remains deferred until a dogfood forces it (the streaming case, which is what has recurred, is covered by this). Backends: interp + C (per-line file input is not implemented on Wasm/LLVM). suite: 2233 passed / 0 failed._
v0.1.71 — 2026-07-28
_C-backend hardening: two latent "compiles-to-C-then-fails" bugs fixed, plus the missing compile-and-run test path that let this family recur. (1) The _as_value closure adapter every top-level fn gets now sanitizes its parameter name through c_safe_name — a source parameter named like a C keyword (case, default) previously emitted an invalid C parameter declaration in the wrapper even when the fn was never used as a value, since wrappers are generated for all top-level fns. (2) An extern used in value position (passed to a higher-order fn, not directly applied) now lowers to a closure adapter __ext_<name>_as_value that calls the raw FFI symbol, instead of the mangled mu_<name> (undeclared — a bare extern is a raw C function, not a closure struct). This is the reference-side twin of the v0.1.61 capture fix; restricted to a simple A -> B signature, a curried/higher-order extern-as-value is now a clear compiler error pointing at fn x -> name x. (3) scripts/ctest.sh + test/ctests/ add a native-backend compile-and-run differential harness: each program is emitted to C, compiled with the C compiler, and (for extern-free programs) run and diffed against the interpreter; it also emits Wasm and assembles it with wat2wasm when available. The in-process suite only ever inspected the emitted C as text, so undeclared-identifier and closure-type failures escaped it; the harness compiles the emitted code for real. The corpus covers the family across both backends — reserved-name/keyword params, extern-as-value and extern-in-closure, top-level-fn-as-value, inner recursive uncurrying, mutual recursion, tuple capture, and multi-variable/deeply-nested captures (the Wasm backend, whose identifiers can't collide with C keywords and which already errors cleanly on extern-as-value, was confirmed free of the two C miscompiles). Two in-process string guards for the fixed regressions were also added. suite: 2232 passed / 0 failed._
v0.1.70 — 2026-07-27
_A referenced-but-never-concretized poly fn is now emitted (tyvars erased) instead of silently dropped. resolve_fn_types treated every fn without a concrete arrow as unused and skipped it — but a fn whose arrow keeps a residual tyvar at every use site (e.g. an unannotated wrapper taking a producer whose result is the bottom type of an endless loop) is NOT dead: call sites still emit a direct call, which failed at the C compile with an undeclared identifier. The fix is a recovery pass after the resolution fixpoint: scan the emitted spine — the program expression minus top-level fn definitions, plus the bodies of emitted and recovered fns (their own fixpoint) — for live references, and emit such fns with deep_erase_tyvars (residual tyvars become int, the representation the v0.1.69 emission erasure already names). The liveness scan must skip fn-definition bindings: scanning the whole expression would resurrect every generic prelude helper referenced from other dropped helpers' bodies. Downstream, the type- instance collectors learned the same erasure so recovered generic bodies register the instances their erased emission references: Channel / Vec / OwnedVec / Map element types erase before the concrete gate in c_type_of, and tuple-shape / mono-variant collection registers erased shapes (dead extras dedup by name). An unannotated polymorphic generator wrapper called with an endless producer now compiles and runs end-to-end. suite: 2230 passed / 0 failed._
v0.1.69 — 2026-07-27
_Residual type variables are erased at C codegen instead of rejected. A type variable that survives to codegen is either dead — the bottom result of a function that never returns, such as an endless generator loop — or genuinely unconstrained; no operation ever inspects such a value, so any representation works. ty_tag now names it int and c_type_of emits long long (the representation the resolver's use-site naming already assumes), instead of raising "unsupported C codegen type element: 'a". This fixes two failures found by a concurrency probe: an endless producer (let rec go = fn a -> fn b -> ... go b (a + b)) killed compilation, and a top-level fn whose body contained such a loop was silently dropped from resolve_fn_types while call sites still referenced its _as_value wrapper (undeclared identifier at the C compile). Both previously needed source workarounds (grounding the type with an unreachable unit branch / eta-expanding at the call site); the natural spellings now compile and run. Two artificial rejections became working programs and their tests were converted to positive assertions: an uninstantiated polymorphic variant now emits a concrete int instance, and a lifted closure capturing a tuple compiles and runs correctly (verified against the interpreter). suite: 2230 passed / 0 failed._
v0.1.68 — 2026-07-27
_C-backend maps get a hash index: map_get / map_has / map_set lookup is now O(1) amortized instead of a linear scan. A concurrency probe (an actor-owned "visited set" benchmark) showed the old cost dominating everything else: 100k check-and-insert ops against a ~30k-entry map took 5.8 s (~58 µs/op, ~300x the 97-entry case) — the map, not the messaging, was the bottleneck. The struct keeps its insertion-ordered keys/values arrays (so map_iter order, map_len, and map_delete's shift-remove behavior are observably unchanged) and adds an open-addressing index of array positions: linear probing, power-of-two capacity, rehash at 0.7 load, rebuilt after a delete (deletes shift positions and are rare). Key hashing mirrors the structural key-equality emitter — splitmix64 for scalars, FNV-1a for strings, recursive combination for tuple / record / variant keys — so keys equal under key_eq always hash equal. Same benchmark after: 10 ms (~580x). A 5k-entry grow/rehash/delete/iter functional check and the full suite verify behavior parity; the interpreter's map is untouched (correctness-first reference). suite: 2230 passed / 0 failed._
v0.1.67 — 2026-07-18
_A caught fail is now silent on the C backend, found by the mere-ruby dogfood's exception milestone. mere-ruby implements Ruby begin/rescue on top of fail + try_or: a raise unwinds via fail, and try_or catches it. But the C runtime's __lang_fail_impl printed fail: <msg> to stderr unconditionally, before checking whether an active try_or would catch the longjmp — so every rescued exception leaked a stderr line, even though the program continued correctly. A caught failure is control flow, not an error; it must be silent. The fix reorders the helper to longjmp first and print only when the failure is genuinely uncaught (about to abort). The LLVM backend and the interpreter were already silent-on-catch, so this also removes a cross-backend divergence. suite: 2230 passed / 0 failed (1 new test)._
v0.1.66 — 2026-07-18
_A C-backend duplicate-definition bug, found by the mere-ruby dogfood's first method milestone. Adding def / method calls turned the interpreter's evaluator into one large mutual-recursion group threading two Maps (locals and methods) through eleven functions, and the C backend refused to compile it: eighteen functions were each emitted twice, a "redefinition" error. The cause was in per-instantiation specialization. A polymorphic function's specialization list is grown, across resolution passes, from one concrete arrow type per use site. When a function is used from many sites, arrows that differ only in a region type variable — which the mangled-name tag erases — accumulate as distinct specs that all mangle to the SAME C symbol, so one fn_decl was emitted per spec and the identical definitions collided. The two later re-scan branches already deduped their arrows; the fix dedups the spec list by its emitted C symbol at the single emission chokepoint, so same-symbol specs collapse while genuinely distinct instantiations are preserved. With the fix mere-ruby's evaluator compiles clean and runs def / recursion / return byte-identical to ruby. suite: 2229 passed / 0 failed (1 new test)._
v0.1.65 — 2026-07-18
_Shortest round-trip float formatting, forced by the newest dogfood. The first program mere-ruby (a Ruby subset interpreter in pure Mere) could not print was puts 0.1 + 0.2: Ruby prints 0.30000000000000004, but str_of_float formatted every float at 12 significant digits, printed "0.3", and the original double was unrecoverable from the string — float_of_str (str_of_float x) was not x. All four backends (the interp's format_float, the C runtime helper, the LLVM IR helper, and the Wasm JS hosts) now format at 12 digits first — every value that 12 digits already represented faithfully keeps its exact old rendering, so nothing else changes — and widen toward 17 until the string parses back to the same double, the same shortest-round-trip contract Ruby, JS, and Python print with. Reading the four implementations side by side also surfaced a real pre-existing divergence: the LLVM helper appended a bare "." to whole-valued floats ("100.") where every other backend renders ".0" ("100.0"). Fixed in the same slice; the four backends were verified byte-identical on a shared corpus. suite: 2228 passed / 0 failed (7 new tests)._
v0.1.64 — 2026-07-18
_A backend gap closed, found by the medit dogfood. read_lines : str -> str list type-checked and ran under the interpreter, but the C backend had no arm for it — so a compiled program that read a file into lines emitted an undefined mu_read_lines and failed to link. The surface language promised something the native target could not deliver, and the type checker could not see it. The fix adds __lang_read_lines, a helper that matches the interpreter's input_line semantics exactly: split on newlines, drop a single trailing empty element when the file ends in '\n', and return the empty list for an empty file — so "a\nb\n" and "a\nb" both give ["a"; "b"], "" gives [], and "\n" gives [""]. Verified line-for-line against the interpreter on those edge cases. suite: 2221 passed / 0 failed (2 new tests)._
v0.1.63 — 2026-07-18
_A native monotonic clock, so Mere can time itself. Every measurement in this project so far shelled out to the time command; now_ms (a self-contained native FFI over clock_gettime(CLOCK_MONOTONIC), emitted like the tcp_ / udp_ externs) returns milliseconds since an arbitrary epoch, and that is enough to build a benchmark harness in pure Mere. The dogfood is mbench: it runs a kernel enough times to span a target wall-clock window, then reports iterations, total ms, and ns/iter, threading a checksum through the loop so the optimizer cannot delete the work. Writing it surfaced the universal benchmark trap first-hand — a kernel that ignores the loop counter is loop-invariant and clang -O2 hoists it out (a plain summation was even strength-reduced to i + C, reported as 0 ns/iter). The fix is the universal one: thread the counter into every kernel's input. A quiet observation falls out — Mere-on-C inherits clang's optimizer wholesale, for better (real kernels run fast) and for worse (arithmetic-reducible kernels vanish). suite: 2219 passed / 0 failed (1 new test)._
v0.1.62 — 2026-07-18
_A native UDP FFI, opened by a DNS resolver. mkv and mhttp used TCP; the new dogfood, mdns, is the first datagram-socket program, and it needed a capability that genuinely did not exist: udp_open / udp_send / udp_recv, connected SOCK_DGRAM sockets that send and receive one datagram at a time through the flat arena, reusing the protocol-agnostic tcp_close and tcp_set_timeout. On top of them mdns builds a DNS query packet byte by byte in the arena — a 12-byte header, length-prefixed labels for the QNAME, QTYPE/QCLASS — sends one datagram to a resolver, and parses the answer section, stepping past compressed names and reading each A record's four IPv4 bytes. The same binary-packet construction and length-prefixed parsing as a gzip block, but on the wire. Verified against dig +short on several names (including this project's own GitHub-Pages A set) and two resolvers. suite: 2218 passed / 0 failed (2 new tests)._
v0.1.61 — 2026-07-18
_An extern-in-closure capture bug, found the first time the client side of the TCP FFI was driven. The dogfood is mhttp — an HTTP/1.1 client in pure Mere over raw tcp_connect / tcp_write / tcp_read (mkv used the server side; this is the first tcp_connect). Its send_all helper called tcp_write from inside an inner recursive closure, and the C backend refused to compile: the closure-lift analysis captured tcp_write as a free variable and referenced it through the env as the namespaced mu_tcp_write, while the extern itself is emitted raw. The cause: the lift's "globals to exclude from capture" set held top-level fns and builtins but not extern fns — so an extern used as a value inside a helper was wrongly treated as a captured local. Externs are globals, referenced directly in the generated C, so they now join that excluded set. With the fix, mhttp parses status lines, case-insensitive headers, Content-Length bodies, and chunked transfer-encoding, verified byte-for-byte against curl (local server, both framings) and against a live server. suite: 2216 passed / 0 failed (2 new tests)._
v0.1.60 — 2026-07-18
_int_of_str semantics pinned across all four backends, caught by a Result-pipeline probe. The probe wrote an ordinary config pipeline — parse three fields, validate, combine — and the same program printed different errors on the interpreter and C: "not a number: abc" versus "out of range: 0". The cause: the interpreter's int_of_str raised on invalid input (so a try_or bridge caught it), while C was a bare atoll, LLVM a bare atoi, and Wasm a hand-rolled stop-at-first-non-digit loop — all three silently returning 0 or a partial prefix. The C emitter's own comment admitted it: "Fail handling is omitted." The shared spec is now strict decimal — optional surrounding whitespace, optional sign, one or more digits, nothing else — and invalid input FAILS on every backend, catchable by try_or: the interpreter validates before parsing (dropping OCaml int_of_string's 0x/0o/0b acceptance, which no compiled backend ever had), C gains a validating __lang_int_of_str over strtoll, Wasm's WAT helper validates and calls $__lang_fail with an interned message, and LLVM gains an IR helper over strtoll + endptr that calls __lang_fail_impl. The probe's other measurements — the Result-chain nesting tax and two phantom-type wrinkles — are design notes, not code changes. suite: 2214 passed / 0 failed (6 new tests)._
v0.1.59 — 2026-07-18
_Streaming file input, forced by a grep. The new dogfood is mgrep — a grep-lite over the backtracking regex engine (examples/regex.mere, whose probe also caught and fixed a real star-backtracking bug in the older contrib/regex engine: a*a failed on "aa" because a single returned end-position cannot give characters back to the rest of a sequence). The measurements came in three acts. Act one: grepping a 94 MB file with whole-file read_file + str_split peaked at 1.26 GB of RSS — thirteen times the file — which forced the new capability: file_open / file_read_line / file_close, an open read handle streaming one line at a time, with EOF as option None rather than read_line's ambiguous "" sentinel (interp + C; Wasm/LLVM are pointed errors; File is Send but not Sync). Act two: streaming alone made it WORSE — 2.4 GB — exposing that the CPS matcher allocates its continuation closure on RSeq entry, before the first character is tested: ~20 bytes of never-freed arena per scanned byte. Act three: first-literal-byte and ^-anchor prefilters (what real greps do with memchr) collapse the churn to match candidates — the 94 MB grep now runs in 100 MB and 1.8 s, output identical to grep -rn. The residual 100 MB ≈ the file size is the cleanest number yet for the known strings-lifetime hole: even perfectly streamed input accumulates its line strings in the program-lifetime region. That, and the closure-on-entry pattern, are the next region-reclamation forcing cases. suite: 2208 passed / 0 failed (3 new tests)._
v0.1.58 — 2026-07-18
_The ray tracer reaches the browser, and an annotation census closes a question. The playground gains /playground/raytrace.html: the same ray tracer as examples/raytrace.mere, compiled to Wasm, drawing to a canvas through two new contrib/dom externs — dom_canvas_fill_style and dom_canvas_fill_rect, the frontend FFI's first pixel-output surface (no compiler change needed; the extern machinery took five-argument imports as is). A headless harness that captures the canvas calls rebuilds the exact PPM the native backends produce, and the page's status line shows the same Adler checksum — cross-backend parity, visible on a page. Separately, a census of the "annotate polymorphic params" wrinkle (T-4) measured what a bidirectional-inference fix would actually buy: of 1,484 parameter annotations across 255 examples, almost all are stylistic (plain int params, vec_get results, and float literals all infer fine unannotated); only two patterns genuinely require an annotation — a float flowing through a polymorphic binding, and a record update on a polymorphic parameter. Both have one-annotation workarounds, so the verdict is no type-system surgery: the float case already had a pointed hint (v0.1.50), and this release gives the record-update case its twin — the error now names the workaround with an example. suite: 2205 passed / 0 failed (1 new test)._
v0.1.57 — 2026-07-18
_The Wasm backend had never actually assembled a float-heavy program with functions, and a ray tracer proved it. The probe itself — three spheres, a mirror bounce, hard shadows, all vec3 math on (float, float, float) tuples — ran identically on the interpreter and C, but the Wasm build died in wat2wasm: local.set expected [i32] but got [f64]. The cause: floats on Wasm are boxed (an i32 pointer to a heap f64), and boxing needs a raw-f64 temp local. The machinery to type those temps existed (local_types, Phase 34.3) and the main-body emitter read it — but the three FUNCTION emitters (top-level, lifted, closure adapter) ignored it and blanket-declared every extra local i32. So float expressions at the top level worked, and any float temp inside a named function produced invalid WAT. All three emitters now declare typed locals. With that fixed, the ray tracer runs on all three executable backends with an identical checksum and a byte-identical PPM (interp / C / Wasm). The checksum had to be Adler-style rather than CRC-32 — 0xFFFFFFFF doesn't fit Wasm's 32-bit int (the v0.1.41 pointed error, working as designed), and a logical shift doesn't exist for it either. The boxing tax, measured: a 96×54 render allocates 26 MB from the never-freeing bump allocator (~5 KB per pixel); at 320×180 the tax exceeds the fixed 64 MB linear memory and traps — the first concrete forcing case for the deferred Wasm memory-growth work (E-2). args is also unsupported on Wasm (pointed error), so the example writes its PPM unconditionally instead of arg-gating it. examples/raytrace.mere. suite: 2204 passed / 0 failed (2 new tests)._
v0.1.56 — 2026-07-17
_Full namespacing of user value/function names in the C backend — the robust end of the reserved-name whack-a-mole. Six times a user name collided with the C namespace (index, remove, acct, dup, run, y0), each patched by adding to a hand-maintained reserved-word list or a missed sanitizer path. That list is now gone: c_safe_name prefixes every user value/function identifier with mu_, so nothing user-named can collide with a C keyword, a libc/POSIX symbol, or a libm function ever again. Two properties make the uniform prefix the real fix rather than a bigger list: additions to libm/POSIX can't reintroduce the bug, and because the prefix is uniform, any emission path that forgets to route a name through c_safe_name fails to compile for every function (not just reserved-named ones), so the test suite surfaces such bypasses immediately — that self-verifying property caught two latent def/use mismatches during this change (pattern-variable binders and lifted-call capture arguments), now fixed. Names emitted directly are unaffected: runtime and generated symbols (__lang_*, __anon_*, __lifted_*, closure_*, mere_*), FFI extern names, and the real int main. TYPE names (records and variants) are a separate C namespace and stay un-prefixed via a new c_type_name, leaving the recursive-variant machinery untouched. The self-hosting byte-identical fixpoint is unaffected — it runs on the WAT backend, which shares no naming with the C backend. suite: 2202 passed / 0 failed (~40 codegen-assertion needles updated to the mu_ forms; behavior byte-identical for programs that never shadowed a builtin). Both halves of the reserved-name problem — Mere builtins (v0.1.54) and C symbols (this release) — are now closed._
v0.1.55 — 2026-07-17
_A reserved-name parameter bug, in the one function-emission path the earlier fix had missed. A date-arithmetic probe wrote fn (y0: int) -> ..., and y0 (with y1, j0, j1, gamma) is a libm Bessel function, already on the reserved list. The interpreter ran it fine, but the C backend failed to compile: the top-level curried function declared its parameter raw as long long y0, while the body — which captures that parameter into the returned closure's environment — referenced the sanitized y0_, an undeclared identifier. v0.1.51 had fixed exactly this mismatch for format_param and the closure adapter after the gzip probe hit it with index, but the plain emit_fn path (a simple top-level curried function, not lifted) still inlined the raw parameter name. It now goes through format_param like the others, so the declaration and every reference agree. The probe itself — day-number conversions, days-between, add-days, all as (y, m, d) tuples since there is no date type — was otherwise new-bug-zero, matching a reference implementation on weekdays, intervals, leap boundaries, and a thirty-thousand-day round-trip, identically on both backends. suite: 2199 passed / 0 failed (2 new tests)._
v0.1.54 — 2026-07-17
_User definitions now shadow builtins at the call site (the recurring reserved-name pain, attacked at its other root). A Scheme-interpreter probe named a function run; on the C backend that call compiled to __lang_run(...) — the shell-exec builtin — because the builtin's direct-call App-arm matched the name before the ordinary user-call path. The interpreter had always shadowed correctly (a later let binding wins), so only C was wrong. This is the same family as the join / is_digit / is_alpha / is_space guards added case-by-case earlier: a builtin App-arm should defer to a same-named user binding. Rather than keep playing whack-a-mole, a single user_shadows helper (local / captured / lifted-inner / top-level) now guards ~30 collision-prone builtin arms (run, spawn, even, odd, abs, show, fail, exit, sqrt, sin, cos, tan, chr, ord, args, len, not, fst, snd, sleep_ms, random_int, file_*, mkdir_p, list_dir, read_line, read_key, tty_*). The guard is strictly safe: it fires only when the user actually bound that name, so programs that don't shadow a builtin are byte-for-byte unaffected. This addresses the Mere builtin half of the reserved-name problem; the C-keyword/POSIX half is still handled by the c_safe_name suffix sanitizer, and full top-level namespacing (which would subsume both) remains deferred. suite: 2197 passed / 0 failed (4 new tests)._
v0.1.53 — 2026-07-17
_Lowercase record types, and one more reserved name (found by a records-heavy ledger dogfood). Mere's convention is lowercase type names with capitalized constructors — type 'a list = Nil | Cons .... Record types followed the same convention at declaration (type addr = { ... } was accepted), but the record literal addr { ... } only parsed for capitalized names, so a lowercase one fell through to a variable followed by a block and failed with "expected ';' or '}' in block" — an error far from its cause. A registered record name of any case followed by { now parses as a record literal; nested updates like { p | home = { p.home | city = ... } } work throughout. Separately, the ledger named a function acct, which collided with POSIX acct(2) at C compile time; a batch of common short POSIX names (acct, dup, read, write, open, close, time, stat, ...) join the reserved-word sanitizer. That list is inherently incomplete — namespacing all user top-level names is the robust fix, deferred as a larger byte-stream change. examples/ledger.mere models double-entry accounting with nested record updates. suite: 2193 passed / 0 failed (3 new tests). One honest wrinkle unchanged: a record parameter that is updated must be annotated, so the update site knows its type — the same "annotate polymorphic params" rule as the numeric overload._
v0.1.52 — 2026-07-17
_Inner functions get uncurried too (the real win the gzip probe was pointing at). v0.1.27 gave curried TOP-LEVEL functions an uncurried __direct twin so a saturated N-arg call skips the closure chain; inner (nested) functions never got it, so a curried inner recursive function compiled to a chain of anonymous closures — allocating a fresh env from the never-freed region on every partial application AND every recursive step. In a hot loop that is catastrophic: a 4-arg curried inner rec fn called a million times allocated 769 MB (the same work with a single tuple arg: 1.4 MB), and gzip's huff_decode made inflating 1 MB cost 484 MB. Now curried inner-lifted functions (≥ 2 params, concrete types) also get a __direct twin, and saturated call sites — including the recursive self-call — use it. Measured: the 1M-iteration microbenchmark 769 MB → 1.46 MB (~530x); gzip inflate of 1 MB 484 MB → 34 MB (~14x), still byte-identical with a verified CRC-32. The single-param closure form stays for partial application, so the change is additive and byte-stable (self-host emission unchanged). suite: 2191 passed / 0 failed (3 new tests). This closes the memory question the C-2 gzip dogfood opened — it was inner-fn currying, not the bytes representation._
v0.1.51 — 2026-07-17
_Three C-codegen bugs a gzip inflater flushed out. Writing a real DEFLATE decompressor (stored + fixed + dynamic Huffman, ~300 lines) exercised the closure-lifting and pattern-matching machinery harder than any prior program, and each bug was an undeclared-identifier compile error the interpreter never saw:_
- _Reserved-name params. A parameter named after a C keyword
(index) was declared raw but referenced via c_safe_name as index_. Fixed on both emission paths — lifted-fn params (format_param) and anonymous-closure adapters — where deeply curried inner functions land._
- _Cross-host capture confusion. A plain local variable
pwas
dropped from its function's captures because a DIFFERENT function had an inner recursive helper also named p: the "exclude inner-lifted fn names from captures" filter used a global, last-write-wins source-name map. Now resolved per-host, so a local and an unrelated inner fn sharing a name stay distinct._
- _Container-typed match fallthrough. A
matchwhose result type
is a pointer container (Vec) emitted (Vec___heap_int){0} — an undeclared struct — for the non-exhaustive default arm, via mono_variant_name mangling. Pointer containers now zero to NULL._
_With all three fixed, the inflater compiles and runs with no workarounds: it decompresses gzip-produced files (1 byte to 1 MB, stored / fixed / dynamic) byte-identically with a verified CRC-32. suite: 2188 passed / 0 failed (5 new regression tests)._
v0.1.50 — 2026-07-17
_The classics quartet (matmul, Game of Life, Sudoku, bignum): four textbook programs aimed at four suspected soft spots — nested Vec[Vec[float]] construction, read-current/write-next generation updates, mutate-and-undo backtracking over vec_set, and digit-vector arithmetic past the fixed-width int. All four ran correctly on interp and C with zero new bugs: the matrix product is exact, the glider translates (+2,+2) in 8 generations, the 9x9 puzzle solves (row0=534678912), and 30! comes out to all 33 digits (after 21! demonstrates the wrap — identically on both backends). After 26 releases of probe-driven fixes, that's a measurement of the suite's reach, and it's recorded as one. The single real pain was an ERROR MESSAGE: when the numeric overload defaults to int through a polymorphic helper (matmul's mat_get, whose element type is still a type variable at inference time), the eventual "expected float, got int" surfaces far from its cause. The unify hint now explains the defaulting and both escapes (annotate a parameter / ascribe an operand). All four programs join examples/ (the Life one as life_glider.mere — game_of_life.mere already existed as the Phase 36 sugar showcase and stays untouched)._
v0.1.49 — 2026-07-17
_A pub/sub broker, and the bug it flushed out. The dogfood set out to force select (waiting on multiple channels at once) — and found it isn't needed: a broker that must react to publishes, subscriptions, and shutdown funnels everything through one command inbox as a cmd variant (the actor pattern), so it never waits on two channels simultaneously. The example also shows channels are first-class message payloads — a Sub command carries a subscriber's Channel[int] through the inbox. What the dogfood did force was a closure-lifting bug in the C backend: a recursive loop that calls a sibling helper whose own nested rec go closes over the helper's locals had those locals (hn, hv) leak into loop's capture set. The transitive-capture fixpoint (which threads a callee's captures through its callers) added a callee's captures without skipping names already bound inside the caller, so loop was emitted as __lifted_loop_N(bag, hn, hv, k) — referencing hn/hv that aren't in its scope ("use of undeclared identifier"). Fixed by skipping any callee capture bound anywhere inside the caller's body. examples/pubsub.mere runs a two-topic broker with two subscribers on interp and C alike (topic0=6 topic1=30). E-1's last piece, select, stays deferred — not from lack of trying, but because the actor pattern subsumes it._
v0.1.48 — 2026-07-17
_Timed receive for supervisors (the second half of the concurrency arc): v0.1.47 let a worker pool shut down cleanly, but a supervisor still had no way to give up on a stuck worker — channel_recv on the results channel blocks forever if a job hangs. `channel_recv_timeout : Channel[a] -> int -> option[a]` blocks up to N milliseconds for a value and returns None on timeout (or once the channel is closed and drained), so a collector records the timeout and moves on instead of hanging the whole run. The C backend uses pthread_cond_timedwait against an absolute CLOCK_REALTIME deadline; the reference interpreter polls at 1 ms granularity (the stdlib has no timed condition wait). interp + C; Wasm and LLVM reject it with a pointed compile error. examples/supervised_pool.mere runs a pool where one job deliberately hangs and the supervisor collects the other five with a 300 ms budget (results=5 timeouts=1) on interp and C alike. Structured-concurrency cancellation is already expressible via channel_close; the one remaining E-1 piece is select over multiple channels, still waiting for a forcing program (a genuine multi-source wait)._
v0.1.47 — 2026-07-17
_Graceful shutdown for concurrency (found by a worker pool): the pool — main pushes N jobs, W workers pull and process, main collects — hit two walls at once. A worker's channel_recv loop blocks forever when the jobs run out, so there was no way to stop a worker and join it; and because the loop never returns, its type is bottom ('a), which the C backend can't emit ("unsupported C codegen type: 'a"). Both are the same missing primitive. `channel_close : Channel[a] -> unit` marks a channel done, and `channel_recv_opt : Channel[a] -> option[a]` blocks for a value but returns None once the channel is closed and drained — so a worker loops match channel_recv_opt jobs with None -> ()
Some j -> ...; loop (), which terminates (returns unit, no longer |
|---|
bottom) and can be joined. channel_recv and channel_send on a closed channel now raise/abort instead of blocking or corrupting. interp + C (the native worker-pool / server target); Wasm and LLVM reject the two with a pointed compile error. examples/worker_pool.mere runs a 4-worker pool over 12 jobs and joins every worker cleanly. Closes the structured-concurrency gap (E-1) that had been waiting for a forcing program since the memory model landed._
v0.1.46 — 2026-07-16
_(Follow-up, no version bump) examples/base64.mere: a composition probe that confirms the day's three separately-shipped capabilities — the bitwise builtins, read_file_bytes, and write_file_bytes — compose in one program. RFC 4648 known-answer vectors pass, and passing a file path round-trips arbitrary binary byte-identically (read_file_bytes → encode → decode → write_file_bytes) on interp and C. No new bug surfaced — the value is the integration check itself._
_Hex literals (a papercut two probes drove into the ground): both the SHA-256 round constants and the East Asian Width range table had to be written in decimal, because 0xFF lexed as the int 0 followed by an identifier xFF ("unbound variable: xFF"). 0xFF / 0Xff now lex as ordinary ints — no separate type, same per-backend width — via int_of_string; a bare 0x with no hex digit still reads as 0 then the identifier x. No octal / binary / digit-separator syntax (not yet forced). The lexer change is one branch; the value is that the next crypto or Unicode probe reads like the reference it's transcribed from._
v0.1.45 — 2026-07-16
_Columns, not codepoints (found by printing a table with Japanese cells): v0.1.38's codepoint view was the right first step and the wrong tool for alignment — utf8_len says 5 for こんにちは, a terminal draws it in 10 columns, and a product table with CJK rows comes out visibly ragged. `utf8_width` is the display width (East Asian Width, wcwidth-lite: CJK / fullwidth / emoji = 2 columns, combining marks = 0, halfwidth katakana = 1), and `pad_right` / `pad_left` pad on it. All three are prelude functions in pure Mere — UTF-8 decoded with plain div/mod arithmetic, the width table a dozen range checks in decimal (the lexer has no hex literals, which is now a recorded papercut) — so they landed on all four backends at once by construction. examples/aligned_table.mere renders a mixed ASCII / Japanese / emoji / halfwidth-katakana table with straight borders on interp and C alike._
v0.1.44 — 2026-07-16
_The picture that fixed the docs (found by a Mandelbrot renderer): the probe went in expecting to measure the "no float infix" tax the docs promised — and the docs were wrong in the language's favor. `+ - * /` and the comparisons had been numeric-overloaded for a while on interp, C, and Wasm; the reference still said prefix-only f_add style, and a docs-faithful reader would write nine needless prefix calls per formula. Three real gaps did surface around the stale entry, all fixed: unary minus was int-only (-2.5 was a type error; negative float literals needed f_neg) — now overloaded like the binary operators on all four backends (fneg / f64.neg); the LLVM backend emitted `add i32` on double operands for float infix (invalid IR, and icmp for float comparisons) — now the fadd family and ordered fcmp; and the write half of the binary path was missing — write_file_bytes : str -> Vec[R, int] -> unit joins v0.1.43's reader, so PPM's raw P6 replaces the 2.6x-larger P3 ASCII escape. examples/mandelbrot.mere renders 400x300 in infix math and writes P6 that is pixel-identical to the P3 version. One honest wrinkle stays: the numeric overload resolves to float only on concretely-float operands, so unannotated fn params default to int — float-heavy code annotates its params. Docs corrected in both places._
v0.1.43 — 2026-07-16
_Bytes get in the door (found by a 30-line CRC-32 tool): the algorithm was trivial on the new bitwise builtins — the discovery was on the input side. `read_file` silently truncates binary data at the first 0x00 byte on the C backend (NUL-terminated char*), while the interpreter, whose strings carry NULs, read the same file correctly: a 25-byte file read as 2 bytes natively and produced a confidently wrong checksum. The str-is-bytes story was only true on interp. `read_file_bytes : str -> Vec[R, int]` is the binary-safe path — one int per byte, 0..255, the whole file, reusing the existing vec machinery instead of introducing a bytes type (8 bytes per byte is the honest cost until a program forces better). It gets the same construction-time region binding as vec_new (without it, the region tyvar stayed unresolved and functions taking the vec were silently never emitted by the C backend — the probe hit that too). interp + C for now; Wasm/LLVM reject it with a pointed compile error. examples/crc32.mere verifies against zlib on both text and NUL-bearing files; read_file's docs now state the truncation divergence plainly._
v0.1.42 — 2026-07-16
_The real ALU (paying off the SHA-256 probe): bitwise builtins on all four backends — bit_and / bit_or / bit_xor / bit_not / bit_shl / bit_shr, on the backend's native int width, with bit_shr as the arithmetic shift. They lower to the machine operation everywhere: &-family operators on C, i32.and-family instructions on Wasm, and i32-family on LLVM, land-family on the interpreter. examples/sha256.mere dropped its div/mod fake ALU for them: one block went from ~29 ms (interpreted, bit-loop emulation) to 6.7 µs native — about 4,300× — with all NIST vectors still passing on interp and C. Cleanups the rewrite surfaced: abs/min/max/clamp still used C int temporaries after v0.1.41 (silent truncation above 2^31, fixed); str_of_int on a variable under a top-level let referenced an undefined show_int on Wasm and LLVM (only show registered the helper, fixed on both); the LLVM backend now rejects out-of-range int literals at compile time like Wasm does — and the v0.1.41 changelog's claim that LLVM was i64 is corrected there: LLVM's int is i32, and widening it to 64-bit remains a deferred item with sha256 as the forcing program._
v0.1.41 — 2026-07-16
_One int, not four (found by writing SHA-256 in pure Mere): the probe aimed at the missing bitwise story and instead hit something under it — the C backend's int was C `int`, 32 bits, while the interpreter tested 63-bit semantics and the docs never said which. SHA-256's round constants (36 of them above 2^31) silently truncated and every digest came out wrong with zero diagnostics; the minimal repro is 2147483647 + 1, which printed -2147483648 natively and 2147483648 under the interpreter. The C backend's int is 64-bit (`long long`) now, with LL-suffixed literals so literal arithmetic doesn't wrap at 32 bits either, %lld show/json formats, and atoll/strtoll parsing. At the extern fn FFI boundary int deliberately stays C int — the functions users declare are libc/POSIX symbols whose ABI type IS the 32-bit int (declaring getpid as returning long long would read undefined upper register bits on arm64). The Wasm backend keeps its i32 int but now says so: an int literal outside -2^31 .. 2^31-1 is a compile-time error with a source location instead of an i32.const 4294967296 that only explodes later inside wat2wasm. The SHA-256 probe passes all NIST test vectors on interp and C; docs state each backend's width honestly. (This entry originally claimed LLVM was already i64 — measuring said otherwise: LLVM's int is i32, so it now gets the same out-of-range-literal compile error as Wasm, and the i64 widening is a known deferred item. The probe also uncovered an unrelated LLVM crash on this program, tracked separately.)_
v0.1.40 — 2026-07-16
_Error-handling ergonomics probe (an 8-step fallible config-loader written three ways): the verdict on the language was mostly good news — the ? / ?! early-return sugar from Phase 36 already turns a seven-level match pyramid into a flat sequence of bindings, and the prelude's result_and_then family covers combinator style. The probe found one genuine inconsistency: the `?` / `?!` lets were the only let form that rejected `;` as sugar for `in` — let x = e?!; rest was a parse error while every other let x = e; rest works. Fixed; both forms now accept both separators._
v0.1.39 — 2026-07-16
_Scale safety (found by sorting a million elements): `list_sort_by` is a stable merge sort now, and the prelude's list functions survive million-element lists. The insertion sort took ~2 s at 20k elements natively and O(n²) beyond — a million-element list_sort now runs in well under a second, still stable (ties keep input order; the merge is tail-recursive via a reversed accumulator, and the split avoids returning a tuple: a struct return compiles to an sret out-parameter in C, which quietly defeats clang's sibling-call optimization — that one cost an AddressSanitizer session to find). Ten more prelude functions were rewritten with accumulators after the probe showed the naive Cons (f h, recurse) shape overflowing the stack near a million elements: list_len, list_map, list_filter-adjacent take/zip, list_append, list_concat, list_flat_map, range, list_max, list_min. The derive family (== on a million-element list) was already safe. list_sort_insert remains for direct users._
v0.1.38 — 2026-07-16
_Unicode (found by ten minutes of typing Japanese at the language): the codepoint view of strings. A Mere str is — and stays — a byte string: str_len "こんにちは" is 15, substring can cut a character in half, and str_rev scrambles multibyte text; all documented rather than changed (byte indexing is what the FFI, the wire protocols, and the existing corpus rely on). What was missing was any way to work with text: two new builtins on all four backends — utf8_len : str -> int (codepoint count) and utf8_chars : str -> str list (split into codepoints; invalid bytes count as single units, so they never loop or throw) — plus prelude compositions utf8_at, utf8_sub, and utf8_rev, written in plain Mere on top of utf8_chars so every backend gets them for free. utf8_rev "aあ😀b" is "b😀あa" on interp, C, Wasm, and LLVM alike — the first new builtin family to land on all four backends at once (str_split's runtime scaffolding made LLVM cheap)._
v0.1.37 — 2026-07-15
_Memory model, ported to Wasm: `region R { }` reclaims on the Wasm backend — the sound version of the save/restore that Phase 16.4 removed as broken. Three parts make it sound where the old attempt was not: the block's result is deep-copied out (per-type $__mcopy_<tag> fns, twice — once above the block's garbage, then down into the enclosing range after the bump restores; the ranges cannot overlap); escaping stores are compile errors (pushing a heap value into a container created outside the block, map_set, strbuf_push on an outer buffer, channel_send, spawn, and externs that register callbacks — a container created inside the block is free to mutate, it dies with the block); and escaping closures/containers/borrows are rejected via the result type. Wasm needs no thread-locals or heap blocks: a mark saved on the value stack and one scratch global do it._
_Measured on the live 2048 with a per-move region around the key handler: the bump pointer stays at exactly 4,544 bytes across 30,000 moves — zero net allocation per move, zero traps. The same game previously burned ~8.4 KB per move and died at ~7,700. The remaining honest gap vs the C backend: no per-container storage (hence the escaping-store errors instead of C's copy-on-store), recorded in memory-model.md §3.5._
v0.1.36 — 2026-07-15
_Library hygiene, applied across contrib: importable libraries are main-free now. Five more libraries carried a demo main at the bottom of the file (the pattern v0.1.35 fixed for contrib/test), so importing them ran the demo — argparse, csv/writer, regex, regex/engine, and time. Each demo moved to examples/<name>_demo.mere and runs standalone. The self-host family (parser / typer / fmt / eval / codegen_wasm) keeps its inline demos deliberately: those are programs whose demo output is the cross-implementation test vector, not libraries._
v0.1.35 — 2026-07-15
_Test-framework dogfood (three small things it surfaced):_
_Generic assertions confirmed working. show (like ==, and like < since v0.1.33) works through type variables — monomorphization plays the dictionary — so contrib/test's assert_eq is genuinely generic: a helper fn s -> fn name -> fn x -> Test.assert_eq s name x x asserts on ints, tuples, nested pairs, and prints failing values with no annotations. No language change was needed; the regression test pins it._
_Library files must not carry a demo main. contrib/test's demo lived at the bottom of the library file, so every importer ran it (noise, an intentional FAIL, and the demo's exit status). The demo moved to examples/test_framework_demo.mere; the library is module-only now, like contrib/xml._
_`-I` now works when running a file. The import search path flag was honored by -c / -l / -w but silently dropped by the interpreter path (mere -I <dir> file.mere failed to resolve imports that mere -c -I <dir> accepted) — the run entry points now pass the search paths through, closing another CLI asymmetry (cousin of v0.1.29's)._
v0.1.34 — 2026-07-15
_Soundness (found by playing the live 2048 for ten thousand headless moves): `&&` and `||` now short-circuit on every backend. The interpreter and the C backend always short-circuited, but the Wasm backend emitted strict i32.and / i32.or and LLVM emitted eager and i1 (behind a comment claiming the "MVP subset has no effects" — long obsolete: a trapping right-hand side IS an effect). The bounds-guard idiom i < len && vec_get v i == x therefore trapped on Wasm only — in production, 97% of the live 2048's keypresses died silently in its stuck-detection (r < 3 && bget b (i + 4) == v), invisible because the DOM glue catches and logs closure exceptions. Both backends now lower &&/|| to their If emission._
_The same probe measured the Wasm page-lifetime allocation model (the memory-model work of v0.1.30–31 is C-only so far): the game burns ~8.4 KB of never-reclaimed bump per move and hits its 64 MB memory at move ~7,700 — a determined player kills the tab in under an hour. That number is now the forcing measurement for porting value reclamation to the Wasm backend._
v0.1.33 — 2026-07-15
_Polymorphic ordering: `<` / `<=` / `>` / `>=` now work through type variables, closing the gap derive-ord (v0.1.11) left open. The design is deliberately not a trait system: the scheme carries no constraint — instead monomorphization plays the dictionary's role. Every compiled instance of a polymorphic comparator compares at a concrete type, where the existing derive machinery (cmp_<tag>) specializes; the interpreter compares structurally at runtime. This is exactly how == has worked through type variables all along — ordering simply joins it (the historical "unresolved comparand defaults to int" rule is gone; programs that used the default still typecheck, since instantiation covers them)._
_Consequences for free: the prelude's list_sort, list_max, and list_min are now generic — list_sort [(3, "c"), (1, "a")] sorts tuples with no annotations and no comparator; a hand-written fn a -> fn b -> a < b instantiates at every use type (the generic pairing-heap example drops its annotated comparator). Instances are structural only — there is no way to override a type's ordering (the derive family's philosophy), _by variants remain for explicit control, and the parity scope is interp / C / Wasm, as with derive-ord._
v0.1.32 — 2026-07-15
_Cleanup release (three small fixes plus doc sync):_
_Top-level / local name collision (invalid C). A local let m = ... inside any function that shared its name with a globalized top-level let m was emitted as an assignment to the file-scope global instead of declaring a shadowing local — the prelude's list_max (local m) plus a program-level let m = map_new () produced C that didn't compile. The global-assignment form now fires only for the exact top-level spine bindings (matched by physical node identity), so same-named locals declare and shadow correctly._
_Tuple exhaustiveness false positive. match (h1, h2) with (HE, _) | (_, HE) | (HN _, HN _) is exhaustive, but no single arm is total, so the checker warned "no wildcard arm for tuple" (found by the generic pairing heap's merge). Tuple scrutinees whose components all range over small finite spaces (bools / unit / registered variants) are now checked by enumerating the product; a genuinely missing combination is reported by example — missing (Greenq, Greenq) — instead of a generic complaint._
_mem_to_str leak. It malloc'd and never freed; it now allocates in the thread's current region, so per-request region blocks reclaim byte-dialect strings too._
_Also: memory-model.md gains §3.5 documenting the implemented v0.1.30-31 reclamation semantics (current region, copy-out, copy-on-store, per-message channel copies, backend notes)._
v0.1.31 — 2026-07-15
_Memory model (stage 2 — the payoff): `region R { }` now reclaims the values its body allocates. Value allocations (strings, cons cells, variant nodes) target a thread-local current region instead of hardcoding the never-freed default region; a region block makes itself current for its body, deep-copies its result out into the enclosing region (stage 1's __mcopy machinery), and releases. Closure envs and container structs deliberately stay in the default region (they carry identity), stores into containers are safe by stage 1's copy-on-store, channel_send deep-copies the payload into a per-message region (freed on recv after copying out into the receiver's current region — a sender's scratch can die while the message is in flight), a container cannot escape as a block result (the typer's region-escape check fires; a codegen guard backs it up), and try_or restores the current region when a fail longjmps past a block. Block regions are heap-acquired with a one-deep per-thread cache, so a per-iteration block costs a pointer swap and a bump reset — and, critically, no stack struct's address escapes, which is what lets clang keep tail-calling. The spawn trampoline frees a finished thread's cached region (_Thread_local has no destructor — a spawn-per-connection server leaked ~1 MB per closed connection without this) (show/to_json/float-formatting helpers are noinline for the same reason: their inlined asprintf(&local) silently broke sibling-call optimization and deep loops overflowed the stack)._
_Measured: the idiomatic line-at-a-time counter — plain read_line + str_len in a per-line region — now runs at 1.5 MB constant RSS over 8M lines (246 MB before; wc -l needs 2.5 MB). A 100k-iteration loop storing every 10,000th string into an outer map keeps exactly the stored data. Long-running servers can finally reclaim per-request memory in the string dialect, not just the byte dialect. Suite: 2093._
v0.1.30 — 2026-07-15
_Memory model (stage 1 of the per-request-reclamation plan): copy-on-store — containers own their contents. map_set deep-copies the key and value into the map's own region, and vec_push / vec_set copy the element, via per-type __mcopy_<tag> functions specialized the same way the derive family (show / json / == / cmp) is: strings copy their bytes, tuples / records / variants copy structurally (cons cells and variant nodes re-allocate in the container's region), scalars and closures pass through, and nested containers copy as pointers (mutable identity and aliasing preserved — they own their own storage). Strings are immutable, so the copies are semantically unobservable; the point is lifetime: a stored value must not dangle when the storer's allocation scope is later reclaimed. This is the prerequisite for scoped string allocation (region R { } capturing str/cons allocations — the next stage), which is what finally makes long-running servers' per-request memory reclaimable. OwnedVec / StrBuf / Channel are deferred to that stage. Today's cost: one copy per store; today's benefit: none visible — by design._
v0.1.29 — 2026-07-15
_Soundness (mkv dogfood P2): sharing a mutable container across threads is now a compile error, and the compile path runs the same safety analyses as the run path. Two fixes:_
_Send/Sync classification. Region-bound mutable containers (Map / Vec / StrBuf) are now explicitly !Send && !Sync — their runtimes are lock-free (linear-scan arrays / bump buffers), so a shared container across spawn is a data race. Previously the classifier fell through to "are all type args Send?", and the region-marker arg is a bare TyVar, judged optimistically — so a shared Map compiled fine and lost ~2% of concurrent writes in a real RESP-server stress test. OwnedVec stays Send/!Sync (drop type: single owner, movable). The blessed pattern is share-by-communicating: Channel remains Send+Sync, and the mkv actor model compiles unchanged._
_The `-c` / `-l` / `-w` paths now run the safety analyses. The compile entry ran type inference only — channel-element Send obligations, borrow-conflict checking, and spawn-capture move analysis were silently skipped, so mere file.mere rejected programs that mere -c file.mere happily compiled (including capturing a region borrow in a spawned thread). All three checks now run before codegen on every backend._
v0.1.28 — 2026-07-15
_Fix (generic-PQ dogfood, two monomorphization bugs): a generic pairing heap (type 'a heap = HEmpty | HNode of ('a * 'a heap list) + comparator closures) ran correctly on the interpreter but failed to compile natively. Two independent root causes, both in the C backend's monomorphization:_
_B-P2 — body-only tuple shapes were never collected. Tuple typedef collection walked main's AST and fn signatures, but not fn bodies — so a tuple that exists only as a body annotation (the (h1, h2) scrutinee of a poly fn's match, concrete only inside a monomorphized instance's cloned body) was referenced in the emitted C without ever being declared. Bodies are now walked too; the concreteness guard still skips unresolved polymorphic shapes._
_B-P2b — no promotion to multi-instance. A poly fn's usage sites inside another poly fn's body only become scannable once that fn resolves. hp_pop was seen at one type (from main), single-resolved by unifying the original skeleton in place — destroying its polymorphism — and the later-discovered second usage (at int, inside drain) was emitted against the wrong instance's struct types. Every skeleton now keeps a pristine clone taken before any unification; single-resolved fns' bodies join the arrow-discovery scan; and a fn already resolved at one type is promoted to multi-instance when a second type shows up._
_With both fixed, the generic heap and a Dijkstra built on it (new examples/generic_heap_dijkstra.mere) run natively, byte-identical to the interpreter. Suite: 2081._
v0.1.27 — 2026-07-14
_Optimization (mlog dogfood P4, the big one): saturated calls to curried top-level fns compile to a direct N-ary C call. Level-by-level application allocated a closure env in the default region per call, through the region lock — measured as O(iterations) permanent memory in every multi-argument hot loop: a byte-at-a-time line counter held 2.1 GB RSS over 8M lines. For each top-level f = fn p1 -> .. -> fn pN -> body (N ≥ 2, concrete types) the backend now also emits f__direct(p1, .., pN) and compiles exactly-saturated call sites straight to it — argument temporaries pin the interpreter's left-to-right evaluation order, and self-recursion becomes a C self tail call. Partial applications and first-class uses keep the curried chain. The same line counter is now 1.5 MB RSS, constant across input size (below wc -l), and 300 MB of input streams in 0.13 s. Constant-memory streaming is genuinely expressible now; what still accumulates is the string dialect's per-line str values (the open type-level lifetime question)._
v0.1.26 — 2026-07-14
_Capability (mlog dogfood P1): `read_line` on the C backend. It was interpreter-only — the sixth member of that family (print_err / file_exists / print_no_nl / random_int / file_size) — so a native streaming line processor could not be written at all (read_stdin slurps the whole input by design). __lang_read_line reads one stdin line without the trailing newline, "" on EOF, matching the interpreter. Found by measuring memory behaviour of line-at-a-time processing for the constant-memory streaming question._
v0.1.25 — 2026-07-14
_Fix (mkv dogfood, long-running processes): regions grow instead of aborting. The region allocator was a single fixed-cap bump block (default region: 4 MB) that aborted with region OOM on overflow — a long-running server's per-command allocations (reply strings, cons cells, tuples) exhausted it after a few thousand requests. A region is now a chain of bump blocks: on overflow a geometrically larger block is chained on. Blocks never move, so existing pointers stay valid, and region R { } frees the whole chain at scope exit. Also hardened the native byte arena: mem_alloc / str_ptr share one bump pointer across spawned threads — it is now mutex-guarded and bounds-checked (it previously raced and silently overflowed past the arena). Under a sustained 80k-command concurrent load the RESP server now runs clean where it previously aborted at ~8k. The honest remaining edge: growth is not reclamation — per-request memory still accumulates for the process lifetime (region-scoped strings need type-level lifetime tracking; see the memory-model open questions)._
v0.1.24 — 2026-07-14
_Capability (mkv dogfood, T4 wire-protocol server): native TCP server primitives. tcp_listen : int -> int (socket + SO_REUSEADDR + bind + listen, returns the listening fd) and tcp_accept : int -> int (blocking accept, returns the client fd) join the existing native_ffi_names, emitted as static impls against the same flat arena + POSIX sockets that back tcp_connect/tcp_read/tcp_write. A Mere program can now be a TCP server, not just a client — the server-side mirror of the pg/redis client FFI. SIGPIPE is ignored so a client disconnecting mid-write drops the connection rather than the whole process. This is the missing capability behind a Redis-wire (RESP) key-value server; the earlier http_serve was HTTP-specific and single-connection._
v0.1.23 — 2026-07-14
_Fix (docs site): the Mere SSG (contrib/site/build.mere) parsed its CLI args assuming args() still prepended the script path — the v0.1.12 args() consistency fix shifted that by one, so input_dir resolved to the output dir and the site built 0 markdown pages (tour.html / tutorial.html etc. 404'd). Updated build.mere to the current args() contract (first positional = input dir). A dogfood consumer that relied on the old behaviour — exactly the interp/native args() mismatch N3 was about, biting a Mere program this time._
Fix: same-named inner functions no longer collide when lifted (2048 dogfood P3). Two inner fns sharing a source name within one top-level function — e.g. a let rec go in each branch of an if — both lifted to the top level, but each backend's inner-fn resolution map is keyed by the source name, so the second go overwrote the first and both call sites dispatched to the wrong one. Cross-backend: the C and Wasm backends both mis-executed (silent wrong results); the interpreter was correct. A new shared pre-pass (Ast.uniquify_inner_fns_program, run next to the par_map lowering) α-renames on collision — the first use of a name keeps it, a later reuse becomes <name>_uq<N> with its references rewritten — fixing every backend in one place. Collision-free inner names (the common case) are untouched, so nothing changes in ordinary code or its pretty-printing.
2069 tests.
v0.1.22 — 2026-07-14
Wasm backend: `spawn` / `join` / `channel_*` now respect shadowing (2048 dogfood P2). A user binding named spawn — a game's tile spawner — was dispatched to the concurrency builtin, silently turning the module into a threaded one (shared-memory import + $mere_spawn), which the plain browser host rejects. The same bug family the C backend fixed for join in the mk dogfood (dd17b8a): the dispatch matched the name without asking whether it was rebound. All five concurrency dispatches now check the local scope / top-level fns / inner-lifted fns first, so a shadowed name falls through to ordinary application while genuine spawn still lowers to $mere_spawn.
Also in the frontend FFI (no compiler change): contrib/dom gained dom_on_key : (str -> unit) -> unit — a global keydown listener passing the key name to a Mere closure; the browser counterpart to native read_key.
2067 tests.
v0.1.21 — 2026-07-14
`file_size` — a binary file's true byte length (mwasm dogfood P1). read_file is binary-safe (the buffer holds every byte and char_at / ord index past NULs correctly, on interp and C native), but str_len is strlen on the C backend and stops at the leading NUL — so a .wasm (magic \0asm) reported length 0, and a binary walk couldn't bound its loop. Added file_size : str -> int (stat's st_size, next to file_mtime), on interp and C. With (buffer, size) carried explicitly, the NUL-safe char_at / ord / substring make binary parsing expressible — no dedicated bytes type needed yet. Driving app: mwasm, a WASM binary inspector that reads the compiler's own output.
2065 tests.
v0.1.20 — 2026-07-14
`random_int` now works on the C backend (mrog dogfood P3). The game's wandering ghost picks a random direction each turn; random_int existed only in the interpreter — the third interpreter-only builtin this dogfood family has flushed out (after print_err, file_exists, print_no_nl). Added __lang_random_int (seeded once from time^pid, uniform [0, n), fails on n <= 0 like the interpreter). mrog M3 — ghost + game over — now runs natively.
2064 tests.
v0.1.19 — 2026-07-13
`print_no_nl` now works on the C backend (mrog dogfood P2). A TUI's cursor-control sequences must be written without a newline and without line buffering; print_no_nl existed only in the interpreter (the same family as print_err / file_exists before it). Added the case (fputs(s, stdout); fflush(stdout)). With it, mrog's full redraw loop — ANSI clear+home, map with @ overlay, hjkl movement, wall collision, gold pickup — runs natively, byte-identical to the interpreter.
2063 tests.
v0.1.18 — 2026-07-13
Interactive terminal: `tty_raw` / `tty_restore` / `read_key` (mrog dogfood P1). Mere had only line-buffered input (read_line waits for Enter, with echo), so an interactive TUI couldn't be expressed at all. Three new builtins — interpreter (Unix termios) and C native (tcgetattr/tcsetattr):
tty_raw : unit -> unit— raw mode on stdin (no echo, no canonical
buffering; ISIG stays on so Ctrl-C works). No-op when stdin isn't a tty, so piped tests behave.
tty_restore : unit -> unit— put back the termios saved by the first
tty_raw.
read_key : unit -> str— blocking single-byte read;""on EOF.
ANSI output already worked (chr 27 ++ "[2J"), so with key input the interactive read → update → redraw loop is now expressible. Driving app: mrog, a tiny terminal roguelike.
2062 tests.
v0.1.17 — 2026-07-13
C backend: closures that call an inner-lifted fn now carry its captures (mk dogfood P5). An inline lambda passed to par_map that captures an enclosing function's parameter gets inner-lifted, and its call sites inject the captured variable as a leading argument. But when that call site sat inside another closure — the par_map lowering's spawn lambda — the spawn closure's env didn't include the injected variable, and the emitted C referenced an undeclared identifier. The anonymous-closure capture computation now unions in the captures of any inner-lifted fn the body calls (one level suffices — lifted captures are already transitively closed by the Phase 45 fixpoint). Found by mk's parallel dependency groups (name [a b c]&: cmd), which now build and run natively: three parallel 0.3s deps complete in ~0.38s, and a failing parallel dep propagates its exit code.
2058 tests.
v0.1.16 — 2026-07-13
`run` is now truly parallel under `spawn` / `par_map` (mk dogfood P4). run was lowered to libc system() (and OCaml's Sys.command, which wraps it) — and on macOS, concurrent system() calls serialize behind a global lock, so par_map (fn c -> run c) cmds executed commands one at a time: three parallel 0.3s sleeps took ~1.0s (interp) / ~1.6s (native). Confirmed with a C probe (3 threads × system("sleep 0.3") = 1.01s; posix_spawn = 0.32s). Reimplemented without system():
- interp:
Unix.create_process "/bin/sh" ["sh";"-c";cmd]+waitpid - C native:
posix_spawn+waitpid(128 + signalon signaled exit)
Three parallel 0.3s commands now take ~0.36s on both backends. Exit-code propagation is unchanged. This is what a parallel task runner needs — the mk dogfood's M5.
2057 tests.
v0.1.15 — 2026-07-13
`file_exists` now works on the C backend (mk dogfood P3). Incremental builds skip a task when its output exists and is newer than its inputs; the "exists" check guards file_mtime (which raises on a missing path). file_mtime was already on C, but file_exists was interpreter-only, so the native build failed with use of undeclared identifier 'file_exists'. Added the case (stat(path, &st) == 0, next to __lang_file_mtime). With this, the mk task runner's incremental mode (name (out: in1 in2): cmd) builds and runs natively — and its float mtime comparison rides the v0.1.11 structural >.
2057 tests.
v0.1.14 — 2026-07-13
`print_err` now works on the C backend (mk dogfood P2). The native backend lowered print to puts but had no print_err, so a compiled CLI couldn't write diagnostics to stderr — a native build using it failed with use of undeclared identifier 'print_err'. Added the case (fprintf(stderr, "%s\n", …), mirroring print → puts); the docs' 3-backend claim for print_err is now actually true.
2056 tests.
v0.1.13 — 2026-07-13
`run` — Mere can start external programs. A new run : str -> int builtin executes a command line through the shell, inherits stdio, and returns the exit code (interpreter via Sys.command; C native via system + WEXITSTATUS). This is the capability the new mk task-runner dogfood needed on day one — a whole class of tools (build systems, task runners, anything that shells out) was previously inexpressible. Exit codes propagate identically under interp and native.
2054 tests.
v0.1.12 — 2026-07-13
Papercut batch — small dogfood findings paid back.
- `args()` is now consistent between the interpreter and native binaries
(mstat N3). Both return only the program's own arguments, dropping the interpreter's script path / the binary name; the CLI entry point hands the post-script args to the args() builtin instead of it reading Sys.argv[1..]. An argument-driven CLI now behaves the same under mere app.mere a b c and the compiled ./app a b c.
- `str_of_float` renders whole-valued floats as `550.0`, not `550.`
(mstat N4). Fixed identically across interp / C / Wasm (and the show path), so all backends still agree and the output round-trips through float_of_str.
Deferred: bare None needing a type annotation is an inference matter, not a papercut, and stays open.
2052 tests.
v0.1.11 — 2026-07-13
derive-ord: structural ordering, the sibling of structural equality. < <= > >= now work on any concrete type, not just int / float / str — completing the compile-time-specialized "derive family" (show / to_json / of_json / == / `<`).
- Structural comparison on tuples, records, lists, and variants, on
interp / C / Wasm, all agreeing byte-for-byte. Lexicographic: tuples and records by declared field order, lists element-wise (shorter prefix is smaller), variants by declaration order then payload. Emitted as a cmp_<tag> function per type (the ordering sibling of eq_<tag>), and as value_compare in the interpreter, ordering variants by the same tag order the codegen assigns.
list_sort_bywith an annotated comparator now sorts a list of any
structural type (float / record / tuple / …), closing the mstat N5 finding's practical half.
- Backward compatible: an unresolved comparator type variable still
defaults to int, so fn a -> fn b -> a < b and the bare list_sort stay int. A fully-polymorphic list_sort needs ad-hoc-polymorphism resolution and remains deferred (documented in the stdlib reference).
2052 tests.
v0.1.10 — 2026-07-12
Bootstrap fixpoint: Mere is truly self-hosting. The Mere-in-Mere compiler, compiled by itself and run as wasm, produces byte-identical output to the reference — and that output runs correctly.
- Self-host TCO (Stage 55f): the self-host codegen now emits
return_call_indirect (guaranteed tail calls) for tail-position closure calls, tracked via a tail flag threaded through if / let / letrec / match. Deep tail recursion in self-compiled code stays stack-flat (a 200000-deep counter completes; it overflowed before).
- Three latent self-compilation bugs fixed (Stage 55g) — found by
trace-bisecting the self-compiled compiler until the bootstrap fixpoint held:
- Pattern checks: a
PConstrpayload sub-check ran eagerly even when
the tag didn't match, dereferencing garbage (out-of-bounds traps). Payload checks now short-circuit.
- Var-vs-var string
==lowers to pointer equality in the un-typed
self-host codegen; member_str (and parser friends) switched to explicit str_eq — ghost closure captures are gone.
- The self-host lexer was missing the
\rescape, corrupting the
data-segment escaper's CR needle ("Err" emitted as "E\0d\0d").
- Fixpoint regression test: the suite now compiles a program with the
interpreter-run compiler AND the self-compiled compiler and asserts the WAT outputs are byte-identical.
- Also:
let recwritten directly in the main expression now lifts on
C + Wasm (mstat N6) instead of erroring.
2035 tests.
v0.1.9 — 2026-07-12
Float operator overloading + libm name collisions — driven by the mstat numeric-CLI dogfood.
- Infix operators on float:
+ - * /and< <= > >=now work on
float, not just int / str. Dispatched on the operand type at codegen (the same compile-time specialization as show / to_json / eq; no trait machinery). Mod stays int-only. All four backends' arithmetic/ordering covered. Also fixes a latent C bug where a whole-valued float literal emitted as 7 (via %.17g), making 7.0 / 2.0 integer division. Caveat: operands must be concretely float-typed — an unannotated fn a -> fn b -> a < b still defaults to int, so the default list_sort stays int (sort floats with an annotated comparator).
- libm / POSIX name collisions: a user fn named
fmin/fmax/ … now
gets rehomed (fmin_) instead of clashing with <math.h> in the C backend (conflicting types for 'fmin'). Same treatment as main.
2033 tests.
v0.1.8 — 2026-07-12
of_json / of_json_opt on the Wasm backend — backend parity.
- Wasm `of_json` / `of_json_opt`: ported the JSON deserializers to the
Wasm backend, so all three shipping backends (interp / C / Wasm) have them — matching to_json's coverage (LLVM excluded, it lacks to_json too). A WAT JSON-parser runtime builds a generic tree in linear memory; per-type $__ojnode_<tag> decoders build the typed value; strict of_json traps on error, of_json_opt returns None. This un-blocks the mere-blog dogfood's wasm deploy path (native-only since it adopted of_json_opt in v0.1.7).
2022 tests.
v0.1.7 — 2026-07-11
of_json (derive-style JSON parsing) + docs push + ergonomics.
- `of_json` / `of_json_opt`: the deserialization mirror of
to_json.
of_json : str -> 'a parses JSON into a typed value, driven by the result type at the call site (an annotation (of_json s : T)) — JSON object → record fields by name, array → list / tuple, null/value → option, string / {"Ctor":…} → variant. Same compile-time specialization as show / to_json; interp + C (native) backends. of_json_opt : str -> 'a option is the non-crashing sibling (returns None on any parse / shape error) — safe for untrusted input like HTTP request bodies. Closed the mere-blog dogfood's request-parsing gap (PAIN B5): its handlers now decode into typed request records instead of plucking string fields, verified end-to-end on the native binary.
- `option` is a transparent JSON nullable:
to_jsonnow encodes
None as null and Some x as x (was the tagged {"Some":x}) on all three backends, the idiomatic API encoding and symmetric with of_json.
- Native `exit n`: the C backend emits libc
exit(), so a native CLI
can set its process exit code (closed mq PAIN P1's last item).
- Trailing commas: allowed in list and tuple literals (
[1, 2, 3,],
(a, b,)); records already allowed them.
- Docs: a one-page Tour of Mere feature showcase, and the
SSG's nav / index are now curated (Start here → tutorials → reference) with real page titles. Site live at merelang.org.
2019 tests.
v0.1.6 — 2026-07-11
to_json (derive-style JSON) + native password-auth Postgres.
- `to_json`: a polymorphic builtin (
forall 'a. 'a -> str, the JSON
sibling of show) that serializes any value structurally — records become JSON objects (dropping the type name), lists/tuples arrays, nullary constructors "Name", and payload constructors {"Name": payload}. Same compile-time-specialization approach as show (no trait machinery); works on interp / C / Wasm. Removes hand-written record→JSON writers (the mere-blog dogfood's PAIN B3).
- Native SCRAM-SHA-256: real SHA-256 / HMAC / PBKDF2 / base64 in the C
runtime, so a native binary authenticates to a password Postgres over plaintext (TLS still pending). Verified against a scram-sha-256 server.
- Native redis/mysql: two arena↔hex helpers complete the byte-buffer
FFI, so the whole contrib/db family — not just pg — compiles to native binaries. Verified driving a real redis.
1992 tests.
v0.1.5 — 2026-07-10
Native full-stack: a web + Postgres app now compiles to a single native binary. Driven by the mere-blog dogfood.
- Native FFI runtime (C backend): the
tcp_*/mem_*/str_ptr
externs that contrib/db (pg / mysql / redis) speak — previously host-provided over the Wasm linear memory — get native implementations: a Wasm-style flat byte arena (32-bit offsets) plus POSIX sockets. So the pure-Mere wire-protocol drivers run in a native binary.
- Native HTTP server:
http_serveruns a POSIX accept loop with the
same handler contract as the Node host ("METHOD URL" + http_set_* / http_get_header / http_current_body).
- Native crypto/util: a real FIPS 180-4
sha256_hexand a
/dev/urandom-backed gen_request_id (password hashing + session ids).
- Result:
mere -c app.mere | clangyields a self-contained native web+DB
server — no Node, no Wasm. (Postgres SSL / SCRAM auth on native are stubbed for now; use trust / plaintext.)
- `let` main diagnostic: a top-level
let main = …now warns on the
compile paths (not just the interpreter) with a message pointing at the entry-point convention, instead of surfacing a cryptic downstream wat2wasm clash.
- Fix: the C backend escaped
\n/\tin string literals but not
\r, so a carriage return broke the emitted C string (hit compiling pg's COPY unescape).
1978 tests.
v0.1.4 — 2026-07-10
Driven by the mere-blog dogfood (a Rails-ish blog on contrib/http + contrib/db/pg).
- `let` constructor/record patterns on all backends:
let Ctor (a, b)
= e and let Rec { f = x } = e now compile on the C, Wasm, and LLVM backends (previously only the interpreter accepted them; the compiled backends handled just P_var / tuple / wildcard). Each backend desugars the general case to a single-arm match.
- `contrib/orm`: a small, DB-agnostic typed layer — row decoders
(Orm.dec_int / dec_str / dec_bool / dec_str_opt + decode_rows) over the str option list rows the contrib/db drivers return, plus matching JSON encoders (Orm.enc_int / enc_str / enc_bool / enc_str_opt / enc_obj / enc_arr). The ML answer to reflection-based ORMs.
1972 tests.
v0.1.3 — 2026-07-10
Closes the last dogfood finding from the mq CLI.
- String ordering:
<,<=,>,>=now work directly onstr,
comparing lexicographically (in addition to int). Previously the typer forced both operands to int, so "a" < "b" failed to typecheck and callers had to route through str_compare/ord. Works across all four backends (interp / C / Wasm / LLVM); the int default for unresolved operands is preserved, so existing code is unaffected.
- contrib/json fix: v0.1.2 claimed the serialiser had moved into
module Json, but the functions were dropped rather than re-added, so the release actually shipped a parser-only json.mere. They are now restored inside the module — Json.to_json_str (Json.parse_json s) type-checks and round-trips as intended.
1961 tests.
v0.1.2 — 2026-07-10
More dogfood-driven fixes (from the mq CLI).
- `read_stdin`: reads all of stdin as a
str(interp + C backend), so
CLIs can filter piped input (echo … | mq '.query').
- contrib/json: the serialiser (
to_json_str/to_pretty_str) moved
into module Json and writer.mere was removed, so parser and writer share one json type — to_pretty_str (parse_json s) now composes.
1947 tests.
v0.1.1 — 2026-07-10
Fixes surfaced by dogfooding two real apps on top of Mere: a realtime collaborative editor (mere-notes, Wasm) and a native jq-like CLI (mq, C backend). Mostly C-backend and contrib hardening.
- Native CLI I/O: the C backend implements
args()(argv → str list),
so a compiled Mere program can read its arguments.
- C backend correctness: respect shadowing of the
joinbuiltin (a
local join no longer compiles to pthread_join); fix cross-host capture merging in inner-fn lifting (composing two modules that each have a same-named inner fn no longer corrupts captures); mask chr's byte index so out-of-range input can't read past the char table.
- C backend parity / ergonomics:
str_eqworks as a function (not
just the == operator); str_of_int pulls in the show_int helper; type annotations accept qualified module types (Module.t).
- contrib hygiene:
contrib/jsonandcontrib/csvno longer run
self-test demos on import (library-clean, module-only).
- Package system v0.2 (from the mere-notes dogfood):
mere install
(manifest + git/subdir deps + lockfile), a [host] entry + mere serve that vendor and run the Node host, and distribution via release.yml + scripts/install.sh.
1945 tests.
v0.1.0 — 2026-07-09 (first tagged release)
First public tagged release of the Mere compiler. What it contains:
- The language: HM inference + let-polymorphism, region / view /
Trivial[R] memory model with refined borrow modes, capability-passing effects, and feature-parity codegen to C / LLVM IR / Wasm alongside the tree-walking interpreter. 1936 tests.
- Self-host: lexer / parser / typer / eval / fmt / codegen are written
in Mere and compile themselves through the Wasm pipeline.
- Concurrency:
spawn/channel/join+par_mapon all four
backends, with a Send / Sync type discipline.
- Package system v0.2:
mere install(manifest + git deps with
monorepo subdir, transitive resolution, mere.lock) and a [host] entry + mere serve that vendor and run the Node runtime host — so an app builds and runs from just an installed mere, no source tree.
- Distribution:
release.ymlbuilds prebuilt binaries for macOS
(arm64 / x86_64) + Linux (x86_64) on each v* tag; scripts/install.sh installs one without an OCaml toolchain.
Work since the entries below (2026-07-07…09): self-host frontier completion (module-import inlining fix; while / brace-block / vec / map builtins), the concurrency stack, and the package-system + distribution tooling above.
2026-07-06 — Tutorial: implement type inference in Mere (roadmap step 4, third of three — series complete)
Third and final tutorial in the initial series (direction paper's educational thread). Builds the unification engine at the heart of Hindley-Milner over a tiny lambda calculus + let.
docs/tutorial-type-inference.md— auto-published. Builds
bottom-up: the expr / ty ASTs (with TVar unification variables), fresh-var supply (single-slot vec), the substitution + apply, the occurs check, unify (tuple-match core), and infer (6 cases). Then the honest HM leap section: explains why the monomorphic let here rejects let id = fn x -> x in id id, and what let-generalization / instantiation add — pointing to the real contrib/typer (which runs in the browser playground).
examples/tutorial_type_infer.mere— the worked example. Verified
end-to-end: fn x -> x : t0 -> t0, fn f -> fn x -> f x : (t6 -> t7) -> t6 -> t7 (arrow domain parenthesized), (fn x -> x) 5 : int, let id = fn x -> x in id true : bool, 1 2 : TYPE ERROR (int isn't a function), id id : TYPE ERROR (occurs check).
The tutorial series now covers all three planned tracks:
- REST API (
contrib/http— routing / path params / CRUD) - Redis client (raw TCP externs — the RESP protocol)
- Type inference (the HM unification engine — self-host compiler
internals)
Together they span the three positioning directions: Wasm-first backend (1), the network/systems layer (2), and the educational PL-implementation angle (3).
2026-07-06 — Tutorial: build a Redis client in Mere (roadmap step 4, second of three)
Second educational tutorial. Builds a minimal Redis client from the raw TCP + memory externs to teach the RESP wire protocol — the layer contrib/db/redis sits on top of.
docs/tutorial-redis-client.md— auto-published (nav + sitemap +
search). Covers RESP in a table (+ simple / - error / : int / $ bulk / * array), then builds bottom-up: the tcp_* + mem_* externs, the reply variant, byte / line / exact-count readers, the first-byte dispatch parser, and command encoding (*N\r\n$len\r\narg\r\n). Ends pointing at the full contrib/db/redis (RESP3, pipelining, TLS, pub/sub) + queue / stream / lock modules + the pg driver (same mem_* pattern).
examples/tutorial_redis_client.mere— the worked example.
Verified end-to-end against redis:7: PING → +PONG, SET → +OK, GET → bulk "hello mere", GET missing → nil, DEL → :1 — one reply type exercised per command.
Teaching point emphasized: bulk strings use a length prefix (not line scanning) because payloads can contain \r\n / NUL — so read_bulk reads an exact byte count via read_exact, unlike the CRLF read_line used for status / length lines.
Note: tcp_* externs need the Node runner's sync TCP worker; they are NOT available on Cloudflare Workers (no raw sockets) — called out in the tutorial.
2026-07-06 — Tutorial: build a REST API in Mere (roadmap step 4, first of three)
First educational tutorial (direction paper's step 4). A guided walkthrough that builds a minimal notes REST API on the contrib/http stack — create / list / fetch / delete over JSON, storage in-memory (no DB to set up).
docs/tutorial-rest-api.md— the tutorial, auto-published to the
docs site (nav + sitemap + search picked it up automatically). Builds the program up in 5 steps (route → store+create → list → path-param fetch → delete), each snippet grounded in real code, then points to next steps (Postgres persistence, ETag concurrency via http_rest_notes, auth, middleware).
examples/tutorial_notes_api.mere— the complete worked example
the tutorial references. Verified end-to-end: create → 201, list → JSON array, fetch → full note, missing → 404, delete → {"deleted":true}, list-after-delete correctly skips the removed note (the list walk gates on map_has, so a deleted id left in the order vector drops out silently).
Teaching points surfaced in the tutorial: the \{ escape for JSON object literals (bare { starts string interpolation), top-level let rec for recursive helpers (Wasm backend disallows let rec nested in a fn body), and route_pattern :id captures working across GET and DELETE.
README gains a pointer under Documentation.
2026-07-05 — Cloudflare Worker: package registry v0.1 (JSON API)
Second CF Worker sample from the direction paper. Read-only JSON API over a static-ish bundled package list — the foundation for mere install speaking a normalized endpoint instead of hitting GitHub directly.
examples/cloudflare-worker-registry/:
main.mere— routes + response builders + naive JSON scan/escapeworker.js— CF entry, exposes bundledpackages.jsonto Mere via
a cf_registry_data () extern
packages.json— v0.1's source of truth (3 sample entries:
mere-http / mere-db / mere-json). To add a package: edit + rebuild
wrangler.toml,build.sh,local_test.js,README.md
Endpoints:
GET /landing HTMLGET /pkgwhole registryGET /pkg/:nameone package's metadataGET /pkg/:name/latestlatest versionGET /pkg/:name/:versionspecific version
Verified via node local_test.js — 21 assertions across 8 request scenarios, all pass:
- Landing 200 + HTML
/pkglists all 3 packages- Package metadata has owner / latest / versions
/pkg/mere-http/latestreturns injected{name, version, tarball, ...}- Specific version endpoint works
- Unknown package → 404
- Unknown version → 404
- POST → 404 (only GET supported)
Two landmines fixed during shipping:
- Balanced-brace parser bug: earlier
whileloop seti = nto
break out but then the "start >= n → empty" check false-negatived every extraction. Restructured with an explicit done flag.
- Unescaped `\n` in 404 body:
resp_not_foundsplicesmsginto
response body JSON without escaping. Added a json_esc pass.
Wasm size: 11 KB. v0.2 roadmap in the README (GitHub tag fetching, KV cache, publish endpoint, mere install CLI).
2026-07-05 — Cloudflare Worker: playground snippet share (KV-backed)
Turned the CF Worker template from "hello, method+path echoed" into a real sample that motivates Workers over static hosting: a playground-snippet share service backed by Cloudflare KV.
Endpoints:
GET /landing HTMLPOST /shareraw code → 8-hex id +KV.put, returns{id, url}GET /s/:idreturns stored snippet, 404 if unknown
The async KV binding on CF is bridged to sync Mere externs via two conventions:
- Pre-fetch (read path): worker awaits
KV.get(id)BEFORE
calling Mere; the value lives in a module-scoped currentKvLookup and Mere reads it via cf_kv_lookup ().
- Outbox (write path): Mere emits
kv_put:{key,value}in the
response JSON; worker honours it AFTER the handler returns via KV.put(key, value).
Body handling uses the same "extern-not-JSON" convention: JS stashes the raw request body in a module scratch, Mere reads it via cf_body (). This sidesteps a JSON-in-JSON double-escape bug where \n inside stored snippets turned into \\n after round-trip.
Local smoke test (local_test.js) verifies six assertions with an in-memory KV mock:
- Landing page 200 + text/html
POST /sharereturns 201 + JSON id/url, KV was writtenGET /s/:idreturns 200 with the ORIGINAL code (newlines
preserved byte-for-byte — regression for the double-escape bug)
- Unknown id → 404
- Empty body → 400
- Unknown route → 404
Wasm size: 5.7 KB → 8.1 KB (added routing + JSON escaper + KV outbox construction).
2026-07-05 — Cloudflare Worker template (roadmap step 2)
Step 2 of the direction-paper roadmap. A minimal, self-contained template that runs a Mere program as a Cloudflare Worker — 5.7 KB compiled wasm, no npm runtime deps, V8-isolate compatible.
examples/cloudflare-worker/:
main.mere— 30-line handler. Registers a request handler via a
new cf_on_fetch: (str -> str) -> unit extern. Handler receives JSON-encoded request, returns JSON-encoded response.
worker.js— CF Worker entry (ES module). Providescf_on_fetch
+ the standard prelude stubs, marshals Request ↔ JSON ↔ Mere closure via the existing __lang_bump + __indirect_function_table machinery.
wrangler.toml— CF Worker deploy config.build.sh—mere -w main.mere → main.wat → main.wasm.local_test.js— Node 22-based smoke test using native
Request/Response (no wrangler/miniflare required for verification).
README.md— layout, build/deploy commands, request/response
protocol, and an explicit "what doesn't work on CF" section (no TCP / subprocess / fs — those are Node-runtime-specific externs).
Verified locally via node local_test.js — three requests round-trip:
GET /→hello from Mere on Cloudflare — GET /GET /hello?name=world→ same shape, path echoedPOST /submit→ method + path echoed
Actual wrangler deploy requires a Cloudflare account and is left to the operator (README.md documents the commands).
Deliberate non-goals for this template: KV / R2 / D1 bindings, Durable Objects, auto-rebuild watcher. All addable incrementally.
2026-07-05 — package system v0.1: .mere_modules/ walk-up resolution
First step of the direction-paper roadmap. Extends the import resolver in lib/parser.ml with Node.js-style node_modules walk- up semantics — a project puts vendored packages under .mere_modules/, and any file in the tree can import "pkg/module.mere" without relative ../ navigation or -I flags.
Resolution order (relative paths only; absolute paths still resolve literally):
<importer_dir>/<path>— historical behaviour<nearest .mere_modules up>/<path>— new (Node-style walk-up)-Idirs +MERE_PATHenv — historical, order preserved
Deliberate v0.1 non-goals (documented in docs/packages.md):
- No
mere.tomlmanifest yet (track deps by git URL / commit) - No
mere installcommand (git clone / submodule / tarball drop) - No central registry (planned for v0.3+, design in internal notes)
- No version resolution (walk-up first-match-wins)
Vendoring workflow — three equivalent options, all documented:
git clone https://github.com/<owner>/<pkg> .mere_modules/<pkg> # or git submodule add https://github.com/<owner>/<pkg> .mere_modules/<pkg> # or curl -L https://example.com/<pkg>.tar.gz | tar xz -C .mere_modules/
New docs page docs/packages.md with layout, semantics, precedence, and a self-contained demo pointer. Demo examples/pkg_demo/:
main.mere— 3 lines,import "hello/greet.mere"; print (greet "world").mere_modules/hello/greet.mere— one-liner greeter package- End-to-end verified:
mere -w examples/pkg_demo/main.mere→
hello, world!
Three new regression tests in test/test_basic.ml:
- Single-level walk-up (
.mere_modules/alongside entry file) - Deep walk-up (entry file in
app/handlers/, modules dir above) - Cross-package imports find the same
.mere_modules/root
All 7 spot-checked existing demos (http_blog, http_admin_dash, http_router_demo, http_ws_chat, db_redis_pubsub, subprocess_demo, gh_stars) recompile unchanged. Test suite: 1846 → 1849.
2026-07-05 — contrib/db/redis_ratelimit: distributed fixed-window limiter
Multi-instance version of contrib/http/ratelimit (which is in-process only — two Mere HTTP servers would each keep their own counter, so a caller can rotate through instances to bypass). This version puts the bucket counter in Redis so N instances share one budget per key.
Standard INCR + EXPIRE pattern:
redis_rate_over_limit fd key window_sec max→ bool
Increments the counter for the current window and returns true if count > max. Attaches TTL on the first hit of a bucket via EXPIRE; subsequent hits are single-INCR calls. Fail-open on network error (returns false).
redis_rate_count fd key window_sec→ int
Peek without incrementing. Useful for X-RateLimit-Remaining headers.
Bucket key layout: <key>:<epoch/window> — all instances at the same wall-clock second share the counter. Not sliding-window (a burst right at the boundary can spike to 2 x max); document for callers who need bursty tolerance.
Demo examples/db_redis_ratelimit.mere — 3-per-2-sec policy: attempts 1-3 return ok, 4-5 return BLOCKED, then a sleep_ms 2200 triggers a window roll and the next attempt returns ok.
2026-07-05 — contrib/os/parallel_map: N shell commands in parallel
Sits on top of contrib/os/subprocess. No new externs. Uses shell backgrounding (&) + wait + tmpfiles to run N children concurrently under the OS scheduler, then reads their stdouts back in index order (not completion order).
The "cheap dogfood" step between the sync subprocess_run primitive and a native worker_spawn / worker_await pair that a future worker_threads shipping will bring.
parallel_map : str list -> str list
Verified end-to-end:
- 4 x
sleep 1 && echo <label>→ 1135 ms wallclock (~max of
individual times, not sum of 4000), results [A; B; C; D] in submitted order
- Mixed timings (0 / 2 / 1 sec) → 2099 ms wallclock, results
[instant; two-sec; one-sec] — the slowest child at index 1 dictates wallclock; ordering follows input order, not completion
Not suitable for streaming (all children must exit before return), very short-lived children (fork overhead dominates), or output containing the fixed sentinel __MERE_PMAP_SEP_9c3d4f7a__. Documented in the module.
2026-07-05 — contrib/os/subprocess: sync shell-out (Q-012 Path A)
First shipping toward the concurrency-primitive design (see design notes in the project's internal notes). Path A of the plan — the "no language change, immediate utility" step before a proper spawn / channel primitive.
Three externs backed by Node's child_process.spawnSync:
subprocess_run cmd stdin -> str— shell-execute, feed stdin,
return stdout. Timeout 30 s, buffer cap 16 MiB per stream.
subprocess_status ()-> int — exit code of the last run
(0 = ok, nonzero = child, -1 = signal / timeout).
subprocess_stderr ()-> str — stderr of the last run.
Blocking by design. subprocess_run holds the whole Wasm frame until the child exits — a Mere HTTP server MUST NOT call it inside a request handler.
Deliberate scope: no async / parallel-collect primitive. For parallelism today, users can shell-background inside one call:
subprocess_run "sh -c '(child1 > /tmp/r1) & (child2 > /tmp/r2) & wait; " ++ "cat /tmp/r1; echo ---; cat /tmp/r2'" ""
The two children run concurrently under the OS scheduler; only collection is serial. A proper worker_spawn / worker_await pair is scheduled for Q-012 step 3 (post worker_threads restructure).
Demo examples/subprocess_demo.mere verifies all four flows:
date -u→ status 0, timestamp captured- text piped into
wc -w→ 5 false→ status 1, stderr captured- two
sleep 1in parallel via shell&→ 1055 ms wallclock
(not 2000+ ms — real OS-level parallelism)
Wired into both run_wasm.js and run_http_server.js via the same factory pattern as http_fetch_env. 1846 tests pass.
2026-07-05 — contrib/http/websocket: RFC 6455 hub
WebSocket support in the standard shape:
- Handshake —
GET /ws/<channel>withUpgrade: websocketand
Sec-WebSocket-Key → 101 Switching Protocols with the standard Sec-WebSocket-Accept: base64(sha1(key + magic)) computation.
- Text frame codec — encode server → client (unmasked), decode
client → server (masked with per-frame XOR key). Both length forms (7-bit / 16-bit / 64-bit) supported.
- Channel pool —
/ws/<channel>sockets go into a per-channel Set;
ws_broadcast writes to every socket, auto-relay writes to every socket EXCEPT the sender.
- Close + ping — client close → echo close + destroy socket. Ping
→ reply pong with same payload.
Public API (contrib/http/websocket.mere):
ws_broadcast channel payload -> unit— server → all clients.ws_client_count channel -> int— for a "0 listeners → skip
work" fast-path.
Deliberate design choice: individual client frames are NOT delivered to Mere. The glue auto-relays them to peers on the same channel (hub pattern), covering chat / cursor-share / collaborative- edit demos without needing an in-Wasm callback per frame. Per-frame Mere handlers would require a callback-into-Wasm design and stay deferred.
Not supported (documented):
- Binary opcodes (0x2) — silently dropped
- Fragmentation (FIN=0 continuation) — every frame treated as full
- Payloads > 2^32 bytes (unrealistic for browser peers)
Demo examples/http_ws_chat.mere — auto-relay chat + admin POST /announce → ws_broadcast. Verified with a native WebSocket probe on Node 22:
- A sends "hello from A" → B receives it, A does NOT (hub excludes
sender)
POST /announce {"msg":"hello everyone"}→{"delivered_to":2},
both A and B receive [admin] hello everyone
All 5 spot-checked existing HTTP demos (router / blog / chat / pubsub_chat / admin_dash) recompile and serve as before — the Upgrade hook is a new event handler on the same server, so non-upgrade requests are unaffected.
1846-test OCaml suite passes.
2026-07-05 — examples/http_admin_dash: integration dogfood
One small admin console exercises six of the modules shipped over the last day in a single mere file (~200 lines):
contrib/http/router—route_prefix "/admin"+ exact routescontrib/http/session— cookie sessions (random 16-hex ids)contrib/http/csrf— synchronizer-token on the "run job" POSTcontrib/http/basic_auth— Prometheus scrape gate on/metricscontrib/http/metrics—/metrics+with_metricsmiddlewarecontrib/http/cache—cache_no_storeon admin pagescontrib/db/redis_lock— "only one instance runs the job" mutex
Feature: press the dashboard's "run job" button. The server acquires a Redis lock, sleeps 500 ms (simulated work), releases. A second instance clicking during the sleep window hits redis_lock_acquire → None and returns 409 "contended".
Verified multi-instance end-to-end (two processes on :8080 + :8081 sharing one Redis at :15650):
- Login flow: admin/adminpw → session cookie → dashboard 200 with
a CSRF token in the form's hidden input.
- Concurrent kick: instance A returns 200
"job ran successfully
(held lock for 500 ms)", instance B returns 409 "contended: another instance is running the job".
- CSRF check: POST without the token → 403.
/metrics: without Basic Auth → 401,-u scraper:s3cret→ 200
with jobs_run_total 1 in the scrape body.
The demo also documents the multi-instance run recipe in the header comments so users can reproduce the race locally with two PORT=… invocations against the same Redis.
2026-07-05 — contrib/db/redis_lock: distributed mutex + gen_request_id shared
Standard SET key <token> NX PX <ttl_ms> acquire with compare-and- delete release via Lua EVAL. Enough for "at most one worker across N processes should be running this job right now"; not enough for critical-section-with-consequences workloads (RedLock, CP consensus).
redis_lock_acquire fd key ttl_ms -> str option
Some fencing token on success, None on contention.
redis_lock_release fd key token -> bool
Compare-and-delete Lua: only deletes if the key's current value matches the caller's token. Prevents "A's TTL expires, B acquires, A's stale Release blows away B's lock" bugs.
Also hoisted gen_request_id (16-hex random) from run_http_server.js into scripts/pg_env.js so CLI Mere programs under run_wasm.js can use it too — the lock's fencing tokens were the immediate trigger, but any test harness minting session ids or correlation ids benefits. All 7 existing consumers recompile unchanged.
Demo examples/db_redis_lock.mere walks the six-step race:
- A acquires (fresh token)
- B tries → None (contention)
- A releases → true (CAS matches)
- C acquires (fresh token)
- Impostor tries release with wrong token → false, lock intact
- E tries → None (C still holds), C releases → true
2026-07-05 — contrib/http/cache: Cache-Control postures + ETag / 304
Rounds out the middleware family (session / basic_auth / csrf / metrics / cache). Three helpers for the three canonical cache postures plus an ETag + If-None-Match short-circuit:
cache_immutable seconds
Sets Cache-Control: public, max-age=N, immutable. For asset URLs with a content hash in the path.
cache_private seconds
Sets Cache-Control: private, max-age=N. For per-session pages that can be briefly re-used.
cache_no_store ()
Sets Cache-Control: no-store, no-cache, must-revalidate + Pragma: no-cache. For login / secrets / POST redirects.
etag body— quoted SHA-256 hex, strong.if_none_match tag— readsIf-None-Match,str_eqcompare.
Doesn't parse * wildcards or comma lists (documented).
Demo examples/http_cache_demo.mere verifies all three postures + the 304 round-trip: matching If-None-Match → 304 with empty body, mismatching → 200 with fresh ETag.
2026-07-05 — contrib/db/redis_stream: consumer groups (XGROUP / XREADGROUP / XACK / XPENDING)
Extends the stream module with the load-balanced worker pattern — Redis' Kafka-consumer-group equivalent.
Added:
stream_group_create fd key group start_id— XGROUP CREATE with
MKSTREAM so producer/consumer bootstrap order is irrelevant. "0" = read from beginning, "$" = only new arrivals.
stream_group_read fd key group consumer count— XREADGROUP
GROUP … > (un-delivered only). Server remembers per-consumer in-flight entries in the PEL.
stream_ack fd key group ids— XACK; returns n acked.stream_pending_len fd key group— XPENDING summary → total
un-acked count.
XCLAIM / XAUTOCLAIM for reassigning stuck entries stays deferred.
Demo examples/db_redis_stream_groups.mere walks the full cycle: one group workers with two consumers A + B share 4 XADD'd jobs. XREADGROUP delivers 1-2 to A and 3-4 to B (no overlap — Redis tracks what's been handed out). PEL sits at 4, then 2 after A ACKs its half, then 0 after B ACKs. A follow-up XREADGROUP > returns empty since the group is drained.
2026-07-05 — contrib/db/redis_stream: XADD / XREAD / XLEN
Third leg of the Redis event story:
redis_pubsub broadcast-and-forget, no history redis_queue exactly-one-worker-claims (BRPOP) redis_stream durable append-only log, replayable
Streams are Redis' Kafka-lite — entries live in an append-only radix tree with server-generated <ms>-<seq> ids. Consumers either resume from a chosen id or use consumer groups (deferred here).
Public API:
stream_add fd key fields -> str option
XADD with * id, returns the new entry id.
stream_read fd key after_id N -> (id, fields) list
XREAD COUNT N STREAMS key after_id. after_id is exclusive; use "0" for a full replay.
stream_len fd key -> int
XLEN, -1 on error.
Out of MVP scope: XREADGROUP / XACK / XPENDING consumer groups, MAXLEN caps, blocking reads (XREAD BLOCK N). Documented in the module header.
Demo examples/db_redis_stream.mere verifies the full flow: 3 XADDs → XLEN=3 → full replay from 0 recovers all fields → resume from mid-stream id yields only the tail → past-the-tail returns empty.
2026-07-05 — contrib/http/csrf: synchronizer-token CSRF middleware
Sits on top of contrib/http/session: the cookie session id is the store key, the token is a fresh 16-hex random via gen_request_id () minted on first csrf_token_for per session and re-used for the lifetime of the session.
Public API:
csrf_new_store ()csrf_token_for store session_id— idempotent per sessioncsrf_validate store session_id tok— boolcsrf_hidden_input token—<input type="hidden" name="_csrf" value="…">snippet
Design choice: kept as primitives rather than a with_csrf middleware because content-type detection (form vs JSON) and body re-parsing are handler-specific concerns; handlers already read the body via form_field / body_field, so passing the value into csrf_validate is a one-liner where the caller already is.
Demo examples/http_csrf_demo.mere — a mutable-message form. Verified: missing _csrf → 403, wrong token → 403, correct token → 303 redirect with the message actually persisting.
2026-07-05 — playground: wordcount demo + build tail-call flag
New live-docs demo — a client-side text stats tool: char / word / line counters computed by a Mere function compiled to Wasm, wired into a textarea + three display slots via contrib/dom. Reuses the Phase 48 C2 frontend FFI (closure dispatch through the exported function table); no new externs.
Files:
contrib/site/playground/wordcount.mere—count_words/
count_lines implemented as manual character scans (folds runs of whitespace into one word boundary; treats \n as line separator so an N-line file reports N).
contrib/site/playground/wordcount.html— form + wire wasm,
matches the styling of the counter / echo demos.
- Nav entry added to all sibling playground pages + the SSG's
playground index.
Build fix: contrib/site/build_full.sh now invokes wat2wasm --enable-tail-call. The wordcount demo emits return_call / return_call_indirect (Wasm tail-call proposal) via its while loop + inner-lifted closures, and the pre-flag site build rejected those opcodes. Enabled by default in Chrome / Safari / Firefox 129+ / Node 22+, so no runtime compatibility loss.
Live path: https://merelang.github.io/mere/playground/wordcount.html after the next Pages deploy.
2026-07-05 — contrib/db/redis_hll: HyperLogLog cardinality estimators
Thin wrappers on Redis's PF* family. Approximate distinct-count with fixed 12 KiB per key regardless of true cardinality (~0.81 % standard error). Complements the exact-set path (SADD / SCARD) for cases where the memory budget matters more than the exact number — unique visitors, distinct URLs, unique IPs per hour.
hll_add fd key values— PFADD; returns1if the estimate
moved, 0 if all values were already there, -1 on error.
hll_count fd keys— PFCOUNT; approximate cardinality. Single
key = that key's count; multiple keys = the union cardinality (server-side merge into a temp HLL).
hll_merge fd dest srcs— PFMERGE; materializes the union of
srcs into dest. Idempotent.
Demo examples/db_redis_hll.mere verifies both the union-via- count and union-via-merge paths: 3 users on shard-a, 3 users on shard-b (one overlap), true distinct = 5 across both, and both merge paths report 5.
2026-07-05 — contrib/log: level filtering + field-taking variants + LOG_LEVEL env
The base log_debug / log_info / log_warn / log_error functions were already there but always printed. Now:
set_log_level "debug" | "info" | "warn" | "error" | "off"sets
the threshold at runtime. Default remains info.
log_from_env ()readsLOG_LEVELfrom the process env. Unset
or empty leaves the default in place — a demo without any explicit configuration still gets info-and-above.
log_debug_f/log_info_f/log_warn_f/log_error_f—
field-taking variants. Same filter applies; structured (str, str) list fields become JSON keys next to msg.
The threshold lives in a single-cell vec_new () allocated once at module-load time (post import-flatten). Note for future contrib authors: module-level mutable state must use ; (top-level decl) rather than let ... in — the latter turns the rest of the file into one expression that import discards. Learned the hard way here; documented in the module.
Demo examples/log_levels_demo.mere exercises all levels + runtime switching. Verified:
- default: info + warn + error + info_f + error_f visible.
LOG_LEVEL=debug: debug included.LOG_LEVEL=warn: only warn + error.LOG_LEVEL=off: silent (until runtimeset_log_levelre-enables).
All 8 existing log consumers (http_users_db, http_jwt_api, http_ci_dashboard, http_feed_reader, http_csv_export, http_wiki, http_file_upload, http_webhook_receiver) recompile unchanged. Test suite: 1846.
2026-07-05 — contrib/http/basic_auth: RFC 7617 Basic Auth middleware
Small addition to gate internal endpoints — /metrics scraping, /admin dashboards, cron-triggered endpoints. Two entry points:
with_basic_auth realm user pass handler— single credential pair
(compile-time constant).
with_basic_auth_pred realm predicate handler— delegate the
credential check to a (user, pass) -> bool predicate. Useful when the accepted set comes from an env var or in-process map.
Missing / wrong credentials → 401 with WWW-Authenticate: Basic realm="…", charset="UTF-8". Handler is NOT called on failure. Simple str_eq compare (not timing-safe) — documented as a gate, not a production auth layer.
Added base64_encode / base64_decode externs to scripts/pg_env.js for utf8 <-> standard-alphabet base64 round-trip (the existing _hex variants take a hex detour that's overkill for Basic Auth's plain user:pass payload).
examples/http_metrics_demo.mere gained a Basic-Auth-gated /metrics route as the first consumer. Verified: no-auth → 401, wrong creds → 401, -u scraper:s3cret → 200 with metrics body. Ungated routes (/, /work) still return 200.
2026-07-05 — Blog-engine papercuts: lexer + typer polish
Two friction points surfaced during the http_blog dogfood get proper first-class fixes now (previously the demo worked around them).
String line-continuation. "foo \<newline> bar" now lexes as "foo bar" — the backslash-newline sequence eats the newline itself plus any leading spaces / tabs on the next line (Python / Rust convention). Long HTML snippets, SQL statements, and log messages can be broken across source lines without smuggling in a \n or indent characters, and without piecing them back with ++ string concatenation. All existing escapes (\n, \t, \r, \", \\, \{) still work identically.
SCREAMING_SNAKE_CASE hint on `let`. let DB_URL = "..." used to fail with a bare type error: unknown constructor in pattern: DB_URL because Mere reserves uppercase-first identifiers for constructors. The typer now recognises the shape (starts uppercase, has no lowercase letters, either ≥ 3 chars OR contains _) and adds:
help: Mere reserves uppercase-first identifiers for constructors. If you meant a value binding, rename to db_url.
The heuristic explicitly excludes single-letter names like let X = … (too plausibly a one-shot constructor placeholder) and still yields to the standard did-you-mean suggestion when one exists (let x = Cnos (…) → did you mean 'Cons'?).
Both changes come with regression tests. Full suite: 1838 → 1846.
2026-07-05 — sse_bridge_from_redis: multi-instance SSE fanout
New extern in contrib/http/sse.mere:
sse_bridge_from_redis channel host port -> unit
Spins up (or reuses — idempotent per channel) a persistent RESP2 subscriber in the Node runner. Every incoming message-shaped reply on channel is forwarded to the JS-side SSE broadcast for the same channel name. Result: N Mere HTTP instances behind a load balancer, all subscribed to the same Redis channel, deliver posted messages to every SSE client regardless of which instance holds the subscription.
Two moving parts:
scripts/sse_redis_bridge.js— new. Async RESP2 subscriber
(Node's net.Socket), auto-reconnect on error / close with a 1 s backoff. Parser handles arrays / bulks / simple strings / integers — enough for the SUBSCRIBE reply shape.
contrib/http/http.glue.js— factored the inner fanout code out
of the Mere-facing sse_broadcast extern into a JS-callable broadcast(channel, payload) helper. makeHttpGlue() now returns { glue, attach, broadcast }; the bridge factory receives broadcast and calls it directly (no Mere-heap ptr boundary crossing).
Demo examples/http_pubsub_chat.mere verifies end-to-end:
- Two instances started on
:8080+:8081against a shared
Redis; both subscribe to chat.
- POST to
:8080returns{"delivered_to":2}(Redis sees two
subscribers) and the message appears on BOTH SSE streams.
- POST to
:8081— same behaviour in reverse.
http_serve and the pubsub subscriber coexist because the subscribe socket lives entirely in JS (Node's event loop), avoiding Mere's single-threaded per-frame constraint.
2026-07-05 — contrib/http/session: consolidate cookie-session pattern
Seven demos (http_blog, http_todo_app, http_users_db, http_todo_pg, http_mini_blog, http_feed_reader, http_cookie_session) all hand-rolled the same five-line dance: map_new (), read session= cookie, look up user, mint id on login, Set-Cookie. Consolidate:
session_new_store ()— opaque store handle (amapunder the
hood; pre-migration demos still compile against map_has etc.).
session_current store— current user id or"".session_login store user— mints a random 16-hex id via
gen_request_id (), sets Set-Cookie: session=…; Path=/; HttpOnly; SameSite=Lax.
session_logout store— removes the entry + emitsMax-Age=0.session_require store login_url— returnsstr option;None
side-effects a 303 to login_url.
Behavioural upgrade: sessions now use gen_request_id () (crypto random) instead of the demos' old "s-" ++ username — non-guessable ids, plus HttpOnly; SameSite=Lax cookie attributes by default.
examples/http_blog.mere migrated as the first consumer. All six CRUD flows still work end-to-end (login → post → view → edit → delete). The other six demos continue to work unchanged and can migrate incrementally.
2026-07-05 — contrib/http/metrics: Prometheus-style metrics + middleware
A small registry of counters and gauges plus a text-format exporter and a GET /metrics handler suitable for direct mount in a route table. Ships an auto-counting middleware with_metrics that increments http_requests_total{method, path} and adds request duration into http_request_duration_ms_sum + _count for every request (Prom's "summary" idiom, no percentiles).
Public API:
metric_declare_counter name help/metric_declare_gauge name help
— register + attach HELP/TYPE metadata (rendered once per name).
metric_inc name labels— counter += 1.metric_add name labels n— counter += n.metric_set name labels v— gauge = v.metrics_render ()— Prometheus text-format string.metrics_handler req— mount asGET /metrics.with_metrics handle— middleware wrapper.
Storage is a plain map_new () keyed by name or name{labels}; values are int (millisecond durations, counts). Float values, configurable histogram buckets, and label-value escaping are out of MVP scope.
Also added now_ms extern to run_wasm.js (previously only in run_http_server.js) so contrib modules that pull it work under either runner.
Demo examples/http_metrics_demo.mere — four routes (/, /work with a 50 ms sleep, POST /error, /metrics) verify the auto- counters, business counters, and duration accumulation. /work's http_request_duration_ms_sum sits at ~55 ms after one hit; errors_total increments only on POST /error.
2026-07-05 — examples/gh_stars: first CLI demo
First Mere program that runs under run_wasm.js (not run_http_server.js) and makes outbound HTTP calls. Fetches https://api.github.com/repos/<owner>/<repo> and prints the star count, using:
arg_get 0forowner/repoargv.getenv "GITHUB_TOKEN"for optional Bearer auth (60 → 5000
req/hour when set).
http_fetch_hfor theAccept: application/vnd.github+json+
User-Agent headers.
http_fetch_response_header "X-RateLimit-Remaining"for the
rate-limit metadata line.
- Naive
"stargazers_count":<n>scanner (avoids pulling in
contrib/json which has a top-level self-test block that would execute on import).
Verified against merelang/mere (0 stars, fresh repo), rust-lang/rust (114325), sindresorhus/awesome (481588), and a 404 path (no-such-owner/no-such-repo-12345 → HTTP 404 with the response body printed).
2026-07-05 — redis_pubsub_run_forever + sleep_ms extern + tcp_worker end-event fix
Three related changes to make a real-world reconnecting subscribe loop possible in pure Mere.
`redis_pubsub_run_forever host port sub timeout_ms retry_ms handler` Opens its own sub fd, sends the SUBSCRIBE / PSUBSCRIBE commands from sub, dispatches messages via handler, and on PSClosed (or redis_connect failure) sleeps retry_ms then starts over. The handler receives PSClosed events too, so it can log / reset metrics / decide to bail (returning false from any invocation ends the loop cleanly). Non-draining redis_pubsub_subscribe variants are used so PSSubscribed events flow through the handler on every reconnect.
Subscription state is captured in a new PubsubSub record — { channels; patterns }.
`sleep_ms` extern — synchronous millisecond sleep via Atomics.wait on a private SharedArrayBuffer. Blocks the whole Wasm frame, so an HTTP server MUST NOT call this inside a request handler. Added to both run_wasm.js and run_http_server.js (both had a no-op sleep).
tcp_worker.js `end`-event handler — with allowHalfOpen: true, a peer FIN emitted end but not close, so a pending tcp_read hung indefinitely. Reproducible via CLIENT KILL TYPE PUBSUB on a subscribed connection. Added an on('end', ...) handler that marks the socket read-closed and wakes any pending read with EOF (respond(0, 0)), matching what the close branch already did.
Demo examples/db_redis_pubsub_reconnect.mere stages the failure in one process: subscribe → publish 2 → 2 deliveries → send CLIENT KILL TYPE PUBSUB → sub fd closes → loop sleeps 500 ms → reconnects + resubscribes → publish 2 more → 2 deliveries → exit.
Verified end-to-end against redis:7, plus the existing base pubsub + queue demos still work unchanged (regression check). 1838-test OCaml suite passes.
2026-07-05 — contrib/db/redis_queue: list-backed work queue
Complements redis_pubsub. Pub/sub is broadcast-and-forget; work queues are exactly-one-worker-claims-each-job. Standard Redis reliable-queue pattern wrapped:
redis_queue_push fd queue payload— LPUSH, returns new length.redis_queue_pop fd queue timeout_s— BRPOP with server-side
block. Some (queue, payload) on delivery, None on timeout. Client-side socket timeout is set to (timeout_s + 5) s as a safety net; timeout_s == 0 blocks forever on both sides.
redis_queue_pop_many fd queues timeout_s— priority multi-queue
BRPOP. Earlier queues in the list win.
redis_queue_len fd queue— LLEN.redis_queue_run fd queues timeout_s handler— event-loop helper
that retries on timeout; handler returns false to break out.
Explicitly out of scope for the MVP: ack / retry semantics (processing-list + RPOPLPUSH reconciliation), delayed jobs, and priorities beyond the multi-queue trick.
Demo examples/db_redis_queue.mere verifies push (returns 1,2,3,4), LLEN=4, FIFO order across three BRPOPs, priority fall- through via pop_many ["jobs.slow"; "jobs"], and the empty-queue timeout returning None.
2026-07-05 — http_fetch shared across both runners
http_fetch and friends now live in scripts/http_fetch_env.js and plug into both run_http_server.js (as before) and run_wasm.js (new). Any Mere CLI that declares extern fn http_fetch: ... can now make outbound calls under the plain runner — previously they had to boot the HTTP server runner just to get the extern env.
examples/http_client_auth.mere dropped its unused extern fn http_serve declaration and runs identically under both runners (verified against httpbin.org).
Also refreshed docs/http-demos.md: added a "Router API" primer covering route / route_pattern / route_prefix, and catalog entries for the recent blog and client_auth demos.
2026-07-05 — contrib/http/client: request + response headers, per-call timeout
The outbound http_fetch was fixed to a bare (method, url, body) shape — no way to attach an Authorization: Bearer … header, no way to read a Retry-After back off a 429, no way to shorten the 10 s default timeout for a cheap probe. Three new externs close that gap without breaking the existing 3-arg call:
http_fetch_add_header name value— attaches a header to the
NEXT fetch (host-side accumulator is cleared once the fetch fires, so a set-and-fetch pair is self-contained).
http_fetch_response_header name— case-insensitive lookup on
the LAST response. Only the final response block is exposed — redirect chains and 100-continue trailers are discarded.
http_fetch_set_timeout ms— one-shot override; 0 restores the
10 s default.
Ergonomic wrappers in contrib/http/client.mere:
http_fetch_h method url body headers— headers as(str * str) list.http_get_bearer url token— sugar over the common auth-header case.
scripts/run_http_server.js runs curl with -i and parses the final response header block (handling redirect / 100-continue prefaces by taking the LAST HTTP/… block) so the host doesn't need a temp file for header capture.
Demo examples/http_client_auth.mere verifies all four features end-to-end against httpbin.org: custom header round-trip, response header read, Bearer token, per-call timeout enforcement.
2026-07-04 — contrib/db/redis_pubsub: dispatch layer
redis.mere already carried the raw SUBSCRIBE / PSUBSCRIBE / PUBLISH primitives, but callers had to destructure the resulting RRArr replies by hand to tell a message from a pmessage from a subscribe confirmation. A separate module now does the classification once and returns a small variant:
type pubsub_msg =
| PSMessage of str * str — (channel, payload)
| PSPMessage of str * str * str — (pattern, channel, payload)
| PSSubscribed of str * int
| PSUnsubscribed of str * int
| PSPong of str
| PSTimeout
| PSClosed
| PSOther of redis_reply
redis_pubsub_next fd timeout_ms— read + classify one reply.
Uses the caller's timeout_ms to disambiguate the "short read" case: > 0 → PSTimeout, else PSClosed.
redis_pubsub_run fd timeout_ms handler— event-loop helper;
handler returns false to break out, loop also exits on PSClosed.
redis_pubsub_subscribe/redis_pubsub_psubscribe— non-draining
variants that leave the confirmation reply on the wire, so the dispatch loop sees each as a PSSubscribed event.
redis_pubsub_open host port— two-fdPubsubClientrecord
(publisher + subscriber connections) encapsulating Redis's "PUBLISH needs its own fd" rule.
redis_pubsub_show msg— one-line pretty-printer for access logs.
examples/db_redis_pubsub.mere rewritten to demonstrate the whole API, including PSUBSCRIBE with a matched-pattern delivery and a PSTimeout tick. Full RESP3 push (RRPush) is also routed through the classifier by recursing into the inner list.
2026-07-04 — contrib/http/router: route_prefix mount points
Third arm of route_entry: REPrefix of str * route_entry list. Declared via route_prefix "/mount" inner_routes, it nests a whole route table at a common URL prefix. Inner entries are stated relative to the mount point ("/" is the mount root, "/login" is "/mount/login", etc.), and if no inner entry matches the request falls through to the next outer entry (rather than the prefix "claiming" the URL).
Made the fall-through work cleanly by refactoring internal _try to return str option — Some body on match, None on no-match — with the top-level router invoking the fallback only if _try returns None. No behavioural change for pure-exact / pure-pattern route tables.
Dogfood in examples/http_blog.mere:
- All 9
/admin/*routes now live underroute_prefix "/admin"—
the admin subtree is declared as a self-contained table and reused as one entry.
- Edit / delete moved to
/admin/edit/:idand/admin/delete/:id
pattern routes — the hand-rolled query-string parse in edit_form_h (that reached into the raw request line because the router had already stripped the query) is gone. Cleaner URLs and one fewer papercut for the next demo author.
2026-07-04 — contrib/http/router: :capture path params
Extended route_entry from a bare tuple to a two-arm variant so the router can dispatch on patterns without breaking the existing exact-match API.
route(backwards-compatible) — exact-path entry, unchanged
signature. Existing 15 demos recompile with zero source changes.
route_pattern method path handler— new. Path segments starting
with : capture one URL segment each. Handler is str list -> str -> str (captures in source order, then req).
- Segment matching splits on
/, ignores leading and trailing
slashes, and requires arity to match exactly (no * glob).
Wired into examples/http_blog.mere — the previous not_found + str_starts_with "/post/" workaround is gone; blog now routes /post/:slug declaratively. examples/http_router_demo gained two-capture /user/:name/pet/:pet for reference.
2026-07-02 — Phase 54.36 runtime codegen bootstrap unblocked
Root-caused the "runtime OOB" that had been the last unresolved self-host gap since Phase 54.20 — turned out not to be a codegen bug but plain memory exhaustion.
Root cause: OCaml-side wasm codegen defaulted to (memory (export "memory") 64) — 64 pages = 4 MiB. Self-host parse_and_emit "42" allocates ~30 MiB at peak (prelude tokens + parsed AST + emit strbuf). The bump allocator has no memory.grow, so writes past 4 MiB trap.
Phase 54.20's 5/6-char boundary observation was a red herring: the allocation crossed the 4 MiB line at a specific input-dependent point that happened to correlate with name length in the isolation harness. Phase 54.23's higher-order-list_map hypothesis was similarly incidental.
Fix:
lib/codegen_wasm.ml— default memory 64 → 1024 pages (64 MiB)contrib/codegen/codegen_wasm.mere— same bump for the self-host
codegen's own memory-line emission (16 → 1024)
test/test_basic.ml— updated the "wasm: memory declared + exported"
snapshot to expect 1024. run_wasm also now passes node --stack-size=65500 because self-host workloads recurse thousands of frames before returning (default Node stack ~500 KB).
Verified: examples/oneshot_codegen.mere (imports the self-host codegen and calls parse_and_emit "42") now runs end-to-end under Node, emits 80,744 bytes of WAT, exits cleanly. Previously trapped with either "call stack size exceeded" or "memory access out of bounds" depending on which limit hit first.
Deferred:
memory.growin the bump allocator. Bumping the default fixes the
common case but doesn't help workloads > 64 MiB. Growth-on-demand needs instrumentation at every bump-alloc site — invasive rewrite in lib/codegen_wasm.ml.
Follow-up (same day): codegen_runtime_bootstrap CI helper added in test/test_basic.ml. Compiles examples/oneshot_codegen.mere via the pre-built _build/default/bin/mere.exe (avoiding nested dune exec inside dune runtest), runs the wasm under Node with a puts hook that captures the auto-printed main result, and asserts the expected value (80746 bytes for parse_and_emit "42"). This closes the previously-deferred CI gap — regressions in the runtime self-host path now fail CI immediately.
dune runtest: 1778 → 1779 passing.
2026-07-02 — Phase 54.35 web backend Stage A (contrib/http)
First Node-hosted HTTP server bindings for Mere. Answers the question "can I write a real web backend in Mere today?" — yes.
Added:
contrib/http/http.mere— five extern fns:http_serve: int -> (str -> str) -> unit— register handler, start serverhttp_current_body: unit -> str— read POST/PUT bodyhttp_set_status: int -> unit— override response statushttp_set_content_type: str -> unit— overrideContent-Typehttp_set_header: str -> str -> unit— add arbitrary response header
contrib/http/http.glue.js— Node glue with per-request slots for
body / status / content-type / headers. Uses the same closure ABI as contrib/dom (Phase 48 C2 MVP): DataView-based {env, fn_idx} dispatch through the exported __indirect_function_table.
scripts/run_http_server.js— reference host that merges standard
env imports (puts, libc stubs, math) with the http glue.
- Four examples exercising the stack:
examples/http_echo_server.mere— minimal echo (~30 LoC)examples/http_echo_body.mere— POST body viahttp_current_bodyexamples/http_json_api.mere⭐ — six-endpoint JSON REST API with
CORS via http_set_header, 404s via http_set_status
examples/http_todo_api.mere⭐ — in-memory TODO CRUD with
routing, top-level mutable Map[str, str] state, POST / GET / PUT / DELETE + 404s on missing ids
- README entries in
contrib/README.mdandexamples/README.md - Detailed
contrib/http/README.mdwith API table, integration
recipe, and MVP limitations
Non-obvious gotcha caught in testing: http_current_body () returns a pointer into a per-request scratch buffer that gets overwritten at the start of the next request. Storing that pointer directly in a Map for later reads returns garbage. Fix: copy the bytes into the stable bump arena via strbuf before storing —
let buf = strbuf_new () in
let _ = strbuf_push buf (http_current_body ()) in
let text = strbuf_to_str buf in
map_set store id text
Documented in contrib/http/README.md.
MVP limitations (documented): Node-only host, no streaming / binary payloads, no custom request-header access, single scratch buffer shared across servers.
Position: Stage 2 contrib (incubation), sibling of contrib/dom on the server side. Graduation target is mere-http (separate repo) once the package manager lands. A future lower-level contrib/net (raw sockets over a C runtime) will slot in below this one.
2026-06-30 → 2026-07-01 — Phase 54 self-host bootstrap loop closes
Over 32 incremental slices (Phase 54.1 → 54.32) the Mere source of the compiler pipeline was made to compile itself. 1622 → 1771 tests. 17 contrib libraries are now self-host-compilable and go end-to-end through parse_and_emit_file → wat2wasm → node.
Milestones achieved:
- Compile-time self-compile loop closes:
codegen_wasm.mere(~2800
lines) compiles itself through parse_and_emit_file to 1,560,495 bytes of valid WAT; wat2wasm accepts the output. CI-verified.
- Runtime self-host of 5 major components:
lexer,parser,
evaluator, type inferencer, and formatter all compile via the self-host pipeline AND run correctly under wasm. Ten bootstrap harness tests exercise real workloads:
tokenize "let x = 1 in x"→ 7 tokensparse_decls (tokenize "let x = 1; let y = 2; let z = 3;")→ 3 declsparse_and_eval "let rec fact = fn n -> if n < 1 then 1 else n * fact (n - 1) in fact 5"→ 120parse_and_infer "let x = 5 in x + 1"→ "int"format_program (parse "1 + 2 * 3")→ "1 + 2 3\n"- 17 contribs self-host-compilable:
ast/lexer/parser/
typer / eval / fmt / json / path / option / regex / regex.engine / argparse / test / toml / markdown/to_html / markdown/to_text / markdown/toc. time.mere still needs float codegen. 10 of the 17 have bootstrap_wat_ok CI checks.
Key infrastructure added:
parse_and_emit_file path(Phase 54.10): recursiveimport "..."inline
with cycle detection + column-0 marker scan.
selfhost_prelude(Phase 54.9 + 54.11 + 54.27): auto-prepended Mere
source with list_map / list_rev / list_fold / list_len / list_append / list_mapi / list_filter / list_iter / list_any / list_all / str_join / str_split / str_trim / str_replace, plus type __list_t = Nil | Cons of int; / option / result so tags register deterministically.
- Constructor-arity rewrite (Phase 54.13): parser post-pass that walks
TopType decls, builds an arity map, and rewrites EApp(EConstr name None, x) → EConstr(name, Some x) when arity is 1 — fixes the Some x bare-app trap the atom-level parser can't disambiguate.
- Stdlib builtins in
codegen_wasm.mere:ord/chr/is_digit/
is_alpha / is_space / str_len / char_at / str_starts_with / substring / str_index_of / str_repeat / int_of_str / str_unescape / str_eq / strbuf_new / strbuf_push / strbuf_to_str / strbuf_len / map_new / map_set / map_get / map_has / read_file / not / fail; every one gets a WAT helper.
- Semantic fixes:
$char_atreturns a 1-byte str (matching OCaml
V_str), ==/!= on any EStr literal lower to $__lang_streq, and str_eq provides explicit content equality for two runtime strings.
- Parser extensions:
module M { }/extern fn/fn _/fn (a: t)/
cons-tail [h, ...t] / 'a tyvar / char literal / 'X' / tuple destructure shorthand / Module.Ctor in patterns and expressions / float literal skip (integer part only) / region R { <expr> } permissive.
Outstanding: runtime self-compile of the codegen itself (parse_and_emit running inside the compiled wasm) traps in an isolated 8-line region — a wasm-level bug that shows up specifically with 6+ character identifier names. Documented reproduction; needs interactive wasm memory inspection to close. Time.mere waits on proper float codegen.
2026-06-22 (cont. — Phase 38.G-1 OwnedVec auto scope-bound Drop)
After Phase 38.C finished, during the public-release prep session we consumed Level 1 of DEFERRED §1.3. 1515 → 1526 tests. Implements N1 of the N1/N2/N3 decomposition that was paper-validated in the design doc (39_nll_linear_design.md).
- Behavior: for
let v = owned_vec_new () in body, if static analysis
can confirm that body does not lexically escape v, we auto-emit free(v->data) at scope end (same shape as Phase 15.13 with).
- Static analysis (new helpers in codegen_c.ml):
no_value_leak v body: checks thatVar vdoes not appear in value
position of Tuple / Constr payload / Record_lit / Record_update / Fun body.
tail_does_not_return_v v body: checks that the tail expression's type
does not transitively contain OwnedVec.
- Both pass → auto-Drop; either fails → fall back to existing registry +
main-end sweep (safe-by-default, conservative).
- Supported backends: C + LLVM. Wasm uses bump-arena and has no
per-allocation free, so Phase 38.G-1 is a no-op there (will enable if GC / linear-memory free arrives).
- Escape patterns (no auto-Drop): tail of body returns
v/v
stashed in a tuple / closure captures v / tail type contains OwnedVec.
- Auto-Drop patterns: build → query → return scalar / each
ifarm is
scalar / nested let chains whose tail is scalar / compatible with Phase 38.C partial application.
- Levels 2/3 (N2 NLL Light, N3 Full Linear, ~5–15 slices) remain
deferred — held back until dogfood actually hurts.
- Relevant commit:
76f00f8
2026-06-22 (cont. — Phase 38.C multi-arg curried builtin first-class)
After Phase 37 finished, the public-release sprint consumed DEFERRED §1.2 A2. Multi-arg curried builtins now work in value / partial-app position on all 3 backends. 1511 → 1515 tests.
- Design call: the originally envisioned per-builtin × per-arity closure
adapter template (extension of Phase 35.1 nullary) was scrapped — boilerplate would explode as builtin × arity × backend. Instead each codegen got an AST-local synthesize helper (synthesize_curried_eta / _llvm / _wasm); the Var handler detects a multi-arg curried builtin in value position and synthesizes a fully eta-expanded fn __arg0 -> fn __arg1 -> ... -> builtin __arg0 ... __argN Fun chain on the spot, then re-feeds it to emit_expr. The existing anonymous-Fun adapter machinery (Phase 5.7-b) builds the closure; the nested inner App hits each builtin's direct-call fast path.
- Supported builtins (9):
owned_vec_push/owned_vec_get/
vec_push / vec_get / vec_set / strbuf_push / map_get / map_has / map_set.
- Examples:
let push_v = owned_vec_push v in let _ = push_v 1 in let _ = push_v 2 in ...
let set_in_m = map_set m in // 1-arg partial of a 3-arg let _ = set_in_m "a" 1 in ...
- Limitation: fully unapplied (
let push = owned_vec_push) becomes
polymorphic after let-poly, so the use site must pin the type with Annot or a concrete argument (same constraint as Phase 35 nullary).
- Slice layout:
46b2704Phase 38.C-1 spike (C / owned_vec_push) /
24ff513 38.C-2 (C / remaining 2-arg) / a6fb4bf 38.C-3 (C / 3-arg) / 8265992 38.C-4/5 (LLVM + Wasm port).
2026-06-22 (cont. — Phase 37 public-release prep)
A prep sprint to public-ize mere after Phase 36 syntactic sugar. LICENSE adopted + CI set up + B/A implementation polish complete. 1488 → 1498 tests.
- LICENSE (MIT alone):
LICENSE(MIT) +CONTRIBUTING.md, with a
contributor heads-up that we may go MIT OR Apache-2.0 dual in the future. Matches the mainstream license of OCaml-family languages (Lua / Zig / Julia / Nim / F#). Strategy notes are in internal design notes Section F.
- GitHub Actions CI: ubuntu + macos × OCaml 5.1/5.4 running
dune build
+ dune runtest. CI / License badges added to README.
- Phase 37.B exhaustiveness Phase 2:
is_total_patternrecurses into
tuple / record ((a, b) and { x = a, y = b } count as total), type hints attached to wildcard warnings for int / str / float / tuple / record ("no wildcard arm for int" etc.). 1488 → 1494 tests.
- Phase 37.A `while` at top-level (3 backends): extended C / LLVM / Wasm
lift_fn_skels so let _ = while cond do body; works directly under main. When Let (P_*, Let_rec (bs, lr_body), rest) is seen, bs is lifted to a top-level fn skel and the value is replaced with lr_body. 1494 → 1498 tests.
- Phase 37.C multi-arg curried builtin first-class: the remainder of
DEFERRED §1.2 A2. Re-estimated implementation size and deferred to Phase 38.C (closure-form for 2-arg curried builtins requires outer/inner adapter generation in two stages, with boilerplate piling up across 10+ builtins like vec_push / map_set × 3 backends).
- `.gitignore` / `.gitattributes`: ignore editor / OS / codegen output;
*.mere linguist-language=OCaml as interim highlighting until Linguist registration.
- CLI ergonomics polish:
--version/-vflag, explicit error for
unknown flags, help text updated to reflect 4-backend feature parity (dropped legacy "Phase N prep, int subset" wording), added pointer to docs / examples at the end of help.
- opam packaging:
(package mere)indune-project+(public_name
mere) in bin/dune. generate_opam_files true auto-generates mere.opam. opam install . works.
2026-06-22 (cont. — Phase 36 syntactic sugar + dogfood examples)
After Phase 32 (FFI), ran straight through Phase 33 (dogfood example batch + did-you-mean expansion), Phase 34 (float on 3 backends + libm dispatch), Phase 35 (DEFERRED §1.2 A1: nullary factory builtin first-class value), and Phase 36 (13 syntactic sugars + 16 prelude entries + 47 examples + 8 DEFERRED fixes). 1486 → 1488 tests, examples 61 → 118 (47 new), the syntactic surface reached practical territory for an ML-family language.
- Phase 36 sugars (13 kinds): range
a..b/ operator section(+ 1)/
cons 1 :: xs / reverse pipe f <| x / apply f @@ x / lambda shorthand \x -> ... / string interpolation "x = {show n}" (lexer re-tokenizes recursively, \{ to escape, nested strings rejected) / ? (Option early-return) / ?! (Result early-return) / list comprehension multi-gen [f x | x <- xs, p x] / if let pat = e then ... else ... / for x in xs do body (→ list_iter) / while cond do body (→ let rec __while_N = fn () -> if cond then body; __while_N () in __while_N ()).
- Phase 36 prelude (16 entries):
range/list_filter/list_take/
list_drop / list_find / list_append / list_concat / list_flat_map / list_zip / list_for_all / list_any / list_member / list_sum / list_product / list_max / list_min (cumulative 34 entries). sum / product / max / min are defined with let rec (looks complex because the test helper codegen_with_decls skips Top_let_rec).
- Phase 36 DEFERRED fixes (8): §1.13 narrowed value restriction (do
not generalize types containing mutable containers) / §1.14 lifted closure capture goes through load / global.get for globals / §1.15 C codegen O(2^N) slowdown on deep list literals (double emit_expr arg inside Constr → cache once) / §1.16 strbuf_to_str inside a region had dangling pointer on region escape (C/LLVM switched to __lang_default_region alloc) / §1.17 C codegen type result shadow blew up List.combine (remove from polymorphic_variants + dedupe variant_decls last-wins) / §1.18 Phase 30.2 top-level global init order (source-order inline init) / §1.19 nested lambda unbound on top-level fn reference (added closure_wrapper_forward_decls in C/LLVM/Wasm; Wasm populates fn_closure_table_idx before emit_fn_def) / §1.20 C codegen forward decl for user record inside polymorphic variant (include mono variant/record bodies in the unified topo sort).
- Phase 36 examples (47): basic dogfood (histogram / traffic_light /
event_counter / html_builder / fallible_lookup / config_loader / csv_writer / markdown_to_text / calendar_lite / matrix_2d / borrow_chain / cache_sim / simple_query / caesar_cipher / fraction / roman_numerals / password_strength / brackets_balance / morse_code / luhn_check / tic_tac_toe / palindrome / anagram / base_conv / rps_game / scoreboard / eight_queens / collatz / bin_tree_traversal / knapsack / factory_value) + sugar showcase (range_demo / sections / cons_pipe_demo / sugar_demo / question_demo / sugar_showcase / comprehension / statistics / if_let_demo / for_loop_demo / while_loop_demo) + 4 big ones (csv_summary / game_of_life / sudoku_check / calc 138 lines / maze_solver BFS).
- Phase 35: extended DEFERRED §1.2 A1 (first-class factory builtin
eta-wrap) to all 3 backends. Added eta_adapters to C/LLVM/Wasm so that unapplied builtins like let mk = map_new work correctly as values.
- Phase 34: float MVP rolled out to 3 backends. Phase 34.1 = C,
Phase 34.2 = LLVM (fadd / fsub / fcmp + @llvm.fabs.f64 + __lang_str_of_float), Phase 34.3 = Wasm (i32 ptr to heap-alloc f64 slot + host import for formatting), Phase 34.4/34.5 = libm dispatch (sqrt/sin/cos/tan/f_pow/atan2) on 3 backends + math_demo example.
- Phase 33: dogfood example batch + did-you-mean expansion. Phase
33.0 expanded did-you-mean to multi-candidate top-3 listing (partially closes DEFERRED §5.1). Phases 33.1–33.7 added D3 option_pipeline / H1 prime_sieve / G5 rate_limiter / C4 stack_calc / G6 markdown_toc / G4 bank_account / H3 graph_bfs working with diff = 0 on 4 backends.
2026-06-22 (cont. — Phase 32 C1 FFI)
Right after Phase 31, ran Outlook §C1 (FFI = calling external C functions) through 5 slices + 1 polish back-to-back. 1480 → 1486 tests, the extern fn <name>: <ty>; syntax lets libc functions be called directly from all 4 backends. A step that takes Mere from "an experimental language that runs by itself" to "a practical language that can talk to the outside world".
- Phase 32.6: multi-arg curried extern (
extern fn setenv: str -> str
-> int -> int;) working on 3 backends. The collect_extern helper walks the App chain to gather all args. Added default JS impls for getenv / setenv / system to scripts/run_wasm.js. Added a 3-arg setenv example in examples/ffi_demo.mere; diff = 0 on 4 backends.
- Phase 32.5: added 4 + 2 tests for §32.1–32.4 + §32.6 (1484 → 1486),
created examples/ffi_demo.mere.
- Phase 32.4: Wasm codegen emits
(import "env" <name> ...)host
import + call $<name>; default JS impls for getpid/getppid etc. injected into scripts/run_wasm.js.
- Phase 32.3: LLVM codegen emits
declare <ret> @<name>(<args>)+ call. - Phase 32.2: C codegen emits
extern <ret> <name>(<args>);decl +
direct call. unit arg → (); unit return → (call, 0) for int-ification.
- Phase 32.1: lexer (T_extern) + AST (Top_extern) + parser + typer +
pipeline + repl + bin + 9 mocks via lookup_extern in eval.ml (getpid / getppid / getenv / setenv / system / sleep / srand / rand / unix_time).
- Phase 32.0: FFI design — fixed syntax / typing /
ABI / per-backend strategy. MVP type range is int / bool / str / unit only; float / tuple / record / variant / callback deferred.
2026-06-22
Ran 11 slices of Phase 29-31 across the night. Starting from 16 examples PERFECT on 4 backends, finished dogfood (toy_sql 1165 LoC) → bug hunt → all fixes → README polish in one day. 1469 → 1480 tests; DEFERRED §1.10 / §1.11 / §1.12 fully resolved; mere reached a state presentable to outsiders.
- Phase 31.1: README updated to reflect Phase 22-31 (1268 → 1480 tests;
3 → 4 backend feature parity; toy_sql 1165 LoC; signature spread / Result helpers / inner-fn lifting / top-level globalization / Wasm runtime execution / str_compare on 3 backends).
- Phase 31.0: ported
str_compareto 3 backends (C / LLVM / Wasm).
Sign-normalized to match interp's OCaml compare s t (-1/0/1) exactly. C uses inline strcmp, LLVM uses strcmp + select, Wasm uses a dedicated runtime helper.
- Phase 30.2c ⭐: Wasm codegen declares non-fn top-level lets as
(global $name (mut i32)), initializes them with global.set $name at main entry. Var emits global.get $name. Works uniformly since all values are i32.
- Phase 30.2b: LLVM codegen declares them as
@<name> = internal
global <ll_type> zeroinitializer, stores init at main entry, Var reference is load.
- Phase 30.2a: C codegen declares non-fn top-level lets as file-scope
static <type> <name>;, initializes at main entry. The heuristic only globalizes lets whose name shows up in skels' free_vars, protecting existing tests. DEFERRED §1.10 fully resolved on all 3 backends.
- Phase 30.1 ⭐: when a captured name in a closure was shadowed by
let, body emission now temporarily removes the shadowed name from current_env_subst. Root cause was not specific to P_tuple — it was env_subst not respecting shadowing. Applied to both Let P_var and Let P_tuple. DEFERRED §1.11 fully resolved.
- Phase 30.0 ⭐: added
when not (Hashtbl.mem toplevel_fn_names ...)
guard to the hardcoded dispatch of builtins (is_alpha / is_digit / is_space). If a user-defined fn shadows them, builtin dispatch is skipped. Same pattern applied to C / LLVM / Wasm. DEFERRED §1.12 fully resolved.
- Phase 29.3 ⭐: implemented nested-loop JOIN in toy_sql + qualify_row
+ project_join + 7 JOIN tests. toy_sql total 1165 LoC, diff = 0 PERFECT on 4 backends, 59 tests (tokenizer 22 + parser 13 + executor 17 + JOIN 7). Final assessment of N1/N2/N3 dogfood: at 1165 LoC the demand never materialized; pain concentrated in codegen plumbing (DEFERRED §1.10–§1.12).
- Phase 29.2: toy_sql executor (Catalog Map[str, table_meta] +
Storage OwnedVec[tagged_row] + WHERE filter + project + 17 tests). Map[K, V=variant] and OwnedVec[variant] codegen worked first try (symmetric to Phase 15.16).
- Phase 29.1: toy_sql SQL parser (AST + continuation flow + 13 tests).
Dogfood findings: C codegen tuple destructure rebind bug (DEFERRED §1.11), Wasm memory expanded from 1 page (64KB) to 16 pages (1MB) for string-heavy apps.
- Phase 29.0: toy_sql foundation (Value variant + Token variant +
hand-written tokenizer + 22 self-tests). Dogfood findings: C codegen record-field × nested-lambda capture bug (DEFERRED §1.10), C codegen shadowing user-defined fn with builtin (DEFERRED §1.12).
2026-06-21
After closing one deferred item in Phase 21, ran Phase 22 → 23 → Phase 24-27 (29 slices straight) to complete 4-backend feature parity, then added 4 dogfood examples in Phase 28. 1268 → 1469 tests passing, DEFERRED §1.7 / §1.8 / §1.9 resolved, 16 examples match diff = 0 PERFECT on all 4 backends.
- Phase 28.1: fix deep nested lambda capture bug in C codegen
(DEFERRED §1.9). Added pattern_vars_with_types helper; Match emit_arms wraps arm body / guard in with_pat scope and prepends pattern bindings to current_var_types. Nested closures in arm bodies now pick up pattern-bound names in free_vars filter and write them into closure env. Same shape as LLVM Phase 25.3 (second N+1 → N backport).
- Phase 28.0: 4 new examples verified on 4 backends:
- D2
chained_parse.mere: Result chain idiom (result_and_then /
- D2
result_map / result_or_else)
- C1
state_machine.mere: variant + match transitions - I1
ini_parser.mere: line parser + Map (Phase 27.1 insertion-order
dogfood)
- C5
regex_lite.mere: recursive AST + backtracking matcher
12 → 16 examples PERFECT-matching on 4 backends. chained_parse surfaced C codegen undeclared identifier 'rest' (DEFERRED §1.9).
- Phase 27.3 ⭐: Wasm ty_tag accepts StrBuf (releases blocker where
Phase 15.9-implemented mere_strbuf_* runtime couldn't be used with StrBuf inside tuple/variant payload). json_writer matches PERFECT on Wasm runtime → 12/12 PERFECT on Wasm → full 4-backend feature parity achieved.
- Phase 27.2 ⭐: Wasm runtime execution verification. Added
scripts/run_wasm.js (Node.js host harness with puts / read_file / write_file imports). Wasm main tail emits show_<main_ty> + puts; add_show_type main_ty forces show emission for main_ty. 11/11 examples match PERFECT vs interp on Wasm runtime.
- Phase 27.1 ⭐: pinned interp Map iter order to insertion order.
V_map changed to (Hashtbl, value list ref); map_set appends new keys; map_iter iterates via the list. All 3 backends now 12/12 PERFECT (C/LLVM 10 → 12; word_freq + mini_shell Map-order cosmetic diff gone).
- Phase 27.0: C codegen prints
"()"for unit main_ty (backport of
LLVM Phase 25.11). template_engine / json_writer / inventory / cap_handler no longer trail () on C; C PERFECT 6 → 10.
- Phase 26 (7 slices): 11/12 examples EMIT + wat2wasm successful on
Wasm codegen. Ported the cumulative Phase 22-25 features (variant boxed payload / stdlib builtins / try_or / inner let-rec lifting / multi-instantiation specialization / str_split / str_join / read_file / write_file / lift_fn_skels non-Fun walk / various polishing) to Wasm one slice at a time.
- Phase 25 (13 slices): LLVM codegen runs 12/12 examples (PERFECT 10).
In parallel with Phase 24.x C features, implemented boxed payload / stdlib / try_or / inner let-rec lifting / multi-instantiation specialization / show_str escape / fn dedup / nested P_constr / missing builtins / various polishing on LLVM side.
- Phase 24 (5 slices): 12/12 examples working on C codegen
(template_engine / json_writer / inventory / cap_handler / word_freq / mini_shell). Variant payload switched to { tag, payload_ptr } boxed representation, unifying polymorphic variant containers across all 3 backends.
- Phase 23 (5 slices): json_parser matches interp 100% on C codegen
(Phase 23.2 added result_map / result_and_then / result_or_else to prelude; Phase 23.3 per-instantiation specialization of polymorphic user let-rec; Phase 23.5 show_str escape — DEFERRED §1.7 fully resolved).
- Phase 22 (5 slices): try_or + str ops (str_split / str_join /
str_count / str_index_of) working on all backends.
- Phase 21 (1 slice): partial resolution of DEFERRED §1.7 (first
stage of polymorphic user let-rec monomorphization on C codegen).
2026-06-20
Started from Phase 15.16, then sprinted through Phase 16 / 17 / 18 in one day. 1268 → 1304 tests, resolved 6 items: DEFERRED §1.4 / §1.5 / §1.6 / §2.1 / §2.5 / §4.1. Reached a state with 4 backends matching exactly on a non-trivial program (todo_app), full coverage of the 10-pair borrow checker conflict matrix, and proper module scoping (M.Red qualified + open A.B; nested paths).
- Phase 18.2: `open A.B;` (open on nested module path) — DEFERRED
§4.1 fully closed. module_bindings registers under both short-name key and full-path key (A.B); parser's T_open refactored to a path parser. Existing open M; follows the same code path (1304 passing).
- Phase 18.1: M-prefix scoping for ctors / records inside modules —
remainder of DEFERRED §4.1. After module M { type T = Red | Blue; }, qualified access M.Red, qualified record literal M.Pt { ... }, and qualified patterns match v with | M.Red -> ... all work. Same-named ctors across two modules can be disambiguated by qualified form. Loose coupling: new AST decls Top_ctor_alias / Top_record_alias + shared alias table (Ast.ctor_aliases) + typer.alias_ctor + eval normalizes to canonical name when constructing V_constr. Bare names still work for backward compat (1301 passing).
- Phase 17.2: full 10-pair borrow conflict matrix + intra-tuple
conflict — resolves DEFERRED §2.5. Of the 4×4=10 conflict pairs, added tests for the 4 untested ones (SW×ER, SW×EW, ER×ER, ER×EW); changed `check_borrows` Tuple branch to sequential threading; added a "Conflict matrix and extension history" section to design doc 08 (1295 passing).
- Phase 17.1: track function-return borrow by let-bound name —
DEFERRED §2.1 fully resolved. For let r = f x in let r2 = &mut R r where f returns &R T, the let-bound name is used as a place and a synthetic borrow is added to active for conflict detection (1287 passing).
- Phase 16 polish: reflected friction points #1/#2/#3/#4 in tutorial
/ patterns ({ t | f = v } partial update, same-name rebinding, type annotation idiom for closure parameters). Phase 16 retrospective document created.
- Phase 16.4: Wasm Region_block bump restore removed — DEFERRED
§1.6. Fixed bug where let v = region R { vec_to_owned ... } in ... allocates inside a region and escapes, but the region exit rewinds bump so subsequent allocations overwrite the escaped value. Aligned Wasm region semantics with arena-leak (1283 passing).
- Phase 16.3: mk_logger / mk_metrics codegen on 3 backends —
DEFERRED §1.5. Brought interpreter-only Logger / Metrics cap builtins to C / LLVM / Wasm parity. Logger = { closure_str_unit info / warn / error }; Metrics = { inc, record (curried str→int→unit) }. Side change: collect_arrow_types (C/LLVM) recursively traverses known record field types → closure typedefs used only via Logger are also auto-emitted (1281 passing).
- Phase 16.2: fix C codegen `let x = f x` same-name rebinding bug —
DEFERRED §1.4. __auto_type x = ...x... hits the C rule "a variable may not reference itself in its initializer" and triggers a clang error. codegen_c.ml Let uniformly expanded to 2-step form ({ __auto_type __let_tmp_<name> = <value>; __auto_type <name> = __let_tmp_<name>; <body>; }); at rhs evaluation the new binding is not yet declared so the old binding is visible (1269 passing).
- Phase 16.1: surface 6 friction points via practical example
todo_app.mere — 110-line TODO app combining OwnedVec[Task] + Logger + vec_map + region. Documented 2 by-design (#1/#2 immutable record update), 2 HM limits (#3/#4 field access inference), 2 real bugs (#5 rebinding, #6 mk_logger codegen), 1 Wasm bug (§1.6) (1268 passing).
- Phase 15 #16: extended Map[R, K, V] K to payload-bearing variants
across 3 backends (Mere's full concrete type set is now usable as a Map key).
2026-06-19
- Phase 15 #16: extended Map[R, K, V] K to payload-bearing variants on 3
backends — extends Phase 15.15 nullary-variant K to also accept ctors carrying payloads. Now Mere's full concrete type set works as Map key. (a) C codegen: extended the variant branch of `key_eq_for` — `(a.tag == b.tag) && (a.tag == TAG_X ? eq_payload_X : a.tag == TAG_Y ? eq_payload_Y : ... : 0)` nested ternaries for per-tag dispatch; nullary ctors short-circuit to `1` (true). Payload recursively calls `key_eq_for`. C codegen accepts different payload types across ctors (leveraging variant's union representation). (b) LLVM IR: extended `emit_map_key_eq_helper_llvm` variant branch — extract tag with `extractvalue`, 0 if tags differ, otherwise extract payload and compare. LLVM MVP restriction: ctors must share the same payload type (MVP variant codegen requires a single payload type). Layered OR of "tag-in-nullary-set" checks for nullary ctors, combined with payload eq. (c) Wasm: extended `emit_map_key_eq_wasm` variant branch — load tag with `i32.load offset=0`, then a nested if/else chain `if (tag == TAG_X) then eq_payload_X else ...`. Last else is `1` (nullary or covered). Wasm also assumes uniform payload type under MVP, like LLVM. `is_key_supported` accepts payload variants on each of the 3 backends, recursively checking payload types. Added 5 tests (1268 passing) — C accepts mixed payload (A int / B str), LLVM/Wasm accept uniform payload (A int / B int / C nullary) + interpreter parity (1502, 603). Side test-helper refactor: changed `vec_codegen_c` / `_llvm` / `_wasm` test helpers to go through `typed_prog` and `Pipeline.process_decls` so Top_type etc. are registered first (programs with type decls used to typer-error in test helpers). Mere's Map key support now covers all concrete types (int / bool / str / tuple / record / nullary variant / payload variant). Remaining: first-class value usage; auto-Drop.
- Phase 15 #15: extended Map[R, K, V] K to record / nullary variant on 3
backends — extends Phase 15.14 (tuple) so records and nullary variants also work as K. Enables meaningful maps with compound keys (e.g. `Pt { x, y } → value`, `Color = Red | Green | Blue → value`). Payload-bearing variants out of scope (per-tag union access is complex, candidate for separate slice). (a) C codegen: extended `key_eq_for` — records use `(a).field_name` for direct field access and recursive compare; nullary variants compare tags only with `(a).tag == (b).tag`. `is_key_supported` allows record / variant in both spots (Map type registration and `map_kv_tags_of`); judgment via `Typer.records` / `Exhaustive.type_variants`. (b) LLVM IR: inside `emit_map_key_eq_helper_llvm` `go` function, records get field via `extractvalue %RecName %r, i`; nullary variants get tag via `extractvalue %VarName %v, 0` + `icmp eq i32`. (c) Wasm: in `emit_map_key_eq_wasm` `build`, records get field via `i32.load offset=4*i` (memory-offset based); nullary variants get tag via `i32.load offset=0` + `i32.eq`. Error messages updated to "int / bool / str / tuple / record / nullary variant". Added 8 tests (1263 passing) — 3 backends × (variant key Color: 9, record key Pt: 1000) accept + interpreter parity. Payload-bearing variants still rejected (DEFERRED §1.1 separately).
- Phase 15 #14: extended Map[R, K, V] K to bool / tuple on 3
backends — extends Phase 15.10 (which had int / str only) to also accept bool / tuple (recursively). Enables compound keys (e.g. coordinates `(x, y) → ...`) with tuples. Key equality expands recursively per K structure. (a) C codegen: refactored `key_eq_expr` into recursive `key_eq_for k a b` — int/bool via `==`, str via `strcmp`, tuples access each field via `(a).f0, (a).f1, ...` and AND them. Tuples are C value types (struct), so direct field access works. (b) LLVM IR: emit one `@mere_map_key_eq_<K>` helper per K (called from `map_set / get / has`). Tuples are decomposed via `extractvalue` and recursively combined with `icmp eq + and i1`. `map_instances` is iterated for unique K and a helper is emitted per unique K in emit_program. (c) Wasm: all values are i32 but tuples access fields via memory offset. Added new `emit_map_runtime_wasm k_ty` function that generates 5 helpers per K (new/set/get/has/len) + `$mere_map_key_eq_<K>`. Phase 15.10 hardcoded `map_int_runtime_wasm` / `map_str_runtime_wasm` removed; `map_key_types : (string, Ast.ty) Hashtbl.t` registers K types → emit_program iterates. Tuple key equality in WAT uses block-scoped local.set + i32.load offset=4*i + recursive call_eq. Added 8 tests (1255 passing) — 3 backends × (bool, tuple key) accept + interpreter parity (bool: 302, tuple: 121). Remaining: extending Map K to record / variant (per-K eq logic is generic so extension is easy, but a separate slice is cleaner).
- Phase 15 #13: scope-bound OwnedVec Drop via `with v = owned_vec_new
() in body** — complements Phase 15.8 process-wide registry (__mere_owned_vec_free_all at main end) by wiring OwnedVec into the with syntax. When written explicitly as with v = owned_vec_new () in body, after body evaluation v->data is freed and the struct's data field is rewritten to NULL. The registry's free_all (at main end) tolerates free(NULL) (C standard no-op) while finally freeing the struct itself. Fits Mere's **"explicit > concise" philosophy** — the user opts into scope-Drop only when needed, safe without Rust-like move semantics or ownership analysis (creating an alias inside with and using it outside is still UB, but typer's Drop-type rule suppresses some of it). **(a) C codegen**: added branch to Ast.With (name, value, body) emission for value.ty = OwnedVec, inserting free(((__mere_owned_vec_base)name)->data); ((__mere_owned_vec_base)name)->data = NULL; after body. The __mere_owned_vec_base is the existing registry { void data; int len; int cap; }` struct — generic free leveraging that all `mere_owned_vec_<T>` share the same leading layout. (b) LLVM IR: emit `getelementptr {ptr, i32, i32}, ptr v, i32 0, i32 0` to access struct field 0 (data ptr), then `load → @free → store null`. LLVM's opaque pointers + shared leading layout means it works without type tags. (c) Wasm: no malloc/free, just a linear-memory bump allocator, so structurally a no-op (process exit collects). No code change, but extended `resolve_vec_let_types` pre-pass to also walk With so typer type info flows correctly (shared across 3 backends). Added 3 tests (1247 passing) — C/LLVM scope-end free emission + interpreter parity (30). Remaining: scope-bound Drop is only on explicit `with`; default `let` still relies on main-end registry sweep. Rust-style auto-Drop requires NLL + move semantics (DEFERRED §1.1).
- Phase 15 #12: added `vec_to_list` + `len` on list to 3 backends —
added the remaining recursive-variant (Nil/Cons chain) construction + traversal in codegen. Parallel to Phase 15.7 vec_to_owned, vec_to_list v converts region Vec to T list (builds Cons chain bottom-up — start from Nil and prepend in reverse). len on list added; other types covered in Phase 15.11. (a) C codegen: vec_to_list inline-expanded in GCC stmt expression, calling mere_vec_<T>_get(v, i) in reverse and writing each into Cons payload tuple_<T>_list_<T> (.f0 = elem, .f1 = acc); new nodes allocated from default region. Cons/Nil tag values resolved at codegen time from variant_tags. Len on list inlined similarly (while loop with __l->tag == cons_tag condition, __l->payload.Cons.f1 for next). (b) LLVM IR: per-T helpers @mere_vec_to_list_<T> and @mere_list_<T>_len, with phi for loop counter (i / acc) and list cursor. Assumes %list_<T>_node = type { i32, %tuple_<T>_list_<T> } exists and accesses payload via getelementptr. vec_to_list_instances : (string, Ast.ty * Ast.ty) Hashtbl.t tracks per-T, deduped in emit_program. (c) Wasm: shared $mere_vec_to_list and $mere_list_len helpers (Wasm values are all i32 and list structure is uniform). Tag values pulled from variant_tags at codegen time and baked into runtime; vec_to_list_used / list_len_used flags for lazy emit. Added 7 example tests (1244 passing) — 3 backends × (vec_to_list / len-on-list) + interpreter parity. v2l_src program: type 'a list = Nil | Cons of 'a * 'a list; ... vec_to_list v ... computing len l + head; 13 on 3 backends + interp. Remaining: Map K extension (tuple / record / variant key); first-class value usage (let f = vec_new in ...); OwnedVec scope-bound Drop.
- Phase 15 #11: 3 backends got `len` ad-hoc polymorphic builtin
codegen — `len : 'a -> int` had runtime dispatch in the interpreter; codegen now uses compile-time dispatch (statically routes to the corresponding `_len` helper based on arg.ty). (a) C codegen: in the `Ast.Var "len"` App handler, walk `arg.ty` for dispatch — `Vec[_, T]` → `mere_vec_<T>_len`, `OwnedVec[T]` → `mere_owned_vec_<T>_len`, `StrBuf` → `mere_strbuf_len`, `Map[_, K, V]` → `mere_map_<K>_<V>_len`, `str` → `((int)strlen(...))`, `TyTuple ts` → static arity constant (`({ (void)(arg); N; })` evaluates side effects). (b) LLVM IR: same pattern; emit `call i32 @mere_vec_<T>_len(ptr %a)` etc. via fresh_reg; str via `@strlen → trunc i64 to i32`; tuple evaluates side effects via emit_expr then returns as constant register via string_of_int. (c) Wasm: Vec / OwnedVec share `$mere_vec_len` (same struct layout in Wasm); StrBuf / Map use their helpers; str via `$__lang_strlen`; tuple via emit_expr + `drop` + `i32.const N`. On each backend, `len` is removed from Var rejection — only first-class value usage is rejected. `len` dispatch depends on arg's static type; if arg is polymorphic like `Vec[__heap, 'a]` the existing `resolve_vec_let_types` pre-pass concretizes it (collection-type support since Phase 15.2). Added 5 tests (1237 passing) — 3 backends × (Vec / str / tuple) dispatch + interpreter parity (vec[3] + "hello"[5] + (1,2,3,4)[4] = 12). Remaining: `vec_to_list` (recursive variant codegen); Map K extension; first-class value usage.
- Phase 15 #10: 3 backends got `Map[R, K, V]` codegen — brought
the region-aware mutable hashmap to 3 backend parity. Scope: K = int / str + V = any concrete type, linear scan (O(n) lookup), on cap-hit allocate new array in region (arena semantics). Brings the 5 interpreter builtins from Phase 12.8 (map_new / map_set / map_get / map_has / map_len) to codegen. (a) C codegen: per-(K, V) mere_map_<K>_<V> struct { K* keys; V* values; int len; int cap; __lang_region* region; } + 5 helpers. Key compare via == (int) or strcmp(...) == 0 (str); set linear-scans for existing key and overwrites value, else appends to tail (on cap-hit, doubles array, memcpy to new region area). (b) LLVM IR: per-(K, V) %mere_map_<K>_<V> = type { ptr, ptr, i32, i32, ptr } + 5 helpers. SSA phi for scan loop; key compare via icmp eq i32 (int) or @strcmp (str). Grow path uses getelementptr ... null, i32 1 → ptrtoint for sizeof(K) / sizeof(V), then @memcpy to migrate parallel arrays. get/has return abort / ret i1 0 from not_found label. (c) Wasm: all values are i32, so per-K only (per-V not needed). 2 sets $mere_map_int_* and $mere_map_str_* (5 fns each); key compare via i32.eq or $__lang_streq. map_int_used / map_str_used flags for lazy emit — only one runtime is emitted if only one K is used. On each backend the App handler unwraps curried Apps; map_new's region pulled from e.ty TyRef marker (same pattern as Vec / StrBuf). Rewrote existing "map: codegen rejection (C)" test to accept; added 3 backends × (str/int) accept + interpreter parity, 8 tests total (1232 passing). Added examples/map_codegen.mere (str→int / int→str / Map inside region combined to return 640; interpreter + 3 backends all 640). Remaining: vec_to_list / len / first-class value usage.
- Phase 15 #9: 3 backends got `StrBuf[R]` codegen — brought the
region-internal mutable string buffer to 3-backend parity. StrBuf is a single non-polymorphic type (no element-type parameter), so per-T monomorphization is not needed; a single runtime helper set (new / push / to_str / len) suffices. (a) C codegen: mere_strbuf struct { char* data; int len; int cap; __lang_region* region; } + 4 helpers; push's realloc within same region (arena semantics); to_str copies null-terminated to region. strbuf_used : bool ref flag for lazy emit (zero overhead in programs that don't use it); added forward typedef. (b) LLVM IR: %mere_strbuf = type { ptr, i32, i32, ptr } + 4 helpers; push calls @__lang_region_alloc + @memcpy; push's resize loop is br-back form (double cap until enough capacity); to_str allocates len+1 bytes + memcpy + null terminator. (c) Wasm: $mere_strbuf_new / push / to_str / len added as an independent runtime block (no closure dispatch, separated from vec_higher_order). $__lang_bump shared; strings copied byte by byte with i8 store/load; resize-time memcpy also hand-written loop. On each backend, App handler unwraps curried form App ({ Var "strbuf_push" }, sb); strbuf_new's region pulled from e.ty TyRef marker (same pattern as Vec). Rewrote "strbuf: codegen rejection (C)" to accept; added 3 backends × accept + interpreter parity, 4 tests total (1225 passing). Added examples/strbuf_codegen.mere (interpreter + 3 backends return 48: len of "hello, world!" + len of string built in another region + sb1 len). Remaining: Map[R, K, V] / vec_to_list / len / first-class value usage.
- Phase 15 #8: main-end batch free for OwnedVec (naive Drop) —
replaces the "leave it to process exit" approach of Phase 15.7 with explicit "batch free at end of main" for heap-allocated OwnedVec. Clean under valgrind / leak sanitizer. Design: all mere_owned_vec_<T> structs share the leading layout { T* data; int len; int cap; }, so generic free works by casting the first field as void* data (free(v->data); free(v);). A process-wide registry (void** items; int count; int cap;) is a file-scope global; each _new helper registers the struct ptr, then main end's __mere_owned_vec_free_all iterates and frees all. (a) C codegen: added owned_vec_registry_runtime block (__mere_owned_vec_register / __mere_owned_vec_free_all + 3 file-scope globals); emit_owned_vec_runtime_for calls __mere_owned_vec_register(v) at end of _new; main end calls __mere_owned_vec_free_all() (only when ≥1 OwnedVec is present). (b) LLVM IR: emit owned_vec_registry_runtime_llvm equivalently; registry expressed via global ptr / i32; @realloc to grow; free_all iterates via phi loop. Each @mere_owned_vec_<T>_new end calls @__mere_owned_vec_register; @main end calls @__mere_owned_vec_free_all. (c) Wasm: no malloc, allocation via $__lang_bump (linear memory); process exit hands the entire WebAssembly instance back to OS, so explicit free is unnecessary / impossible — registry / free_all not emitted (preserves current behavior). Remaining limit: process-wide, not scope-bound, so memory grows monotonically for long-running programs that create many OwnedVecs. Real scope-Drop with NLL / move semantics is future work. Added 4 tests (1222 passing) — C / LLVM assertContains for registry + free_all calls; Wasm negative test confirms no registry emitted.
- Phase 15 #7: 3 backends got `OwnedVec[T]` + `vec_to_owned` /
owned_vec_to_vec — brought interpreter-only heap-allocated OwnedVec to 3-backend parity, including round-trip (deep copy) with region Vec. Drop processing omitted in this minimum scope (process exit collects). (a) C codegen: generates per-T `mere_owned_vec_<tag>` struct + 4 helpers (new/push/get/len) via `emit_owned_vec_runtime_for`; allocates with `malloc / realloc`. vec_to_owned / owned_vec_to_vec inlined in GCC stmt expression; the latter extracts the target region from e.ty TyRef marker (active region). `c_type_of` walks `OwnedVec[T]` → `mere_owned_vec_<tag>*` in parallel with Vec; forward typedefs added. (b) LLVM IR: per-T `%mere_owned_vec_<tag> = type { ptr, i32, i32 }` + 4 helpers; `getelementptr ... null, i32 1 → ptrtoint` for sizeof(T); push's realloc uses declared `@realloc(ptr, i64)`. Conversion helpers per-T `@mere_vec_to_owned_<tag>` / `@mere_owned_vec_to_vec_<tag>` implemented with SSA phi loops. (c) Wasm: values are all i32 and `$__lang_bump` is shared, so OwnedVec runtime is physically the same as Vec — owned_vec_new / push / get / len thin-alias-routed to `$mere_vec_*`; conversions use newly added `$mere_vec_clone` helper for deep copy (allocate new vec, loop element-push). Wasm owned_vec only retains drop_types' region-placement rejection; runtime representation distinction not needed. Extended `resolve_vec_let_types` pre-pass to also handle `Ast.TyCon ("OwnedVec", _)` on C / LLVM. Added `examples/owned_vec_codegen.mere` — vec → owned → vec round trip + fold returning 67 (interpreter + 3 backends all 67). Added 12 tests (1218 passing) — 3 backends × (owned_vec / vec_to_owned / owned_vec_to_vec) codegen-symbol emit + 3 interpreter parity. Remaining: real Drop (per-instance free); `vec_to_list` (recursive variant construction); `StrBuf` / `Map` / `len` / first-class value usage.
- Phase 15 #6: 3 backends got `vec_map` / `vec_filter` — all 5 main
Vec higher-order APIs are present — follows Phase 15.5 (vec_set / iter / fold) with the two region-preserving ones. Both APIs build a new Vec in the same region as the input (vec_map converts element type T → U; vec_filter keeps only elements where predicate is true). (a) C codegen: GCC/Clang stmt expression inlining; pull the original Vec's region from `__vc->region` to create new Vec via `mere_vec_<U>_new(__vc->region)`; expand closure dispatch in-line into a loop. vec_filter uses `__auto_type __x = mere_vec_<T>_get(...)` (compiler infers C type) and conditionally pushes via `mere_vec_<T>_push` based on predicate's if branch. (b) LLVM IR: vec_map per-(T, U) helper (`@mere_vec_<T>_map_<U>`); vec_filter per-T helper (`@mere_vec_<T>_filter`). Both pull the input Vec's region field (offset 12 = idx 3) via `getelementptr + load` and call corresponding `@mere_vec_<U>_new` / `@mere_vec_<T>_new` to make new Vec. phi manages loop counter; vec_filter conditional-pushes via `br i1` on predicate's i1. `vec_map_instances` / `vec_filter_instances` tables dedupe. (c) Wasm: all values are i32, so `$mere_vec_map` / `$mere_vec_filter` added to `vec_higher_order_runtime`. Both call `$mere_vec_new` (no region parameter in Wasm); apply closure to elements via `call_indirect`; push to new Vec via `call $mere_vec_push`. Added `examples/vec_map_filter_codegen.mere` (interpreter + 3 backends return 226). Added 9 tests (1206 passing) — 3 backends × (vec_map / vec_filter) codegen-symbol emit + LLVM's per-(T, U) per-T branch confirmation + interpreter parity. Now all 5 main Vec higher-order APIs (set / iter / fold / map / filter) work on 3 backends, with almost no gap to the interpreter. Remaining: `vec_to_list` / `vec_to_owned` / `OwnedVec` / `StrBuf` / `Map` / first-class value usage.
- Phase 15 #5: 3 backends got Vec higher-order APIs (`vec_set` /
vec_iter / vec_fold) — Vec[R, T] working on 3 backends since Phase 15.2 / 15.3 / 15.4; this slice brings interpreter-only main higher-order APIs to parity. (a) C codegen: vec_set is a per-T runtime helper (`mere_vec_<T>_set`); vec_iter / vec_fold are inlined at call site (GCC/Clang stmt expression `({ ... })` writes local + for loop + closure dispatch directly). Side bug fix: anonymous Fun in main_body wasn't draining closure adapter — added `drain ()` after `let main_body = emit_expr body_expr in` in emit_program to re-collect `pending_closures`. (b) LLVM IR: vec_set is per-T helper; vec_iter is per-T helper (`@mere_vec_<T>_iter`); vec_fold is per-(T, U) helper (`@mere_vec_<T>_fold_<U>`). Hand-written SSA with basic blocks managing loop state (i, acc) via phi. (c) Wasm: all values are i32, so all 3 helpers shared single runtime (`$mere_vec_set / $mere_vec_iter / $mere_vec_fold`). `vec_iter / vec_fold` helpers reference `(type $cl)` + `call_indirect`, so even programs whose closure values aren't in the table need `(table 0 funcref)` empty-declared; isolated via `vec_higher_order_used : bool ref` flag + separate runtime block. On each backend, App handler unwraps curried Apps (vec_set / vec_fold are 3-arg = 2-stage unwrap; vec_iter is 2-arg = 1-stage). Added `examples/vec_higher_order_codegen.mere` (interpreter + 3 backends return 1234 demo). Added 12 tests (1197 passing) — 3 backends × (vec_set / vec_iter / vec_fold) codegen + interpreter parity. Remaining: `vec_map` (region-preserving new Vec creation) / `vec_filter` (dynamic size calc) / `vec_to_list` / `vec_to_owned` / `OwnedVec` / `StrBuf` / `Map` / first-class value usage.
- Phase 15 #4: Wasm codegen supports `Vec[R, T]` — full 3-backend
feature parity — followed Phase 15.2 (C) / 15.3 (LLVM) and ported Vec to Wasm. In Wasm Mere values are all 4-byte i32 (scalar direct for primitives; structured types are linear-memory offsets), so per-T monomorphization (as in C / LLVM) is not needed. Design call: single `$mere_vec_new / $mere_vec_push / $mere_vec_get / $mere_vec_len` runtime handles all element types. lib/codegen_wasm.ml: (1) added `vec_used : bool ref`, emit_expr sets true when going through vec_*; (2) 4 fns + struct layout `{data:i32, len:i32, cap:i32, _pad:i32}` (16 bytes) written into `vec_runtime` literal in WAT; push's realloc allocates from single `__lang_bump` = arena semantics; (3) `ty_tag` catch-all relaxed to allow TyRef _ R TyUnit (region marker); explicit Vec rejection removed; (4) Var handler's vec_* rejection retained only for first-class value usage; (5) 4 special-cases added to emit_expr — `App (App (Var "vec_push", v), x)` unwrapped to runtime call; `vec_new`'s region argument ignored (Wasm bump is global); (6) introduced `resolve_vec_let_types` pre-pass same as Phase 15.2 / 15.3 (concretizing binding type doesn't directly affect Wasm code but maintained for consistency). Added `examples/vec_codegen_wasm_typed.mere` (int / str / tuple / variant 4 types = 252). Added 4 tests + rewrote existing Wasm rejection test (1185 passing). Now `Vec[R, T]` works on all 3 backends (C / LLVM IR / Wasm) — the constraint "Vec / OwnedVec / StrBuf / Map are interpreter-only" is fully gone for Vec[R, T]. Remaining: higher-order APIs / first-class value usage / OwnedVec / StrBuf / Map codegen remain interpreter-only (see DEFERRED §1.1).
- Phase 15 #3: LLVM IR codegen supports `Vec[R, T]` (C feature
parity) — ported the same monomorphization pattern as Phase 15.2 (C version) to LLVM IR. lib/codegen_llvm.ml: (1) added `vec_instances : (string, Ast.ty) Hashtbl.t`; (2) `emit_vec_runtime_for_llvm` emits one set per element type of `%mere_vec_<tag> = type { ptr, i32, i32, ptr }` + 4 helpers (`_new` / `_push` / `_get` / `_len`) (using LLVM's `getelementptr ... null, i32 1 → ptrtoint` idiom for sizeof(T), allocates via region; push's realloc within same region = arena semantics); (3) `llvm_ty_of` walks `TyCon ("Vec", args)`, returns Vec value as LLVM opaque ptr (`ptr`) and registers element type in `vec_instances`; (4) `ty_tag` catch-all relaxed to allow `TyRef _ R TyUnit` (region marker); (5) Var handler's vec_* rejection retained only for first-class value usage; (6) 4 special-cases (`vec_new` / `vec_push` / `vec_get` / `vec_len`) in emit_expr — `vec_elem_tag_of` reads element type; unwrap curried App (`App(App(Var "vec_push", v), x)`) and call `@mere_vec_<tag>_*`; `vec_new` pulls active region from `current_regions` and passes `@__lang_default_region` or `%__region_R`; (7) introduced `resolve_vec_let_types` pre-pass same as Phase 15.2 — connect let-poly generalized binding and use tyvars with `Typer.unify`; once any use site resolves, chain-propagates to all sites. Added `examples/vec_codegen_llvm_typed.mere` (mixes int / str / tuple / variant 4 types in one program; total 252). Added 5 tests (1182 passing) — confirms emit of mere_vec_T_new runtime for 4 patterns Vec[R, int] / str / tuple / region R inside. Remaining: Wasm backend Vec[R, T] (Phase 15.4 candidate) / higher-order APIs / first-class value / OwnedVec / StrBuf / Map.
- Phase 15 #2: C codegen generalizes element type T of `Vec[R, T]`
— extends Phase 15.1 (Vec[R, int] only) to support any concrete element type supported by codegen: int / bool / str / tuple / record / variant. Monomorphize emits mere_vec_<tag> runtime struct + 4 helpers (_new / _push / _get / _len) per element type (e.g. mere_vec_int / mere_vec_str / mere_vec_tuple_int_int / mere_vec_Tag). lib/codegen_c.ml: (1) added vec_instances table; c_type_of / emit_expr register T encountered in Vec[_, T] sanitized via ty_tag; (2) emit_vec_runtime_for : Ast.ty -> string generates C runtime block per element type; (3) emit_expr's 4 special-cases (vec_new / vec_push / vec_get / vec_len) routed to mere_vec_<tag>_* helper names via vec_elem_tag_of helper; (4) let v = vec_new () in body generalized binding (Mere has no value restriction; generalized to forall T. Vec[..., T]) leaves App's own .ty TyVar unresolved; added resolve_vec_let_types pre-pass — for each Let(P_var name, value, body) where value.ty is Vec, connect all Var name in body to binding side via Typer.unify; once any use site (e.g. vec_push v 10) resolves, chain-propagates to all sites; (5) element type's C struct may be forward-referenced by later closure typedef etc.; insert typedef struct mere_vec_<tag> mere_vec_<tag>; forward typedef after tuple/record/variant bodies. Added examples/vec_codegen_c_typed.mere (mixes int / str / tuple / variant in one program; total 252). Added 2 tests + rewrote existing "Vec[R, <non-int>] reject" test to "str / tuple accept" (1178 passing). Remaining Vec codegen listed in §1.1 (higher-order APIs / first-class value / LLVM/Wasm / OwnedVec / StrBuf / Map).
- Phase 15 #1: C codegen for `Vec[R, int]` (DEFERRED §1.1 partial
resolution) — first step toward native-izing interpreter-only Vec in the smallest scope (element type int / C backend only). Added `mere_vec_int` struct + `mere_vec_int_new / push / get / len` helpers to `lib/codegen_c.ml` runtime (region-allocated; push's realloc allocates new buffer in same region; old buffer reclaimed at region free = arena semantics). Fixed `c_type_of` to walk `Ast.walk` TyCon args, then map `TyCon ("Vec", [_; TyInt])` to `"mere_vec_int*"`. Added 4 special-cases to `emit_expr` `App` handler (`vec_new` / `vec_push v x` / `vec_get v i` / `vec_len v`) — vec_new reads active region binding via `Ast.walk e.ty` (outside → `__heap` = `__lang_default_region`; inside region R → `__region_R`) and expands to `mere_vec_int_new(&...)`. Remaining 3 unwrap curried form (`App (App (Var "vec_push", v), x)`) via inner/outer combo to runtime helper calls. Relaxed `ty_tag` catch-all rejection to pass only `TyRef` (region marker). Var handler's vec_* rejection kept only for first-class value usage (`let f = vec_new in ...`); direct application changed to pass. Added `examples/vec_codegen_c.mere`: returns 95 computing `vec_new () + push×5 + get / len` in outside-region (verified working via `clang` native binary). Added 6 tests (1177 passing): C codegen accepts Vec[R, int]; runtime helpers emitted; binds to `__lang_default_region` outside / to `__region_R` inside; non-int like Vec[R, str] still rejected; LLVM / Wasm continue rejecting all Vec. Remaining Vec codegen listed in §1.1 (higher-order APIs / first-class value / LLVM·Wasm support / OwnedVec / StrBuf / Map / element types other than int).
- Phase 14 #2: rename codebase from working name lang-ml → Mere —
followed Phase 14.1 name fixation (internal design notes) and changed code body / extensions / docs to Mere across the board. dune library lang_ml → mere (lib/dune); executable main → mere (bin/dune); bin/main.ml → bin/mere.ml (git mv); Lang_ml.* → Mere.* (bin/mere.ml / lib/codegen_llvm.ml / lib/repl.ml / test/test_basic.ml); examples/.lang → .mere (37 files, git mv); updated internal .lang references to .mere (comments in examples / import "..." paths / docs / repl_session.md); CLI usage lang-ml → mere; REPL startup message updated. Updated all Lang / lang-ml / .lang notation in docs / README / CLAUDE.md. Lang in sentences ("Lang program", "of Lang", etc.) also changed to Mere. Intentionally left design context directory internal design notes as-is (historical record). All 1171 tests pass. DEFERRED §7.1 (rename work) moved to fully resolved. Remaining GitHub repo rename (lang-ml → mere) is a user manual operation.
- Phase 12 #10: reverse `owned_vec_to_vec` (DEFERRED §3.6 fully
resolved) — follows Phase 12.11 one-way (`vec_to_owned`) with reverse `owned_vec_to_vec : OwnedVec[T] -> Vec[R, T]`. Region R injected from `active_regions` at call site as App-handler special-case same as `vec_new` / `strbuf_new` / `map_new` (outside → `__heap` default). Eval is `Array.copy` for deep copy (V_vec shared, copy alone yields independence). 3-backend codegen interpreter-only stub. Verified: outside → `Vec[__heap, T]`; `region R { owned_vec_to_vec o }` → `Vec[R, T]` (escape check works); deep copy means subsequent owned-side push doesn't affect vec. Added 5 tests (1171 passing). DEFERRED §3.6 fully resolved.
- Phase 13 #1: type error UX continued — did-you-mean for record
field / view field / qualified name — partially consumes DEFERRED §5.1. Switched `Field_get` family errors (view / record) and `Record_update` field mismatch errors in `lib/typer.ml` to go through `raise_with_suggestion`: passes the corresponding record / view's declared field name list as candidates and adds nearby names by Levenshtein distance as `did you mean \`X\`?` in help: message. Qualified name typo (e.g. `Math.factrial` → `Math.factorial`) needs no implementation change — when env lookup for `Var "Math.factrial"` fails, existing `Var` branch uses entire env (including M-prefixed bindings inside Module) as candidates and calls suggest_name, which works naturally. Verified: `Pt { name, value }` then `p.namee` → `did you mean \`name\`?`; same for view fields; same for `{ p | namee = ... }` record update; `Math.factrial 5` → `did you mean \`Math.factorial\`?`. Added 4 tests (1166 passing). Remaining DEFERRED §5.1 (type variable rename hint / N-best candidate display) in separate slice.
- Phase 12 #9: `vec_filter` / `vec_to_list` / `vec_to_owned` —
consumes DEFERRED §3.5 remainder and §3.6. Added 3 builtins: vec_filter : Vec[R, T] -> (T -> bool) -> Vec[R, T] (region-preserving, keeps only elements where predicate is true); vec_to_list : Vec[R, T] -> T list (converts to 'a list = Nil | Cons of 'a * 'a list, builds Cons chain via Array.fold_right); vec_to_owned : Vec[R, T] -> T OwnedVec (Array.copy deep copy, returns OwnedVec independent of source — a way to extract region-internal Vec to heap). All schemes region-polymorphic; vec_to_owned result is drop_types-registered OwnedVec type so cannot be placed in region (region R { ... vec_to_owned v ... &R ... } auto-rejected as Trivial[R] violation). 3 backend codegen interpreter-only stubs for all 3 builtins. Added 10 tests (1162 passing): 3-scheme type inference; filter behavior / empty result; list conversion + empty Vec → [] display; deep copy to OwnedVec + independence from source mutations; region escape rejection. DEFERRED §3.5 fully resolved; §3.6 updated to one-way (Vec→Owned) resolved (reverse Owned→Vec needs region context, separate slice).
- Phase 9 #5: precise import paths (importer-relative +
canonicalisation) — consumes DEFERRED §4.2. Phase 9.2 introduced cwd-relative `import "path";`; changed to importer-relative (resolved from the file containing the import statement). Added `Parser.current_base_dir : string ref`; `parse_program ?(base_dir = Sys.getcwd ())` for initial value. `import` branch: relative path via `Filename.concat !current_base_dir path`; canonicalized via `Unix.realpath`; during recursive parse swap `current_base_dir := Filename.dirname canonical` (restored on exception). Added `?base_dir` to Pipeline.process; CLI (bin/main.ml) passes `~base_dir:(Filename.dirname path)` in file mode. Canonicalisation makes different relative forms (e.g. `/tmp/foo.mere` vs `./foo.mere`) refer to same file → accurate cycle guard. Verified: `import "./sub/inner.mere"` resolves from main.mere's dir; nested imports (main → middle → sub/inner) work from each step's dir; same file via different relative forms loaded once. Added 3 tests (1152 passing). DEFERRED §4.2 updated to resolved.
- Phase 9 #4: `type` / `record` declaration inside modules —
consumes last 1/3 of DEFERRED §4.1. Extracted T_type branch logic inside parse_decls (including record / variant / alias disambiguation) into helper parse_type_decl_after_keyword; added T_type branch to parse_module_body calling same helper. As a slice-1 limitation, type / record / constructor names are not M-prefixed and enter global registry — declaring same-named type in different modules conflicts (proper scoping in subsequent slice). Verified: module M { type Pt = { x: int, y: int }; let mk = fn p -> Pt { ... } }; compute p.x + p.y from M.mk (3, 4); module M { type 'a opt = ... }; M.unwrap (S 42) dispatches via variant; type and let mix OK. Added 3 tests (1149 passing). DEFERRED §4.1 fully resolved (3/3).
- Phase 9 #3: nested modules + `open M;` — consumes 2/3 of
DEFERRED §4.1 (remaining: type / record inside module). Refactored parse_module_body to take cur_path parameter; handles T_module T_ident inner T_lbrace recursively. Registers both short name (inner) and full path (outer.inner) to module_names; qualified access from both inside and outside works. Newly added module_bindings : (string, string list) Hashtbl.t registry — inside prefix_module_decls, records direct binding names (only names without dots); used to expand open M;. Added open keyword + T_open token to lexer; added T_open T_ident name T_semi branch to parser's parse_decls: extract module_bindings[m_name] and expand to chain of Top_let (P_var n, Var "M.n") aliases; unregistered module is parse error. Nested module direct binding names containing dots are excluded from open expansion (e.g. module M { module N { ... }; let g = ... } with open M; brings in only g; N exports referenced as M.N.foo). Verified: module M { module N { let f = ... }; let g = N.f + 1 }; M.N.f + M.g works; shortcut access after open M; coexists with M.foo qualified access. Added examples/module_nested.mere. Tutorial 10.5 updated: nested + open usage + constraints. Added 7 tests (1146 passing). DEFERRED §4.1 updated to "2/3 resolved" (type / record inside module is future work).
- Phase 11 #7: borrow checker refinement (3) — borrow propagation
from match arms — continues DEFERRED §2.2 (match patterns). Added Match case to `extract_borrows`: union of `extract_borrows` from each arm body (which arm runs is runtime-dependent, so conservatively treat all arms as active). Guards are side conditions so not subject to extraction. While we're at it, extended `Let_rec` / `With` / `Region_block` bodies to also traverse recursively (these values can leak borrows when let-bound). Verified: `let r = match v with | N -> &R x | S _ -> &R x in let m = &mut R x in 0` → conflict; `let r = match v with | N -> &R x | S _ -> &R y in let m = &mut R y in 0` → conflict (else branch equivalent &R y also active); unrelated `&mut R z` OK. Added 3 tests (1139 passing). Remaining borrow checker DEFERRED: §2.3 NLL only.
- Phase 11 #6: borrow checker refinement (2) — borrow propagation
through if branches — consumes DEFERRED §2.2. Up through Phase 11.5, only `Let (P_var _, Ref ..., body)` patterns added borrow to active set; couldn't detect cases where if expression result leaks the borrow, like `let r = if cond then &R x else &R y in ...`. Added helper `extract_borrows : Ast.expr -> (region * place * mode * loc) list`: Ref to single-element list; If(cond, t, e) to union of extracts from t/e; Let(_, _, body) recurse from body; Annot recurse from inner; otherwise empty list. Refactored `check_borrows` `Let` branch: pass value through `extract_borrows` to get borrows propagating up; conflict-check each one and add to active set; pass union to body. Verified: `let r = if c then &R x else &R y in let m = &mut R y in 0` → conflict (else branch from y also active); `let r = if c then &R x else &R y in let m = &mut R z in 0` → OK (z unrelated); nested let-in-if recurses properly. Added 5 tests (1136 passing). Next stage is §2.3 NLL (Non-Lexical Lifetimes) — releasing borrow at "the moment it stops being used", equivalent to liveness analysis.
- Phase 11 #5: borrow checker refinement (1) — tracking complex
expressions (field chain) — consumes DEFERRED §2.1. Phase 11.4 only tracked simple Var for `x` in `&[mode] R x`; extended to identify field chains like `p.field` / `p.q.r`. Added `place_id : Ast.expr -> string option` helper (Var → Some name, Field_get inner f → Some "<inner>.<f>", otherwise None). Replaced Var-only checks in `check_borrows` `Ref` / `Let` branches with place_id based. Non-place expressions (function call results, literals etc.) continue to be skipped (None). Error messages also display dotted paths like `&R p.x`. Verified: `&R p.x + &mut R p.x` → conflict; `&R p.x + &mut R p.y` → OK; `&R p.x + &R p.x` → OK (shared read each other); `&R o.inner.v + &mut R o.inner.v` → conflict (nested chain); `&R p + &mut R p.x` → OK (whole p and p.x are separate places). Added 6 tests (1131 passing). Remaining borrow checker DEFERRED: §2.2 control flow analysis (separate borrow sets per if branch) and §2.3 NLL in separate slices.
- Phase 12 #8: `Map[R, K, V]` (region-aware mutable map) —
Minimum harness for design doc 13_region_std_types.md §5 Map[R, K, V]. Same construction-time binding pattern as Vec[R, T] / StrBuf[R]. Type is 3-arg TyCon ("Map", [TyRef BorrowedRead R TyUnit; K; V]). Eval has V_map of (value, value) Hashtbl.t (OCaml polymorphic hash/eq) + 5 builtins (map_new / map_set / map_get / map_has / map_len). map_get on missing key is eval error; map_has for safe check. Typer has 5 schemes (region / K / V each as TyVar for polymorphism); types["Map"] = 3; App (Var "map_new", _) special-cased pulls region binding from active_regions (empty → __heap). Ast.pp_ty has 3-arg Map[R, K, V] bracket display (TyRef-of-unit / polymorphic both handled). Added V_map case to Phase 12.6 len builtin for polymorphic len. All 3 backend codegen interpreter-only stubs for Map type / 5 builtin names. Added examples/map_basics.mere: simple str→int, has-safe lookup, int→str (type reversal), short-lived inside region — 4 patterns demo. Tutorial 10.6 added Map API table + caveats (closure / ref as key identified per-ref). Added 10 tests (1125 passing): 5-scheme type inference; basic set/get; has branch; len with duplicate key; polymorphic type (int → str); eval error on missing key; region escape rejection; outside-region default; polymorphic len integration; codegen rejection. Now Q-010 main collections (Vec / OwnedVec / StrBuf / Map) all work in interpreter. Remaining: trait system proper (§3.1), unified Allocator trait API (§3.4), OwnedVec / Vec round-trip (§3.6), 3-backend codegen (§1.1).
- Phase 12 #7: Vec higher-order APIs (iter / map / fold / set) —
Implemented higher-order functions intended for Vec API in design doc 13_region_std_types.md §3. All region-polymorphic + element type polymorphic. vec_map result Vec bound to same region as source (region-preserving). Schemes: vec_iter : Vec[R, T] -> (T -> unit) -> unit; vec_map : Vec[R, T] -> (T -> U) -> Vec[R, U]; vec_fold : Vec[R, T] -> U -> (U -> T -> U) -> U; vec_set : Vec[R, T] -> int -> T -> unit. Eval calls user functions (V_closure / V_builtin) via apply_value_ref pattern (same as flip / try_or / iter_n etc.); placement after apply_value_ref definition. vec_set is in-place mutation; out-of-range index is eval error. 3 backend codegen interpreter-only stubs for all 4 names. Added examples/vec_higher_order.mere: int→int map / int→str map / fold for sum and max / set + iter / chain inside region — 5 patterns demo. Tutorial 10.6 section added higher-order API table + usage examples. Added 12 tests (1115 passing): 4-scheme type inference; map (incl. element type conversion); fold (sum); set + out-of-range; iter side effects via separate Vec; region-preserving behavior; codegen rejection. Remaining Q-010: Map[R, K, V]; Allocator trait; Vec / OwnedVec / StrBuf codegen support.
- Phase 12 #6: `StrBuf[R]` (Q-010 narrowed — region-internal mutable
string buffer) — Minimum harness for design doc 13_region_std_types.md §4 `StrBuf[R]`. Same construction-time binding pattern as `Vec[R, T]` (Phase 12.3); type is 1-arg `TyCon ("StrBuf", [TyRef BorrowedRead R TyUnit])` (region marker only, same convention as view types). Added `V_strbuf of Buffer.t` to eval (internal storage in OCaml Buffer); `to_string` formats as `StrBuf["..."]`. Builtins: `strbuf_new : unit -> StrBuf[R]`, `strbuf_push : StrBuf[R] -> str -> unit`, `strbuf_to_str : StrBuf[R] -> str`, `strbuf_len : StrBuf[R] -> int`. Added 4 schemes to typer in polymorphic-region form (TyVar in region position); pre-register `types["StrBuf"] = 1`; `App (Var "strbuf_new", _)` special-cased same as vec_new pulls region binding from active_regions (empty → __heap). Added polymorphic `StrBuf[a]` bracket display to `Ast.pp_ty`. Added `V_strbuf` case to Phase 12.6 `len` builtin for length via polymorphic. 3 backend codegen rejects both type / builtin as interpreter-only. Added `examples/strbuf_basics.mere`: outside-region (default `__heap`) / inside region (auto-bound to `StrBuf[R]`) / polymorphic `len` — 3 patterns demo. Tutorial 10.6 updated: StrBuf[R] explanation + constraints. Added 9 tests (1103 passing): type inference; push/to_str round-trip; empty len; inside region binding; escape rejection; polymorphic len integration; codegen rejection. Remaining Q-010: `Map[R, K, V]`; Allocator trait; Vec/OwnedVec/StrBuf codegen support.
- Phase 12 #5: ad-hoc polymorphic `len` (Q-010 narrowed / lightweight
unified trait-style API) — Minimum practical alternative to a full trait system planned for `trait Collection { fn len(self) -> usize }` in design doc 13_region_std_types.md §6. Instead of introducing a full trait system (~500 LoC), added `len : 'a -> int` as an ad-hoc polymorphic builtin in the same frame as `show : 'a -> str`. Single scheme in typer (`'a -> int`); eval dispatches based on runtime value variant: `V_vec` (shared by Vec[R, T] and OwnedVec[T]) → array length; `V_str` → byte length; `V_tuple` → arity; `V_constr (Nil/Cons chain)` → list traversal counts elements; otherwise eval error. Single API for Vec[R, T] / OwnedVec[T] / `'a list` / `str` / `tuple`. 3 backend codegen reject `len` as interpreter-only stub. Added 8 tests (1094 passing): type inference; behavior for str / Vec / OwnedVec / tuple / list; eval error for unsupported value (int); codegen rejection. Full trait system introduction in future slice — whether trait's implicitness fully aligns with Mere's design philosophy (explicit > concise) is on hold.
- Phase 12 #4: `OwnedVec[T]` (Q-010 narrowed (b) separate type) —
Implemented "separate type" portion of design doc 13_region_std_types.md §9 "(b) separate type + trait for unified API". Added OwnedVec[T] (heap-allocated, has Drop) in contrast to Vec[R, T] (region-internal, Trivial). Added owned_vec_new / push / get / len schemes (1-arg, 'a OwnedVec form) to typer; types["OwnedVec"] = 1 + registered in `drop_types` so that region-placement triggers automatic rejection by contains_drop_type (Trivial[R] violated: cannot place value of type \'a OwnedVec\ into region — type contains a Drop type). Eval shares V_vec (only type system treats them as different; internal implementation is the same mutable array). 3-backend codegen rejects both owned_vec_ builtins and OwnedVec type as interpreter-only (unified message `Vec / OwnedVec builtins are interpreter-only`). Added `examples/vec_vs_owned_vec.mere`: contrasts short-lived region Vec and long-lived OwnedVec in one program. Tutorial 10.6 updated: OwnedVec[T] explanation + how to choose vs Vec[R, T]. Added 6 tests (1086 passing): type of owned_vec_new; polymorphic push/get/len; region rejection via Drop; contrast that Vec[R, T] can be placed in region; 3-backend codegen rejection. Remaining Q-010: `StrBuf[R]` / `Map[R, K, V]`; unified Allocator trait API (trait-based unification of read API); Vec / OwnedVec codegen support.
- Phase 12 #3: semantic backing for `Vec[R, T]` (Q-010 narrowed →
implementation stage 3) — Gives type system that actually tracks region to `Vec[R, T]` syntax that was parse-only in Phase 12.2. Changed Vec arity from 1 → 2; internal representation unified to `TyCon ("Vec", [TyRef BorrowedRead R TyUnit; T])` (region marker convention same as view types). Parser: `Vec[R, T]` emitted as 2-arg; legacy `T Vec` (1-arg postfix) auto-filled with default region `__heap` and expanded to 2-arg form (forward-compat). With TyVar in region position of scheme, region-polymorphic APIs are realized through scheme machinery as-is (`vec_push : forall T R_marker. Vec[R_marker, T] -> T -> unit`); R_marker unifies with concrete region marker at call site. Added special handler to `Typer.infer` App case: `App (Var "vec_new", _)` reads innermost active_regions and directly binds region of `Vec[R, T]` (same shape as view construction); empty → `__heap`. Added bracket display for 2-arg Vec to `Ast.pp_ty` (`Vec[R, int]` / `Vec[__heap, 'a]` / `Vec['a, 'b]` etc.). Verified: `vec_new ()` outside → `Vec[__heap, 'a]`; `region R { vec_new () }` → `Vec[R, 'a]` (escape is static error); `fn (v: Vec[R, int]) -> vec_len v` → `(Vec[R, int] -> int)`; `fn (v: int Vec) -> vec_len v` → `(Vec[__heap, int] -> int)`. Updated `examples/vec_basics.mere`: demonstrates auto-bind of region for `vec_new ()` inside region. Tutorial 10.6 updated: noted that region got semantic backing + explicit escape check. Added 3 tests + updated 7 existing tests to new format expectations (1080 passing). Remaining Q-010: explicit distinction from OwnedVec[T]; StrBuf[R] / Map[R, K, V]; unified Allocator trait API; Vec codegen support.
- Phase 12 #2: `Vec[R, T]` syntax (Q-010 narrowed → implementation
stage 2, lightweight) — Forward-compatible slice that accepts the notation `Vec[R, T]` from design doc 13_region_std_types.md into parser. Added `T_ident name :: T_lbracket :: ...` branch to `simple_ty` in `lib/parser.ml` (name is uppercase): parses bracket-delimited argument list; region marker (bare uppercase ident yielding TyCon name=[]) dropped; remaining type arguments passed to `expand_alias_or_tycon name type_args`. Result is that `Vec[R, int]` is internally identical to `int Vec` (1-arg TyCon) — generates same TyCon. Region R is a documentation marker currently with no semantic backing (region-aware allocation / lifetime tracking implementation planned in future slice). Updated `examples/vec_basics.mere`: demonstrates `(vec_new () : Vec[R, int])` annotation inside region. Tutorial 10.6 section updated: `Vec[R, T]` syntax can now be written; current R is documentation only; both forms (`int Vec` / `Vec[R, int]`) produce equivalent types. Added 3 tests (1077 passing): type annotation parse; str version; `int Vec` and `Vec[R, int]` produce same type. Implementation scale: only ~25 lines added to parser.ml. Next slice (12.3) gives R semantic backing: reflect active_regions in vec_new return type (view construction pattern).
- Phase 12 #1: `'a Vec` minimum harness (Q-010 narrowed →
implementation stage 1) — Adds basic variable-length vector as polymorphic builtin under name `'a Vec`, the most basic of design doc `13_region_std_types.md` region-version std types. Phase 12 total (Vec[R,T] / OwnedVec[T] / StrBuf[R] / Map[R,K,V] / Allocator trait etc.) narrowed to MVP; syntax for region parameters in type and distinction from OwnedVec come in subsequent slices. Added `V_vec of value array ref` (storage in OCaml mutable array; push appends with reallocate) + 4 builtins (`vec_new : unit -> 'a Vec`, `vec_push : 'a Vec -> 'a -> unit`, `vec_get : 'a Vec -> int -> 'a`, `vec_len : 'a Vec -> int`) to `lib/eval.ml`; `to_string` formats as `Vec[...]`. Added 4 schemes (`vec_new_scheme` etc.) to `lib/typer.ml`; `Hashtbl.replace types "Vec" 1` pre-registers as arity-1 polymorphic type. Registered in `initial_env`. Trivial[R] check works because existing `contains_drop_type` walks recursively, so placing `Conn Vec` (where Conn is a drop type) in region is auto-rejected. Added explicit stubs to Var handlers of codegen (C / LLVM / Wasm) raising `Codegen_error` when they see `vec_new` / `vec_push` / `vec_get` / `vec_len` (all 3 backends emit `interpreter-only` message). Added `examples/vec_basics.mere`: basic operations on int / str Vec + Vec inside region demo. Added 14 tests (1074 passing): type inference for 4 builtins; len of empty Vec; len/get after push; polymorphic (str Vec); region placement OK; Conn Vec rejected with Trivial[R]; eval error for out-of-range get; 3-backend codegen rejection. Future slice candidates: Vec[R, T] with region as parameter + Allocator trait + distinction from OwnedVec[T].
- Phase 11 #4: borrow checker minimum harness — Slice that
consumes Q-004 "remaining implementation TODO". Added check_borrows : (string * string * borrow_mode * Loc.t) list -> Ast.expr -> unit to lib/typer.ml. Threads borrows for the same (region, var name) as active set through lexical scope; rejects coexistence of conflicting modes with Type_error. Coexistence allowed pairs defined in borrows_compatible: only shared read with shared read (BorrowedRead + BorrowedRead) and shared write with shared write (SharedWrite + SharedWrite); all else conflicts (exclusive family doesn't coexist with anything; shared read + shared write also rejected due to invalidation risk). AST walk: when discovering Let (P_var p, Ref (mode, region, Var v_name), body), adds (region, v_name, mode, value.loc) to active set and recurses on body; free-standing &[m] R v also conflict-checks with active. Pipeline.process calls Typer.check_borrows [] (Ast.desugar_program prog) after Typer.infer to inspect program in one pass. Added examples/borrow_conflict.mere (intentional failure demo: taking &mut R v after &R v). Error message includes "previous borrow at line N, col N" note. Verified: let a = &R v in let b = &mut R v / let a = &mut R v in let b = &mut R v / let a = &R v in let b = &shared write R v / let a = &exclusive R v in let b = &R v all reject as conflict; let a = &R v in let b = &R v / 2 shared write / different variables OK. Added borrow checker explanation + conflict example output to docs/tutorial.md 10.4 section. Added 8 tests (1060 passing). Currently tracking is limited to simple Var for x in &[m] R x — complex expressions (&R rec.field etc.) in future. Now Q-004 design (b) borrow annotation refinement is complete in both "can be written as types + machine-verifies conflict".
- Phase 11 #3: auto-deref for field access through `&R T` — At
Phase 11.1 borrow annotation introduction, field access like lg_ref.info "hi" was crashing with field access on non-record value. Added strip_refs helper to Field_get case of lib/typer.ml (recursively peels TyRef wrappers); changed to perform existing view / record judgment on type after peeling. Borrow mode remains static contract; eval side already passed &R v through, so zero runtime changes. Result: method calls work directly through any of &R Logger / &mut R Logger / &shared write R Logger, like lg.info "msg". Fully rewrote examples/borrow_modes.mere: rewrote signature-only demo to actually call cap methods (log_action, db_run, show_config) across borrow; prints mk_logger's [INFO] output + DbHandle's exec call + AppConfig's name/threads read. Added 5 tests (1052 passing): field access on Pt record through &R; through &mut R; through &shared write R; type inference for user-defined Lg11; type confirmation extracting field from &R Lg11r.
- Phase 11 #2: borrow annotation realistic example + tutorial 10.4
section — Milestone showing "what is it good for" of the 4 modes added in Phase 11.1 (`&R T` / `&mut R T` / `&shared write R T` / `&exclusive R T`). Added `examples/borrow_modes.mere`: realistic demo constructing 3 kinds — Logger (shared write) / DbHandle (exclusive write) / AppConfig (shared read) — inside region, then borrowing each cap with appropriate mode and passing to handler. Run prints "[logged] save_order" / "[exclusive] UPDATE ..." / "[read]". Added `examples/borrow_modes_typeerror.mere`: intentionally fails with type error demo passing `&R db` to `&mut R DbHandle` parameter (displays as documentation that `expected \`&mut R DbHandle\`, got \`&R DbHandle\`` is shown). Added 10.4 "Borrow annotation" section to `docs/tutorial.md` (4-mode table + usage examples + mode mismatch error example + current limitations (borrow checker exclusion rules and `&R T` field auto-deref are future work)). Also added 2 new examples to section 12 examples list. No test count change (1047 still). Phase 11.1 brought "writable as type" state; Phase 11.2 brought "readable with understood meaning" state. Next slice candidates: borrow checker (exclusion rules) and `&R T` field auto-deref.
- Phase 11 #1: borrow annotation refinement (Q-004 narrowed →
implementation stage 1) — Minimum harness for narrowing (b) borrow annotation refinement in design doc 08_effect_granularity.md down to implementation. Added `borrow_mode = BorrowedRead | SharedWrite | ExclusiveRead | ExclusiveWrite` to AST; signatures for `TyRef of borrow_mode * string * ty` (type level) and `Ref of borrow_mode * string * expr` (value level) changed to 3-arg. 4 new syntaxes in parser: `&R T` (default = BorrowedRead); `&mut R T` (ExclusiveWrite); `&shared write R T` (SharedWrite); `&exclusive R T` (ExclusiveRead). Value level `&R v` / `&mut R v` / `&shared write R v` / `&exclusive R v` similarly. `mut` / `shared` / `write` / `exclusive` are contextual keywords (regular idents in lexer; parser recognizes only after `&`). Typer's unify changed to require "region and mode equality" for `TyRef (m1, r1, t1) ↔ TyRef (m2, r2, t2)` (strict, no subtyping). pp_ty handles `&R T` / `&mut R T` / `&shared write R T` / `&exclusive R T`. Codegen (C / LLVM / Wasm) ignores mode — pointer representation is the same; only static guarantee. Verified: `fn (x: &mut R int) -> ...` type display OK; passing `&R 5` to `fn (x: &mut R int) -> 1` is type error `expected \`&mut R int\`, got \`&R int\``; calls with same mode pass; `(&R 5 : &mut R int)` annotation mismatch is type error. Logger problem (shared write representation) solved at syntax level; borrow checker (exclusion rules) in future slice. Added 14 tests (1047 passing).
- Phase 10 #1: aggregating where we are — tutorial / README / new
examples / SUMMARY — Milestone with 1033 tests / 3 backends / REPL / module / import in place; arranging outward-facing documentation. Added 10.5 "Modules and import" section and 11.5 "Using the REPL" section to `docs/tutorial.md`; rewrote 13 "Native compilation" from C-only to 3-backend (C / LLVM / Wasm); updated closing remark from "memory model is not implemented in codegen" to "works in all backends". Full rewrite of `README.md`: status as of 2026-06-19 (1033 tests / 3 backend parity / module / import / REPL commands); added rows for module, import, REPL command, error UX to features table; added LLVM / Wasm build paths to build examples. New examples: `examples/module_basic.mere` (`module Math { let inc / square / pow / inc_then_square ... }` + shortened internal reference demo); `examples/lib_list_ops.mere` (decls-only library exporting `module ListOps { sum / length / map }`); `examples/import_demo.mere` (imports lib with `import "examples/lib_list_ops.mere";`); `examples/repl_session.md` (Markdown showing `:type` / `:env` / `:show` / `:load` / `:reset` / multi-line in dialog session format). Created new `internal design notes`: restructured destinations of Phases 1-9 as "outward-facing" (5-min status delivery to future self / sharing partners); aggregates feature coverage, history phase table, what's missing, next directions. No test count change (1033 still).
- Phase 9 #2: file split — `import "./other.mere";` — Added
import keyword + T_import token to lexer. Added imported_files : (string, unit) Hashtbl.t registry and parse_decls T_import T_string path T_semi branch to parser: reads target file with In_channel.with_open_text, recursively calls Lexer.tokenize + parse_program_internal, mixes resulting decls into current decl stream with List.rev_append (discards main expression). Skips same path if already registered (cycle prevention). Split parse_program into parse_program_internal (recursive worker) + parse_program (top-level wrapper, runs worker after Hashtbl.reset imported_files) — top-level cycle guard accumulator extends throughout recursive imports while being fresh per top-level call. Parser registries (constructors / records / module_names / aliases) are shared across recursive calls, so types / records / modules defined in imported files are visible from importer side. Verified: import "/tmp/lib.mere"; helper base references helper / base from another file; import "/tmp/lib_mod.mere"; Math.sq (Math.dbl 5) qualifiedly references module in import; mutual cyc_a ↔ cyc_b imports yield a_val + b_val = 30 (no infinite loop thanks to cycle guard); diamond pattern (importing lib via both A and B) loads once without duplication; missing file is parse error. Added 6 tests (1033 passing). Base path resolution is cwd-based; symlinks / different relative forms treated as different files (canonicalisation in future).
- Phase 9 #1: minimum module harness — `module M { let f = ...; }`
+ M.f reference — Next milestone for language surface. Added `module` keyword + `T_module` token to lexer; added `module_names : (string, unit) Hashtbl.t` registry and `parse_module_body` to `parser.ml` (slice 1: only `let` / `let rec`; terminates at `T_rbrace`); added `prefix_module_decls` (rewrites binding names and free Var references in body with `M.` prefix). Newly implemented `Ast.rename_free_vars`: shadowing-aware AST walker that excludes bind names computed by `pattern_vars` from shadow list in `Fun (param, ...)` / `Let (P_var p, ...)` body / `Let_rec [(n, _); ...]` / `With (n, ...)` body / `Match` arm patterns. Extended parser's `field_chain`: if lhs is `Var "M"` and `M ∈ module_names`, emits `Var "M.f"` instead of `Field_get`. uppercase ident atom_base also checks `module_names` before constructor / record judgment. Added decls-only mode to `parse_program` (main = `()` if only T_eof); removed `Repl.prepare_input`'s `; ()` hack (made no-op, left as identity wrapper for compatibility). Verified: `module M { let answer = 42; let add = fn x -> fn y -> x + y; }; M.add M.answer 8` → 50; internal `inc (inc x)` shortened references rewritten as `M.inc (M.inc x)`; `let rec fact = fn n -> ... fact (n-1)` M.fact self-call works; `module M; module N;` same-name bindings don't conflict; `p.x` regular field access unchanged. In REPL also can write `module M { ... }` multi-line directly; `M.f` appears in `:env`. Added 7 tests (1027 passing). Types / records / nested modules in future slices.
- Phase 8 #2: REPL continued — `:show NAME` + `:reset` — Added 2
new commands to lib/repl.ml. (1) :show NAME outputs type and value at once: format_show eval_env type_env name helper pulls scheme from type_env and value ref from eval_env respectively, returns string in val NAME : TY\n = VAL format (uses Eval.to_string, so closures are <closure:p>, str is quoted, numbers / records / variants in same formatter). Unbound name yields unbound name: NAME. print_show is print entry of same content. (2) :reset rewinds both envs to Eval.initial_env / Typer.initial_env via do_reset eval_env type_env; displays (envs reset). Added 2 lines to help text. Verified: let x = 42; let g = "hi"; :show x → "val x : int\n = 42"; :show g → "val g : str\n = \"hi\""; :show inc (closure) → "val inc : (int -> int)\n = <closure:n>"; :show nope → "unbound name: nope"; after :reset env cleared, :env → "(no user bindings)". Added 5 tests (1020 passing; split I/O of format_show / do_reset to directly assert pure parts).
- Phase 8 #1: REPL UX improvement — multi-line input +
Diagnostic.format integration + :env / :load — 4-point enhancement to `lib/repl.ml`. (1) Switched to loop accumulating multiple lines with `read_logical_input`: if tentative parse after input yields "error at T_eof location", treats as incomplete and prompts `..>` for continuation; returns `Some input` on parse success. `is_unfinished ~source` judges by whether `Parser.Parse_error` loc matches T_eof loc in tokenize result (`eof_loc` helper + `loc_eq`); `Lexer.Lex_error "unterminated string literal"` also treated as unfinished. Empty line in continuation is `(input aborted)`; line starting with `:` interrupts multi-line buffer for standalone command execution. (2) Replaced `format_exn` with `format_diag ~source`; passes each error (`Lexer / Parser / Typer / Eval`) through `Diagnostic.format ~source ~filename:"<repl>"` — REPL also displays with Rust-style code frame, same as file mode. (3) Added `:env` command: `user_bindings` helper excludes builtin names of `Typer.initial_env` and returns only user-added bindings in insertion order, listed as `val name : type`. (4) Added `:load FILE` command: reads file, adds decls to eval/type env through `process_decl`; displays added bindings as `val name : type` then `(loaded path)`. Updated help text for new commands. Verified: can directly write multi-line `let rec` like fib/factorial in REPL; type error `let x = 5 + "hi" in x` displays caret + help:; `:load /tmp/foo.mere` loads definitions and they can be confirmed with `:env`. Added 9 tests (1015 passing) — REPL helpers (probe_unfinished detects each pattern; user_bindings insertion order / empty user env).
- Phase 7 #7: type error UX — hint expansion + App type error
direction fix — Expanded coverage of `Typer.type_conversion_hint`: (1) `expected int, got bool` → `use \`if b then 1 else 0\` to get an \`int\` from a \`bool\``; (2) `TyTuple ts1` vs `TyTuple ts2` arity mismatch → `tuple lengths differ — expected N element(s), got M`; (3) per-direction branching for `expected fn, got value` (extra arg / partial application); (4) `TyCon (n1, _)` vs `TyCon (n2, _)` name difference → `these are different named types (\`n1\` vs \`n2\`)`. Further restructured `Typer.infer` `Ast.App (f, arg)` case into 3 sub-cases: (a) `tf = Ast.TyArrow (param_ty, ret_ty)` → caret at arg.loc + `expected param_ty, got ta` via `unify arg.loc param_ty ta`; (b) `tf = TyVar _` → fresh var + whole unify as before; (c) others (extra arg case where `inc 3` portion of `int 3 4` is `int` etc.) → dedicated error `expected a function (\`'a -> 'b\`), got \`<actual>\`` + `help: you may be passing one too many arguments (...)`. Verified: `inc 3 4` → "expected a function, got int / help: too many arguments"; `add "hi" 3` → "expected int, got str / help: use str_len" (caret at arg.loc); `add 1 + 2` (= `add 1` arrives at int) → "expected int, got (int -> int) / help: missing an argument"; `true + 1` → "expected int, got bool / help: use if b then 1 else 0"; `f (1, 2, 3)` (where f is `(int, int) -> ...`) → "expected (int * int), got (int * int * int) / help: tuple lengths differ — expected 2, got 3"; distinct named records → "expected BarN, got FooN / help: different named types (BarN vs FooN)". Added 6 tests (1006 passing).
- Phase 7 #6: type error UX — type conversion hint — Added
Typer.type_conversion_hint t1 t2 -> string option helper; appends help: ... after base message in unify error (via with_hint). Covered cases: expected str, got int/bool → use \show x\; expected int, got str → use \str_len s\ ...; expected bool, got int/str → wrap in a comparison; expected fn, got value → you may be missing an argument; expected value, got fn → you may have passed a partially-applied function. Other cases get no hint. Verified: "answer: " ++ 42 → help: use \show x\; 5 + "hi" → help: use \str_len s\; if 1 then ... else ... → help: wrap in a comparison. Added 4 tests (1000 passing — milestone).
- Phase 7 #5: type error UX — source span (caret range display
with token width) — Extended `Loc.t` from `{ line; col }` to `{ line; col; width }`; added `Loc.mk ?(width=1) ~line ~col ()` helper (default width = 1 for backward compatibility; `Loc.dummy` has width = 0). In lexer's `tokenize`, attached token char count to pos via `with_width pos w` at output of each token: identifier / tyvar / string literal / int literal / float literal / 1-3 char operator (existing kept at 1). In `Diagnostic.format`, extended caret to multiple chars with `String.make (max 1 width) '^'`; applied bold-red ANSI color to all carets. Verified: in `let y = x + "hello"` error from `^` alone to `^^^^^^^^^^` (10 chars); in `factrial` identifier error to `^^^^^^^^` (8 chars); in `add "hello"` `add` to `^^^` (3 chars). Added 3 tests (996 passing).
- Phase 7 #4: type error UX — ANSI coloring — Added
Diagnostic.use_color : bool ref (default false); CLI (bin/main.ml) sets to true when Unix.isatty Unix.stderr && not NO_COLOR. ansi/red/blue/cyan/bold/bold_red/bold_cyan helpers selectively insert escape codes (\027[CODEm ... \027[0m). In Diagnostic.format, kind is bold-red; line number, |, -->, = in gutter are blue; caret ^ is bold-red; help: / note: keywords are bold-cyan. When use_color = false, everything passes through (test compatibility). Also respects NO_COLOR env var (https://no-color.org/). Verified: when run via TTY (via script), colored; plain when piped; plain when NO_COLOR=1. Added 5 tests (993 passing).
- Phase 7 #3: type error UX — suggesting typo corrections via
Levenshtein — Added `Typer.levenshtein` (edit distance calculation, O(la*lb) DP), `Typer.suggest_name` (`max_dist` based on length, 3/2/1), `Typer.with_hint` / `raise_with_suggestion` helpers. Changed Type_error raises in `unbound variable` / `unknown constructor` / `unknown record type` (both in expression and in pattern) to go through `raise_with_suggestion`; appends `help: did you mean \`<name>\`?` if there's a close candidate. Extended `Diagnostic.format`: splits msg by `\n`; headline goes beside caret of code frame; rest (help:/note:) renders after code frame in `= help: ...` format. Verified: `factrial + 1` (factorial in scope) → "unbound variable: factrial / help: did you mean `factorial`?"; `Greeen` (Color = Red | Green | Blue) → "unknown constructor: Greeen / help: did you mean `Green`?"; `zzzzzz` (no close name) → no hint. Distance threshold adjusts by name length (stricter for short names); tie-break prefers shorter. Added 4 tests (988 passing).
- Phase 7 #2: type error UX — "expected X, got Y" form + audit of
unify call order — Changed `Typer.unify` error wording from `"type mismatch: \`X\` vs \`Y\`"` to `"expected \`X\`, got \`Y\`"` (X=expected, Y=actual). At the same time, unified unify loc t1 t2 calls across Typer to (expected, actual) order: primitive type checks for Neg / Bin (+, -, *, /, %, ++) / Logic / If condition swapped to `unify ... Ast.TyXxx actual` (TyXxx=expected); Fun annotation `unify t' alpha` (annotation=expected); Match guard `unify TyBool tg`; each Match arm `unify result_var tb` (first arm is expected); Record_lit / Record_update field `unify exp_ty t` (declared=expected); Field_get / Record_update base `unify result_ty t_base`. Constr arg `unify exp ta` (param=expected). Symmetric cases (`==` lhs/rhs; if branch then/else; P_or bs1/bs2; let-rec alpha vs body) preserve meaningful order. Pattern checks (P_int/Bool/Str/Unit/constr/tuple/record) were originally `unify expected XXX` (scrutinee=expected) so no change needed. App preserves original `unify tf (TyArrow (ta, result))` (recursive structural unify compares tf.param and ta yielding "expected param_ty, got arg_ty"). Verified: `let y = x + "hello"` → "expected `int`, got `str`"; `add "hi"` (add: int->int) → "expected `int`, got `str`"; `if cond then "yes" else 42` → "expected `str`, got `int`"; record field → "expected `int`, got `str`". Added 4 tests (984 passing).
- Phase 7 #1: type error UX improvement — Rust-style code frame
— Rewrote Diagnostic.format in lib/diagnostic.ml to Rust-style multi-line code frame: header (kind: msg); location pointer --> filename:line:col; line-numbered margin (1 | ...); caret + message below error line ( | ^ ...); context of 2 lines before + 1 line after. Changed terminal error message of Typer.unify from "cannot unify X with Y" to "type mismatch: \X\ vs \Y\" (type names enclosed in backticks, neutral order). At zero-loc, 1-line fallback as before. Verified: let y = x + "hello" displays as type error: type mismatch: \str\ vs \int\ --> file:2:13 | 1 | let x = 5 in | 2 | let y = x + "hello" in | | ^ type mismatch: ... | 3 | y. Parse error / unbound variable error etc. output in common format. Added 6 tests (980 passing). Phase 7 started — improving language surface developer experience.
- Phase 6 #12: Wasm codegen special-cases `'a list` show in
[a, b, c] form — Wasm version of LLVM Phase 5.14. In `emit_show_fn`'s variant branch, processes `TyCon ("list", [elem_ty])` as special-case before others: loop scan with cur / acc / first / tag / pl / h locals. `block $end` + `loop $lp` loads tag from head, break on Nil; on Cons, loads payload (tuple offset) → head = `i32.load offset=0 payload` → concat `, ` if needed (first flag) → concat `show_<elem_tag>(h)` → cur = tail = `i32.load offset=4 payload` → loop. After end, concat `]`. `[` / `]` / `, ` deduped via `intern_show_str`. Verified (wat2wasm + Node.js): `show [1, 2, 3]` → `[1, 2, 3]`; `show (Nil : int list)` → `[]`; `show ["hello", "world"]` → `["hello", "world"]`. Added 3 tests (974 passing). 3 backends (C / LLVM / Wasm) fully parallel — the same Mere program runs on each of 3 backends as native binary / WAT.
- Phase 6 #11: Wasm codegen show general builtin — Wasm version
of LLVM Phase 5.12. Wasm has no asprintf equivalent so all hand-rolled: show_int performs int→decimal string conversion on Wasm (allocates 16-byte buffer from bump pointer → writes digits right-to-left → prepends - if needed → returns pointer to first digit); show_bool registers true / false in data segment and branches with select; show_str is 2-stage concat wrapping with "; show_unit is const offset of (); show_tuple_X_Y concatenates (, each element show, , , ) via __lang_str_concat; show_<R> concats R { f1 = , each field show, , f2 = , }; show_<V> is tag dispatch (nested if/else of i32.load + i32.eq) → each ctor: data ptr direct if nullary; concat ctor_name + " " + recursive payload show if payload. show_types Hashtbl + collect_show_types + add_show_type registers types + recursively registers dependent types (cycle guard). subst_params helper applies args of polymorphic record/variant (Wasm also emits separate function per mono instance; layout is shared). intern_show_str dedupes literals to save data segment. App (Var "show", arg) dispatches to call $show_<ty_tag arg.ty>. Verified (wat2wasm + Node.js): show 42 → "42"; show true → "true"; show "hi" → "\"hi\""; show (1, "hi") → (1, "hi"); show (SS 42) → "SS 42"; show (Pt { x = 3, y = 4 }) → Pt { x = 3, y = 4 }; show (Cons (1, Cons (2, Cons (3, Nil)))) → Cons (1, Cons (2, Cons (3, Nil))) (recursive variant works naturally). Added 8 tests (971 passing). 'a list special-case [a, b, c] form in future slice.
- Phase 6 #10: Wasm codegen complex patterns (P_int / P_str /
P_bool / P_unit / P_record / P_as / nested ctor / or / guard) — Wasm version of LLVM Phase 5.11. Rewrote `compile_pat` as fully recursive `(cond_local_slot, bindings)` function: P_int → `i32.eq`; P_bool → `i32.eq`; P_str → `call $__lang_streq` (new runtime helper, byte-by-byte compare yielding i32 boolean); P_unit → constant true; P_record → declared field order `i32.load offset` + sub-pattern recurse (handles both record / view); P_as → inner pattern + whole value bind; P_tuple → each element `i32.load offset=i*4` + recurse; P_constr → tag test (`i32.load offset=0 + i32.eq`) + sub-pattern recurse (nested OK). Multiple sub-tests chained with `combine_and` helper via `i32.and`. Or-patterns pre-flattened with `expand_or`. Guard evaluated in arm's bindings scope, AND with cond, short-circuit with `if/else` (no guard eval if cond is false). Added `@__lang_streq` runtime helper (block + loop with sequential byte_a / byte_b compare). Verified (wat2wasm + Node.js): `match 3 with | 0 -> 100 | 1 -> 200 | _ -> 300` → 300; `match "hello" with | "hi" -> 1 | "hello" -> 2 | _ -> 9` → 2; `match Cons (SS 5, Nil) with | Cons (SS n, _) -> n` → 5 (nested ctor); `match Pt { x = 3, y = 4 } with | Pt { x = a, y = b } -> a + b` → 7; `(a, b) as p → fst p + snd p + a + b` → 6; `LCgA | LCgB -> 1` → 1 (or); `when n < 10 -> 200` → 200 (guard). Added 8 tests (963 passing).
- Phase 6 #9: Wasm codegen polymorphic variant / record + recursive
variant + P_tuple sub-pattern — Wasm memory layout is uniform (every value is i32 = 4 bytes), so LLVM-style (Phase 5.9 / 5.10) monomorphization is not needed. `'a opt`, `'a Box`, `'a list = Nil | Cons of 'a * 'a list` all work via same code path as mono variant/record. Removed `params <> []` check in `Constr` and `r_params <> []` check in `Record_lit` (Wasm doesn't emit type-specific struct typedefs, so same code works for multi-instantiation). To expand `'a list` Cons (tuple payload `('a, 'a list)`) in Match, added `P_tuple` sub-pattern to `compile_pat` equivalent in `Match`: loads each element from payload tuple offset via `i32.load offset=i*4` into fresh local and binds (`Cons (h, t)` → h, t each loaded into separate locals). Verified (wat2wasm + Node.js): `type 'a opt; match LSome 42 with | LSome n -> n` → 42; `type 'a Box; let bi = Box { v = 42 } in let bs = Box { v = "hi" } in str_len bs.v + bi.v` → 44; `type 'a list; sum [1,2,3,4,5]` → 15; `length ["a","b","c","d"]` → 4. Added 4 tests (955 passing). Wasm backend's advantage: layout uniformity makes monomorphization unnecessary.
- Phase 6 #8: Wasm codegen Region_block + Ref + with Drop + view
construction + Unit_lit — Wasm version of LLVM Phase 5.13. Wasm's linear memory + `__lang_bump` global already acts as one region, so user's `region R { body }` is implemented in LIFO: save current value of `__lang_bump` to local at entry → evaluate body → stash result in another local → restore bump to saved value → push result back. This way allocations within region scope are "freed" at scope end (subsequent allocations can overwrite as bump pointer returns). `Ref (R, v)` (`&R v`) evaluates inner + bump 4-byte alloc + `i32.store offset=0` + push base. `With (c, v, body)` saves v to local + evaluates body + after body, if v's record has `close: unit -> unit` field, pulls env/fn_idx from closure value via `i32.load` + auto-invokes with `i32.const 0` (unit arg) + `call_indirect (type $cl)`, drops result, pushes body value. `view V[R] of T { ... }` Record_lit handled separately by view-name (same memory layout as record, bump alloc + i32.store); Field_get of view value uses field index from `Typer.views.v_fields` with `i32.load offset=idx*4`. `Unit_lit` → `i32.const 0`. Verified (wat2wasm + Node.js): `region R { let x = &R 5 in 42 }` → 42; `with c = mk 7 in c.id * 10` (close prints "closing") → 70; `view Cell[R] of int { v: int }; region R { let c = Cell { v = 7 } in c.v }` → 7. Added 6 tests (951 passing). Wasm backend covers all memory model features, on par with C / LLVM.
- Phase 6 #7: Wasm codegen first-class fn + closure —
Wasm-specific constraint handling: function pointers are not memory ptr but function table indexes; indirect calls go through call_indirect (type $sig). Declared (type $cl (func (param i32) (param i32) (result i32))) at module top; adapters registered in table starting from index 0 via (table N funcref) + (elem (i32.const 0) ...). closure value is 8-byte memory struct { env_offset, fn_table_idx }. Auto-generated env-ignoring adapter (func $f_closure (param i32) (param i32) (result i32) local.get 1; call $f) for each top-level fn f + table registration; recorded index in fn_closure_table_idx. At Var name value position, if fn_closure_table_idx is registered, memory-allocs closure value (env=0, fn_idx=N) and pushes. Indirect App: save closure to local → load env / arg / load fn_idx → call_indirect (type $cl). Anonymous Fun: compute free variables via free_vars → capture only those registered in locals → register fresh adapter anon_N_fn in table → push to pending_closures queue → at construction site, memory-alloc env (store each capture in sequence), alloc closure value + push. Adapter body entry loads captures from env into local slots via i32.load offset=N*4 before evaluating body. Drain loop in emit_program processes pending. Added pattern_vars + free_vars helpers. Verified (wat2wasm + Node.js): let inc = fn x -> x + 1 in let apply = fn f -> f 5 in apply inc → 6; (make_adder 5) 10 → 15; compose inc dbl 5 → 11; twice inc 5 → 7. Added 7 tests (945 passing).
- Phase 6 #6: Wasm codegen variant + match (monomorphic, single
payload type) — Variants also laid out in linear memory: 4 bytes (`{ i32 tag }`) if nullary-only; 8 bytes (`{ i32 tag, i32 payload }`) if payload. `variant_tags : (cname, int) Hashtbl` populated at start of emit_program from `Exhaustive.type_variants`; `variant_payload_ty` helper detects payload type (single type shared by all payload-bearing ctors; Codegen_error if differ). Compiled `Constr cname (arg)` to bump alloc + `i32.store offset=0` (tag) + (if needed) `i32.store offset=4` (payload) + push base. `Match` saves scrut to local, loads tag/payload via `i32.load offset=0/4`; each arm compiles to nested chain of `local.get tag; i32.const N; i32.eq; if (result i32) ... else ... end`; fallthrough traps with `unreachable`. Pattern subset: P_constr / P_var / P_wild; payload bind uses payload local slot. Verified (wat2wasm + Node.js): `type Color = R | G | B; match G with | R -> 0 | G -> 1 | B -> 2` → 1; `type Stat = Ok | Err of str; match Err "boom" with | Ok -> 0 | Err msg -> str_len msg` → 4; `let v = ISome 42 in match v with | INone -> 0 | ISome n -> n` → 42. Added 6 tests (938 passing). guard / polymorphic / recursive / nested pattern / or-pattern continue to be Codegen_error (future slices).
- Phase 6 #5: Wasm codegen record (monomorphic) — Same linear
memory layout as tuple (Phase 6.4). Stores Record_lit (name, fields) in Typer.records.r_fields declaration order (reconstructed even if source field order differs): base = bump → immediately advance bump by 4N (reserve) → write each field via `i32.store offset=i4 → push base. Field_get (inner, fname) pulls index from record name of inner type → i32.load offset=idx4`. `Record_update (base, updates)` allocates new buffer with bump; for each field, writes new value if in updates, else copies from source via `i32.load offset=...`; returns base of new buffer. Functions that take / return record also work naturally (record is also passed as i32 offset; signature unchanged). Verified (wat2wasm + Node.js): `type Pt = { x: int, y: int }; let p = Pt { x = 3, y = 4 } in p.x + p.y` → 7; `{ p | x = 100 }.x .y → 400; record-returning fn let mk = fn x -> Pair { a = x, b = str_len x } in print ((mk "hello").a) → "hello". Polymorphic record / view continue to be Codegen_error (future slices). Added wasm_with_decls test helper. Added 4 tests (932 passing).
- Phase 6 #4: Wasm codegen tuple — Tuple laid out in linear
memory: each element 4 bytes (Mere int / bool / str all in i32 / offset representation). Tuple [e1; e2; ...] construction: base offset = bump; bump += 4N immediately reserves memory area; write each element via `i32.store offset=N4 at base-relative position; finally push base. Important to reserve first — nested tuple or ++ inner emit advances bump further (during implementation, fixed bug where ((1,2), 3) summed to 22 because reserve was after writing). fst / snd builtin dispatched to i32.load offset=0 / offset=4. Tuple-arg / tuple-return functions also work naturally (tuple is i32 offset, no signature change). Verified (wat2wasm + Node.js): let p = (1, 2) in fst p + snd p → 3; let p = ("hello", 42) in print (fst p) → "hello"; ((1, 2), 3) sum → 6; tuple-arg fn sum_pair (10, 20) → 30. Added 5 tests (928 passing).
- Phase 6 #3: Wasm codegen string support — Implemented
architecture for handling strings via Wasm's linear memory. (memory (export "memory") 1) declares 1-page (64 KB) memory + exports; (global $__lang_bump (mut i32) (i32.const N)) is bump pointer for dynamic alloc (mutable global). Str_lit lifted as (data (i32.const offset) "...\00") data segment; wasm_string_escape escapes \HH. fresh_str_offset helper assigns unique offset to each literal; accumulates in str_data_decls ref. $__lang_strlen (block + loop searches null byte) and $__lang_str_concat (2 strlen calls + 2 copy loops + null terminator + bump update) defined inline in WAT (emitted as runtime_helpers in one go). print s delegated to host (Node.js) via host import (import "env" "puts" (func $puts (param i32))); value is i32 0; Node.js side accesses memory to decode + console.log. str_len s dispatched to call $__lang_strlen; ++ to call $__lang_str_concat. Functions taking / returning str also work naturally (Wasm also treats str as i32, so signature unchanged). Verified (wat2wasm + Node.js with puts that decodes memory): str_len "Hello, world!" → 13; str_len ("hello, " ++ "world!") → 13; print "Hello, Wasm!" → "Hello, Wasm!"; let greet = fn name -> "Hello, " ++ name ++ "!" in print (greet "world") → "Hello, world!". Added 9 tests (923 passing).
- Phase 6 #2: Wasm codegen function lifting + recursion — Top-
level let f = fn x -> ... and let rec lifted as (func $f (param i32) (result i32) ...). fn_skel / lift_fn_skels / find_concrete_arrow / resolve_fn_types implemented in codegen_wasm.ml in parallel, same shape as LLVM Phase 5.2. emit_fn_def puts each fn in independent locals/instrs scope: param in slot 0 (Wasm positional locals); let bindings minted as slot 1, 2, ...; local_counter / locals / instrs saved / restored per-fn. Compiled App (Var name, arg) to <arg push> + call $name (only names registered in toplevel_fn_names get direct call). Wasm allows forward reference in same module, so C/LLVM-style forward declaration / mutual recursion special handling not needed. Verified (wat2wasm + Node.js): factorial 10 → 3628800; fibonacci 15 → 610; is_even 7 (mutual recursion) →
- Added 5 tests (914 passing).
- Phase 6 #1: Wasm (WAT) codegen MVP — Started on the third
design target (Wasm). Implemented emit_program : ?main_ty:ty -> Ast.program -> string in new lib/codegen_wasm.ml; emits subset (int / bool / arith / cmp / logic / Neg / If / Let (P_var) / Var / Annot) as WAT (WebAssembly Text format, S-expression form). Wasm is a stack-based VM (different from LLVM's SSA) — each expression pushes operands in sequence, opcode consumes from stack + pushes result. Compiled Bin (op, a, b) to sequential emit_expr a; emit_expr b; <opcode>. If to if (result i32) ... else ... end block. Let (P_var n, value, body) to combination of (local i32) (fresh slot assignment) + local.set N + local.get N. Comparison via i32.lt_s / i32.gt_s / i32.eq etc.; bool widened to i32 (i32.const 0/1); Neg expressed as 0 - x. main function emitted as (func $main (export "main") (result i32)); local decls consolidated at function head. Added -w <file> / -we <expr> flags to CLI; infer_program helper shared across 3 backends (C/LLVM/Wasm). Verified (via wat2wasm .wasm binary + Node.js WebAssembly.instantiate): let a = 10 in let b = 20 in if a + b > 25 then a * b else 0 → 200; if 3 > 2 then 100 else 200 → 100; let x = 5 in x * x + 1 → 26; true && (false || true) → 1. Added 14 tests (909 passing). Functions / strings / record / variant / closure / region etc. in subsequent slices of Phase 6.
- Phase 5 #14: LLVM IR codegen `'a list` show special-case
([a, b, c] form) — Equivalent to C codegen Phase 4.16. Special-cases `TyCon ("list", [elem_ty])` (when recursive list) before variant branch in `emit_show_fn`: scans from head with alloca/load/store + loop blocks (`loop_test` / `loop_body` / `loop_iter` / `loop_end`); stringifies each element via `show_<elem_tag>`; concats with `", "` between via `__lang_str_concat`; appends `"]"` at end. Pre-registers `@.s_lbracket` ("["), `@.s_rbracket` ("]"), `@.s_comma_space` (", "). Side: (1) `add_show_type` registers in `mono_variant_instances` / `mono_record_instances` when encountering polymorphic TyCon (struct typedef emitted for cases like `show (Nil : int list)` where mono instance can't be collected via Constr); (2) `collect_tuple_shapes` end walks substituted payload of mono variant instances (emits tuple shape of Cons payload `(int, int list)` of `int list` even without Cons in AST); (3) moved `collect_show_types` before typedef emission (so instance flow propagates correctly). Verified (clang native): `show [1, 2, 3]` → `[1, 2, 3]`; `show (Nil : int list)` → `[]`; `show ["hello", "world"]` → `["hello", "world"]`. Added 4 tests (895 passing). Phase 5 (LLVM backend) covers all C codegen (Phase 4) features — int / fn / str / tuple / record / variant / closure / region / poly / recursive variant / complex pattern / show / all memory model / list pretty-print.
- Phase 5 #13: LLVM IR codegen Region_block + Ref + with Drop +
view construction + Unit_lit — Implemented all Mere memory-model features in LLVM backend in one slice; equivalent to C codegen Phase 4.17 user-side region + 4.18 with Drop + 4.19 view construction. `current_regions : (name * register) list ref` tracks region scope. Compiled `Region_block (R, body)` to `alloca %__lang_region` + `__lang_region_init(ptr, 1MB)` + body + `__lang_region_free`. Compiled `Ref (R, v)` (`&R v`) to inner evaluation + sizeof (`getelementptr null` + `ptrtoint`) + `__lang_region_alloc` + `store` to write to region buffer; ptr return. `With (c, v, body)`: `let c = v` + body evaluation; after body, if v's record has `close: unit -> unit` field, auto-invokes via `c.close.fn(c.close.env, 0)` (`extractvalue` separates closure value → env/fn + call). At `Record_lit`, if `name in Typer.views`, view construction: get region name from `e.Ast.ty`'s `TyCon (V, [TyRef (R, ...)])` → build record value with `insertvalue` in declaration order → place in region with `__lang_region_alloc` + `store` → ptr return. At `Field_get`, if inner type is `is_view_type`, `getelementptr %V, ptr %x, i32 0, i32 idx` + `load` to get field. Added `TyRef _ → ptr` and `TyCon (n, _) when Typer.views n → ptr` to `llvm_ty_of`. `Unit_lit` emitted as `i32 0` (needed for `fn () -> ()`). Verified (clang native): `region R { let x = &R 5 in 42 }` → 42; `region R { let pair = &R (1, 2) in 99 }` → 99; `type Pt = { x: int }; region R { let p = &R Pt { x = 42 } in 100 }` → 100 (record also placeable in region); `drop type Conn = { id, close }; with c = mk 7 in c.id * 10` → "close 7\n70" (close called correctly at scope end); `view Cell[R] of int { v: int }; region R { let c = Cell { v = 7 } in c.v }` → 7. Added 7 tests (891 passing). LLVM backend covers all memory- model features, on par with C backend (Phase 4.21).
2026-06-17
- Phase 5 #12: LLVM IR codegen show general builtin — LLVM
version of C codegen Phase 4.12. Specializes show : 'a -> str per-call from arg type's show_<ty_tag>; generates dedicated function for each type. Added @asprintf(ptr, ptr, ...) to runtime_decls. show_types Hashtbl + collect_show_types walks AST to find App (Var "show", arg); add_show_type recursively registers arg type + dependent types (tuple elem / record field / variant payload), with Hashtbl guard so recursive variant 'a list etc. doesn't infinite-loop. emit_show_fn emits specialized fn per type: int → @asprintf("%d", x); bool → select i1 for @.s_true / @.s_false; str → @asprintf("\"%s\"", x); unit → const @.s_unit; tuple → call each element show_T → @asprintf("(%s, ..., %s)", ...); record (mono / poly) → each field show + @asprintf("Type { f = %s, ... }", ...); variant (mono / poly / recursive) → tag dispatch (icmp eq + br + phi) → each ctor: @.s_ctor_<name> direct if nullary; recursive payload show + @asprintf("Ctor %s", ...) if payload. Format strings and ctor name strings pre-registered at start of emit_program for what is needed (mint_show_global / mint_show_format helpers). App (Var "show", arg) dispatched to call ptr @show_<ty_tag arg.ty>(arg). Verified (clang native): show 42 → "42"; show "hi" → "\"hi\""; show true → "true"; show (1, "hi") → (1, "hi"); show (SS 42) → "SS 42"; show (Pt { x = 3, y = 4 }) → Pt { x = 3, y = 4 }; show (Cons (1, Cons (2, Cons (3, Nil)))) → Cons (1, Cons (2, Cons (3, Nil))). Added 9 tests (884 passing). 'a list special-case [1, 2, 3] form (equivalent to Phase 4.16) in future slice.
- Phase 5 #11: LLVM IR codegen complex patterns (P_int / P_str /
P_bool / P_unit / P_record / P_as / nested / or / guard) — LLVM version of C codegen Phase 4.14 + 4.15. Rewrote `compile_pat` as fully recursive `(test_cond, bindings, var_types)` function: P_int → `icmp eq i32`; P_bool → `icmp eq i1`; P_str → `@strcmp(ptr, ptr)` + `icmp eq i32 result, 0`; P_unit → constant `1`; P_record → declared field order `extractvalue` + sub-pattern recurse; P_as → inner pattern + whole value bind; P_tuple → each element `extractvalue` + recurse; P_constr → tag test + sub-pattern recurse (payload via GEP+load if recursive variant, else extractvalue). Multiple sub-tests chained via `and_cond` helper with `and i1`. Or-patterns pre-flattened with `expand_or` (typer guarantees both branches' bound names match, body duplicable). Guard evaluated in arm's bindings scope; if true → body, if false → next_label (= try next arm). Added `@strcmp` to runtime_decls. Verified (clang native): `match 3 with | 0 -> 100 | 1 -> 200 | _ -> 300` → 300; `match "hello"` str match → 2; `match Cons (SS 5, Nil)` nested ctor → 5; `match Pt { x=3, y=4 } with | Pt { x=a, y=b }` → 7; `(a, b) as p` → 6 (`P_as`); `match LCgB with | LCgA | LCgB -> 1 | LCgC -> 2` → 1 (or); `match 7 with | n when n < 5 -> 100 | n when n < 10 -> 200 | _ -> 300` → 200 (guard). Added 8 tests (875 passing).
- Phase 5 #10: LLVM IR codegen recursive variant + P_tuple sub-
pattern — Switched variants with self-referential payload (`type ilist = INil | ICons of int * ilist`, `'a list = Nil | Cons of 'a * 'a list`) to heap-allocated node + ptr representation. `recursive_variants` set + `variant_is_recursive` / `mono_variant_is_recursive` helpers for judgment. Populated in 2 stages within emit_program: at decl registration (source-level) + at mono instance collection (substituted). `emit_variant_typedef` / `emit_mono_variant_typedef` emit `%V_node = type { i32, T }` (on-heap node) if recursive; `llvm_ty_of` returns `ptr` if name in recursive_variants, so value type is transparent ptr. `Constr` recurse: `__lang_region_alloc` allocates node in default region; `getelementptr` + `store i32 tag` + `getelementptr` + `store T payload` write; ptr return. `Match` recurse: get tag from scrutinee ptr via `getelementptr` + `load i32`; payload of each arm similarly via `load`. In pattern compile, expand `P_tuple` sub-pattern (`Cons (h, t)`) into chain of `extractvalue` of payload tuple struct; bind each element to fresh register. `pattern_var_types` helper adds concrete types of pattern bind names to current_var_types (so polymorphic recursive calls don't leave `'a list` as-is). Match scrutinee type fallback to current_var_types if Var; same for direct-call App arg type. Reordered typedef emission to `collect_mono_instances` + recursive judgment → tuple/record/variant typedef emit (so recursive_variants state affects tuple emit). Verified (clang native): `type ilist = INil | ICons of int * ilist; sum (ICons (1, ICons (2, ICons (3, INil))))` → 6; `type 'a list = Nil | Cons of 'a * 'a list; sum [1,2,3,4,5]` → 15; `length ["a","b","c","d"]` → 4 (poly recursive list). Added 5 tests (867 passing).
- Phase 5 #9: LLVM IR codegen monomorphization of polymorphic
variant / record — C codegen Phase 4.11 + 4.13 implemented on LLVM side in one slice. `polymorphic_variants` / `polymorphic_records` Hashtbl defer declarations (walk `Exhaustive.type_variants` + `Typer.records` at start of emit_program; register only poly ones); recover poly variant param names via constructor's `params`. `mono_variant_instances` / `mono_record_instances` accumulate found instances; `collect_mono_instances` walks AST + fn signature to find `(name, args)`. `subst_params` / `subst_variants` substitute type vars → concrete types; `mono_variant_name n args` / `mono_record_name n args` produce specialized names (`opt_int`, `Box_str` etc.). `emit_mono_variant_typedef` determines payload type from substituted payload type via `variant_payload_ty_of`; emits `%opt_int = type { i32, T }`. `emit_mono_record_typedef` emits `%Box_int = type { ... }` with substituted field types. `llvm_ty_of (TyCon (n, args))` maps to mono name if name in `polymorphic_variants/records`. `Constr` emit: pull mono name from `e.ty`; determine payload type with `variant_payload_ty_of`. `Record_lit` / `Field_get` / `Record_update` similarly use `mono_record_name` + substituted fields for poly records. `Match` scrutinee type, if poly, uses mono name + substituted variants. Verified (clang native): `type 'a LCgOpt = LCgN | LCgS of 'a; match LCgS 42 with | LCgN -> 0 | LCgS n -> n` → 42; `type 'a Box = { v: 'a }; let b = Box { v = 42 } in b.v` → 42; specialize both types `let bi = Box { v = 42 } in let bs = Box { v = "hi" } in str_len bs.v + bi.v` → 44 (both `%Box_int` and `%Box_str` emitted). Added 7 tests (862 passing). Recursive poly variant (`'a list`) requires recursive variant support → Phase 5.10.
- Phase 5 #8: LLVM IR codegen default region runtime + closure/
string alloc via region — Implemented work equivalent to C codegen Phase 4.17 + 4.20 + 4.21 on LLVM side in one slice. `%__lang_region = type { ptr, ptr, i64 }` struct + `@__lang_default_region = internal global %__lang_region zeroinitializer` file-scope global + 3 helper functions `__lang_region_init/alloc/free` defined inline in LLVM IR (`region_runtime_helpers`). `__lang_region_alloc` uses 8-byte aligned bump pointer (`(n + 7) & -8` implemented with `and i64 ..., -8`; advances top via gep i8, store). Calls `__lang_region_init(@__lang_default_region, 4194304)` (4 MB) at `@main` entry; calls `__lang_region_free` before final `ret i32 0`. Replaced `malloc` in `__lang_str_concat` with `__lang_region_alloc(@__lang_default_region, ...)`; closure env (anonymous Fun) `malloc(sizeof)` similarly replaced. Added `@free` to runtime_decls; inserted region_runtime_helpers in emit order right before str_concat_helper. Verified (clang native): `(make_adder 5) 10` → 15; `compose inc dbl 5` → 11; concat like `"hello, " ++ "world"`; only `malloc` call in generated IR is one spot inside region init (one-shot free at program end, valgrind clean). Added 8 tests (855 passing). LLVM backend memory model reached the same level as C backend (Phase 4.21).
- Phase 5 #7 Phase B: LLVM IR codegen anonymous Fun + closure-
with-captures — Handles internal `fn x -> ...` in expression position. Computes free variables of AST with `free_vars` / `pattern_vars` helpers (excluding bound names, preserving order); filters to those registered in `current_var_types` (excluding top- level / builtin). For each capture, gets concrete type from `current_var_types`; generates `%anon_N_env = type { T1, T2, ... }` env struct typedef; pushes `anon_N_fn` adapter to `pending_closures` queue. At construction site: allocates env with `malloc(sizeof(%anon_N_env))` (LLVM `getelementptr null` trick + `ptrtoint` calculates size); writes each capture to env field via `getelementptr %env, ptr %p, i32 0, i32 idx` + `store`; assembles closure value with `insertvalue %closure undef, ptr %env, 0` + `insertvalue ..., ptr @anon_N_fn, 1`. `emit_anon_adapter` invoked by `emit_program` draining `pending_closures` (iterative loop also processes new pending added during drain); in adapter body entry block, pulls each capture from env_self with `getelementptr` + `load` into fresh register, binds to env, then emits original Fun body. Added `current_expected_ty : ty option ref`: lets parent context type serve as fallback when AST's Fun.ty is polymorphic (resolves cases where inner Fun's type stays `'a -> 'a` in let-poly curried polymorphic HOFs like `fn f -> fn x -> f (f x)`; emit_fn_def / emit_anon_adapter set return_ty at body start / restore at end). Extended Let case to add value type to current_var_types (so closure can capture variables of outer let). Verified (clang native): `let make_adder = fn n -> fn x -> x + n in (make_adder 5) 10` → 15 (capture); `let twice = fn f -> fn x -> f (f x) in twice inc 5` → 7 (curried HOF + polymorphic); `let apply = fn f -> fn x -> f x in apply (fn n -> n * 3) 7` → 21 (anon Fun passed as arg); `let compose = fn f -> fn g -> fn x -> f (g x) in ((compose inc) dbl) 5` → 11 (3-level nested closure + 2 captures). Added 7 tests (847 passing). env currently leaks via `@malloc` — default region-ization in future slice.
- Phase 5 #7 Phase A: LLVM IR codegen first-class top-level fn —
Lowers T1 -> T2 type as %closure_T1_T2 = type { ptr, ptr } (env, fn pointer). closure_struct_name helper for closure type name; collect_arrow_types walks AST + fn signatures to gather all used arrow types; emit_closure_typedef generates typedef. Auto-generates env-ignoring adapter define T2 @<name>_closure_fn (ptr %env_unused, T1 %x) { ret T2 @<name>(T1 %x); } for each top-level fn (emit_closure_adapter). At emit_expr Var name: if no shadowing in env and registered in toplevel_fn_names, inline-constructs closure value with insertvalue %closure undef, ptr null, 0 + insertvalue ..., ptr @<name>_closure_fn, 1. App f arg: existing direct-call path preserved (for known top-level fn); otherwise dispatches via extractvalue %closure %c, 0/1 to get env/fn, then call T2 %fn_ptr(ptr %env, T1 %arg) (no fn pointer type cast needed via opaque pointer). Added current_var_types : (string * ty) list ref: for polymorphic Var in fn body (parameter staying as 'a -> int after let-poly), can pull concrete type from resolve_fn_types-derived (sets param's concrete ty at start of emit_fn_def, save/restore). Verified (clang native): let inc = fn x -> x + 1 in let apply = fn f -> f 5 in apply inc → 6; let apply2 = fn f -> f (f 5) in apply2 inc → 7. Added 7 tests (840 passing). Anonymous Fun (inner fn x -> ...) and closure-with-captures (Phase B) in separate slice.
- Phase 5 #6: LLVM IR codegen variant + match (monomorphic, single
payload type) — Lowers monomorphic variant to LLVM named struct: if all ctors nullary, `%V = type { i32 }`; if payload exists, `%V = type { i32, T }` (`variant_payload_ty` detects single payload type shared by all payload-bearing ctors; Codegen_error if differ). `variant_tags` Hashtbl holds constructor → int tag; set as side effect of `emit_variant_typedef`. `collect_variant_names` walks AST + fn signature + Constr's type_name to gather used variant types (only `Typer.types` arity 0 ones). `Constr cname arg_opt` → `%t0 = insertvalue %V undef, i32 tag, 0` → optional `%t1 = insertvalue %V %t0, T arg, 1` chain constructs SSA struct value. `Match` gets scrutinee's tag with `extractvalue %V %s, 0`; tests each arm sequentially with `icmp eq i32 %tag, N` + `br i1`; fallthrough is `@abort()` + `unreachable`; merges all arm results with `phi <result_ty>` at end. Pattern is P_constr / P_var / P_wild only; payload bind creates payload register with `extractvalue %V %s, 1` and adds to bindings. Added @abort declaration to runtime_decls. Verified (clang native): `type Color = R | G | B; match G with | R -> 0 | G -> 1 | B -> 2` → 1; `type Status = Ok | Err of str; match Err "boom" with | Ok -> 0 | Err m -> str_len m` → 4; `type IntOpt = INone | ISome of int; let v = ISome 42 in match v with | INone -> 0 | ISome n -> n` → 42. Added 9 tests (833 passing). Guard / polymorphic variant / recursive variant / nested pattern / or-pattern continue to be Codegen_error.
- Phase 5 #5: LLVM IR codegen record (monomorphic) — Lowers
monomorphic record (type Pt = { x: int, y: int }) to LLVM named struct (%Pt = type { i32, i32 }). Added TyCon (name, []) when Hashtbl.mem Typer.records name -> "%" ^ name to llvm_ty_of; record_fields / field_index helpers pull declaration-order fields from Typer.records. Record_lit emit constructed with insertvalue chain in declaration order (even if source field order differs from declared, pulls values with List.assoc_opt and stacks in declaration order). Field_get is extractvalue %R %p, idx; Record_update starts from base value and stacks each update field via insertvalue. collect_record_names walks AST + fn signature to gather all used record types (polymorphic records excluded for now, separate slice). emit_record_typedef generates %Name = type { T1, T2, ... }. Via bin/main.ml infer_program helper, so Typer.records is already populated. Added llvm_with_decls test helper (parallel to codegen_with_decls). Verified (clang native): type Pt = { x: int, y: int }; let p = Pt { x = 3, y = 4 } in p.x + p.y → 7; Record_update { p | x = 100 } x y → 400; record-returning fn `let mk = fn x -> Pair { a = x, b = str_len x } in print ((mk "hello").a)` → "hello". Added 6 tests (824 passing). Polymorphic record (`type 'a Box`) stays Codegen_error.
- Phase 5 #4: LLVM IR codegen tuple — Lowers tuple to LLVM named
struct (%tuple_int_str = type { i32, ptr }). ty_tag / tuple_struct_name helpers (same naming convention as codegen_c generates symbols like tuple_int_str); collect_tuple_shapes walks AST + fn signature to gather all used tuple types; emit_tuple_typedef generates %name = type { T1, T2, ... }. Tuple node emit constructs struct value in SSA register with insertvalue chain (starts from undef, stacks each element via insertvalue %T %prev, Tn vn, idx). fst / snd builtin compiled to extractvalue %tuple_X %p, 0/1 (struct name resolved from arg's .ty). llvm_ty_of (TyTuple ts) returns %<tuple_struct_name>, so tuple-arg / tuple-return function signatures automatically take correct form (define %tuple_int_int @split(ptr %s), define i32 @sum_pair(%tuple_int_int %p)). Nested tuple (((1, 2), 3) → %tuple_tuple_int_int_int = type { %tuple_int_int, i32 }) auto-generated. Verified (clang native): let p = (1, 2) in fst p + snd p → 3; let p = ("hello", 42) in print (fst p) → "hello"; let split = fn s -> (s, str_len s) in print (fst (split "hello")) → "hello"; nested tuple ((1,2), 3) sum → 6; tuple-arg fn sum_pair (10, 20) → 30. Added 8 tests (818 passing).
- Phase 5 #3: LLVM IR codegen strings + print + ++ + str_len +
str-taking/returning functions — Maps `TyStr` to LLVM `ptr` (opaque pointer). Lifts `Str_lit s` as private constant global `@.str_N = private constant [N x i8] c"...\00"`; generated via `fresh_str_global` helper; escapes non-printable ASCII with `\HH`. Uses global symbol directly as ptr for value (no GEP needed with opaque pointer). Compiles `Bin (Concat, a, b)` to `call ptr @__lang_str_concat(ptr %a, ptr %b)`; `__lang_str_concat` defined inline in LLVM IR (combination of `malloc` + `strlen` + `memcpy` + GEP + `store i8 0`). Compiles `print` builtin to `call i32 @puts(ptr %s)` (discards return value; Mere value is 0); `str_len` to `call i64 @strlen(ptr %s)` + `trunc i64 ... to i32`. Added `TyStr → ("ptr", "%s")` to `main_format_of`; generates `@.fmt_s = c"%s\\0A\\00"` global. str-taking/returning functions auto-lowered correctly (`define ptr @f(ptr %s)`). Runtime helpers (`declare ptr @malloc(i64)` etc.) and `__lang_str_concat` body emitted in emit_program in one go; `.ll` file is self-contained. Verified (clang native): `print "Hello, LLVM!"` → "Hello, LLVM!"; `"hello, " ++ "world!"` → "hello, world!"; `str_len "Hello, world!"` → 13; `let greet = fn name -> "Hello, " ++ name ++ "!" in print (greet "world")` → "Hello, world!"; `let exclaim = fn s -> s ++ "!" in print (exclaim "wow")` → "wow!"; `let pick = fn n -> if n > 0 then "positive" else "negative" in print (pick 5)` → "positive". Added 10 tests (810 passing).
- Phase 5 #2: LLVM IR codegen function lifting + recursion —
Lifts top-level let f = fn x -> ... and let rec f = ... and g = ... as LLVM define iXX @f(iYY %x) { ... }. Implemented fn_skel / lift_fn_skels / find_concrete_arrow / resolve_fn_types in codegen_llvm.ml in parallel, same shape as C codegen (combined with LLVM-specific llvm_ty_of). emit_fn_def emits each function as independent SSA scope (reg_counter / label_counter reset per-function; instrs save/restore). At App (Var name, arg), if name is registered in toplevel_fn_names, compiled to %t = call iZZ @name(iYY %arg) (closure-as-value in future slice). LLVM IR allows forward reference within same module, so forward declaration needed in Phase 4 is unnecessary (mutual recursion works as-is). Verified (clang native): factorial 10 → 3628800; fib 15 → 610; is_even 7 (mutual recursion) → 0. Added 6 tests (800 passing).
- Phase 5 #1: LLVM IR codegen MVP — Started second backend that
compiles Mere to native binary. Implemented emit_program : ?main_ty:ty -> Ast.program -> string in new lib/codegen_llvm.ml; converts subset (int / bool / arith / cmp / logic / Neg / If / Let (P_var) / Var / Annot) to LLVM textual IR. Hand-written text generation (no dependency on opam's llvm package; directly compile with clang out.ll). Name management via SSA register counter (%t0, %t1 ...) and basic block label counter; If goes through br i1 + label/phi; comparison via icmp slt/sgt/eq/...; bool computed in i1 and zext-extended to i32 at main end for output via @printf (@.fmt_d = c"%d\\0A\\00"). Added -ll <file> / -lle <expr> flags to CLI; shared infer_program helper for both C / LLVM backends. Verified (clang native execution): let a = 10 in let b = 20 in if a + b > 25 then a * b else 0 → 200; if 3 > 2 then 100 else 200 → 100; let x = 5 in x * x + 1 → 26; true && (false || true) → 1. Added 15 tests (794 passing). Functions / strings / record / variant / closure / region etc. now Codegen_error (same scope as Phase 4 MVP).
- Phase 4 #21: strings + recursive variant nodes also moved to
default region — Unifies remaining 2 malloc sites under `__lang_default_region`. Replaced `malloc(la + lb + 1)` in `__lang_str_concat` runtime helper with `__lang_region_alloc (&__lang_default_region, la + lb + 1)`; replaced `malloc(sizeof(T_node))` in recursive variant Constr emit (self-referential variant like `Cons (h, t)`) with `__lang_region_alloc(&__lang_default_region, sizeof(T_node))`. Reordered helper ordering in `emit_program` to `region_runtime_helpers → str_concat_helper` so str_concat helper can reference `__lang_default_region` symbol (ordering issue). Now the only remaining malloc on C side is base buffer allocation inside `__lang_region_init`; all user-visible alloc sites ride on bump arena. Batch free with `__lang_region_free(&__lang_default_region)` at `main` end; valgrind clean. Verified (clang native): `let greet = fn name -> "Hello, " ++ name ++ "!" in print (greet "world")` → "Hello, world!"; `sum [1, 2, 3, 4, 5]` → 15 (Cons of recursive list all in region alloc). Added 2 tests + updated 1 (779 passing; renamed "Constr mallocs node" to "Constr uses default region").
- Phase 4 #20: closure env moved to default region — Added
program-lifetime arena __lang_default_region at file scope (static __lang_region __lang_default_region;); calls __lang_region_init(&__lang_default_region, 1 << 22) (4MB) at start of main, __lang_region_free at end. Switched anonymous closure env struct alloc from malloc(sizeof(...)) to __lang_region_alloc(&__lang_default_region, sizeof(...)). Closures can outlive user's region R { ... } (carried out like make_adder 3 |> add3 4), so don't coexist with user region; needed to be in separate program-lifetime arena. Per-closure malloc cost gone; batch-freed at main end; valgrind also clean. Verified (clang native): let make_adder = fn n -> fn x -> n + x in let add3 = make_adder 3 in add3 4 → 7; let compose = fn f -> fn g -> fn x -> f (g x) in compose (fn n -> n + 1) (fn n -> n * 2) 5 → 11 (nested closure with captures all in default region). Remaining leaks: string concat (++) and recursive variant node (Cons). Added 5 tests + assert_no_contains helper (777 passing).
- Phase 4 #19: region-izing view construction — Codegen places
view V[R] of T { ... } on region's bump allocator. View value represented in C as V* (pointer type); at construction, allocates in region via __lang_region_alloc(&__region_R, sizeof(V)), copies content, returns pointer. c_type_of (TyCon (V, [TyRef R TyUnit])) -> V*; is_view_type helper distinguishes record / view; Field_get uses -> for view value. View value's lifetime matches region scope (combined with Phase 2.1 escape check + Phase 4.17 region runtime) — memory model's view feature works fully at runtime level. Verified (clang native): view Cell[R] of int { v: int }; region R { let c = Cell { v = 7 } in c.v } → 7. Added 3 tests (772 passing; added Top_view handling to codegen_with_decls helper).
- Phase 4 #18: `with` Drop execution codegen + typedef ordering
cleanup — C codegen for `with c = v in body`: at scope end, auto-calls c's `close` field via `c.close.fn(c.close.env, 0)` (only when `close: unit -> unit` field exists in Drop type; skip if absent. Multiple `with` are nested in AST, so naturally LIFO). Side: reorganized typedef structure to "all forward decls → closure typedefs → all struct bodies". Logic: for cases where record has `closure_T1_T2` type like `close: unit -> unit` field of Drop type, closure typedef needs record's full definition as function-pointer return; but C can use forward-declared struct as function pointer return type, so closure typedef can be emitted if forward decls come first. Split all variant / record / tuple typedefs into 2 stages of forward decl + body; reorder them in emit_program. Verified (clang native): `drop type Conn = { id: int, close: unit -> unit }; let mk = fn id -> Conn { id = id, close = fn () -> print ("close " ++ show id) } in with c = mk 7 in c.id * 10` → "close 7\n70" (close called correctly at scope end). Added 3 tests + updated 6 typedef snapshots to new format (769 passing).
- Phase 4 #17: region runtime (bump allocator) — Codegen
region R { body } as a real bump allocator. Added new C runtime helper __lang_region ({ char* base; char* top; size_t cap; }) + __lang_region_init/alloc/free injected into generated source. emit_expr Region_block outputs statement expression ({ __lang_region __region_R; __lang_region_init (&__region_R, 1<<20); __auto_type __r_result = body; __lang_region_free(&__region_R); __r_result; }). emit_expr Ref (R, v) emits ({ __auto_type __ref_v = v; typeof(__ref_v)* __p = __lang_region_alloc(&__region_R, sizeof __ref_v); *__p = __ref_v; __p; }) (bump alloc + copy + return pointer in region). c_type_of (TyRef _ inner) to inner*. Combined with escape check (typer), memory is batch-freed on region scope exit, but type signature guarantees &R T doesn't leak (Phase 2.1 escape check) for safety. Milestone where memory model went from "type level label" to "real bump allocator". Verified (clang native): region R { let x = &R 5 in 42 } → 42; region R { let pair = &R (1, 2) in 99 } → 99; type Pt = { x: int }; region R { let p = &R Pt { x = 42 } in 100 } → 100 (record also placeable in region). Added 5 tests (766 passing).
- Phase 4 #16: `'a list` show in `[a, b, c]` form + variant
payload tuple shape collection — Special-cases `TyCon ("list", [elem_ty])` in `emit_show_fn`; generates specialized function that strings the whole list with a while loop (`"[]"` if Nil; `[1, 2, 3]` format if Cons; matches Mere interpreter output). Side: extended tuple shape collection to include mono variant payload (`tuple_int_list_int` etc. referenced even in cases like `show ([] : int list)` that doesn't include Cons construction; fixed build failure where necessary struct typedef wasn't emitted). Verified (clang native): `show [1, 2, 3]` → `[1, 2, 3]`; `show ["hello", "world"]` → `["hello", "world"]`; `show ([] : int list)` → `[]`. Added 2 tests (761 passing).
- Phase 4 #15: C codegen or-pattern + match guard — Flattens
| pat1 | pat2 -> body into multiple arms via pre-pass expand_or of Match emit (constraint that both branches bind same name set guaranteed by typer). Body is duplicated to both but safe as pure expression. when ... guard evaluated in arm's bindings scope; falls through if false (test ? ({ bindings; guard ? body : next; }) : next). Verified (clang native): type Col = R | G | B; match G with | R | G -> 1 | B -> 2 → 1; match 7 with | n when n < 5 -> 100 | n when n < 10 -> 200 | _ -> 300 → 200. Nested or-pattern (constructor etc. inside or) continues to be Codegen_error. Added 4 tests + updated 1 (759 passing; replaced "guard rejected" with "guard accepted").
- Phase 4 #14: C codegen complex patterns — Rewrote Match
pattern compilation as fully recursive compile_pattern. Decomposes each pattern into (test_expr, bindings_str); supports nesting constructor / tuple / record inside constructor; implements P_int / P_str (strcmp == 0) / P_bool / P_unit / P_record (named field destructure) / P_as (whole-value bind). is_ptr_ty / payload_ty_for_ctor / field_ty helpers resolve sub-value types and recursively decompose patterns. Verified (clang native): match 3 with | 0 -> 100 | 1 -> 200 | _ -> 300 → 300; match "hello" with | "hi" -> 1 | "hello" -> 2 | _ -> 3 → 2; match Cons (Some 5, Nil) with | Nil -> 0 | Cons (None, _) -> 1 | Cons (Some n, _) -> n → 5 (nested poly variant); match Point { x = 3, y = 4 } with | Point { x = a, y = b } -> a + b → 7. Or-pattern and guard continue to be Codegen_error. Added 6 tests + updated 4 substrings to new format (755 passing).
- Phase 4 #13: C codegen polymorphic record monomorphization —
Specializes type 'a Box = { v: 'a } etc. polymorphic records per type (Box_int, Box_str etc.) using same pattern as variant's Phase 4.11. polymorphic_records Hashtbl defers declarations (emit_record_typedef defers if r_params != []); extends collect_mono_variant_instances to also cover records; emit_mono_record_typedef concretizes field types with subst_params and generates typedef struct { int v; } Box_int;. Record_lit emit pulls mono name from Record_lit's .ty and emits compound literal (((Box_int){.v = 42})). Field_get and Record_update naturally work via __auto_type. Verified (clang native): type 'a Box = { v: 'a }; let b = Box { v = 42 } in b.v → 42; let bi = Box { v = 42 } in let bs = Box { v = "hi" } in show (bi.v, bs.v) → (42, "hi") (specializes both Box_int and Box_str). Added 3 tests + updated 1 (749 passing; replaced "polymorphic record reject" with "specialize verification").
- Phase 4 #12: C codegen `show` general builtin — Auto-generates
per-type specialized show_T C functions for show : 'a -> str by collecting per-call arg types from AST. collect_show_types finds App (Var "show", arg); add_with_deps recursively registers types arg type depends on (tuple elem / record field / variant payload) (with cycle guard; doesn't infinite-loop on self-referential payload of recursive variant). emit_show_fn generates specialized fn per type — int/bool/str/unit trivial; tuple/record composes element show; variant (mono + polymorphic instantiation + recursive) is tag dispatch + payload show. emit_expr App's Var "show" dispatches to show_<tag>(arg) call resolved by arg type's ty_tag. Verified (clang native): show 42 → "42"; show (1, "hello") → (1, "hello"); show (Some 42) → "Some 42"; show [1, 2, 3] → "Cons (1, Cons (2, Cons (3, Nil)))". Based on asprintf (malloc leak but consistent with other codegen). Added 7 tests (747 passing).
- Phase 4 #11: C codegen polymorphic variant monomorphization
— Implemented monomorphization that collects concrete instantiations from AST and fn signatures for type 'a opt = None | Some of 'a or type 'a list = Nil | Cons of 'a * 'a list etc. polymorphic variants and emits specialized struct (opt_int, list_int etc.) per instance. polymorphic_variants Hashtbl defers declarations; mono_variant_instances accumulates found instances; subst_params / subst_variants for param→arg substitution; mono_variant_is_recursive for recursion judgment on concrete types. Extended c_type_of and ty_tag to handle TyCon (n, args) with args (int list → list_int etc.). Constr emit pulls mono name from Constr's .ty; Match's is_ptr judgment also recursion-checks with mono name. Verified (clang native): type 'a opt = None | Some of 'a; let v = Some 42 in match v with | None -> 0 | Some n -> n → 42; type 'a list = Nil | Cons of 'a * 'a list; let rec sum = fn xs -> match xs with | Nil -> 0 | Cons (h, t) -> h + sum t in sum [1, 2, 3] → 6 (list literal + recursive sum; [1, 2, 3] is parser-desugared to Cons (1, Cons (2, Cons (3, Nil)))). Added 4 tests (740 passing).
- Phase 4 #10: C codegen recursive variant + P_tuple pattern —
Switched variants with self-referential payload (e.g. type ilist = INil | ICons of int * ilist) to heap-allocated node + ptr typedef (typedef struct ilist_node ilist_node; typedef ilist_node* ilist; struct ilist_node { ... };). variant_is_recursive detects self-reference in payload; registers in recursive_variants Hashtbl. Constr emit malloc-returns ptr with ({ ilist_node* __p = malloc(...); __p->tag = N; __p->payload.CTOR = ...; __p; }). Match emit switches . vs -> based on scrutinee's type. Expands P_tuple sub-pattern (CgCons (h, t)) into .f0 / .f1 binding sequence. Circular typedef dependency resolved by emitting forward decl + ptr typedef first, then struct body after tuple/record typedefs. Verified (clang native): type ilist = INil | ICons of int * ilist; let rec sum = fn xs -> match xs with | INil -> 0 | ICons (h, t) -> h + sum t in sum (ICons (1, ICons (2, ICons (3, INil)))) → 6 (linked list sum). Added 5 tests (736 passing).
- Phase 4 #9 Phase B: C codegen anonymous Fun + closure-with-
captures — Lifts anonymous Fun in expression position as heap-allocated env struct + adapter + closure construction. Capture vars rewritten to `__env_self->name` via `current_env_subst` map; capture types resolved by traversing scope via `current_var_types` (workaround for polymorphic residual problem after let-poly). Closure typedefs emitted in inner→outer order (post-order walk) to avoid circular references. `current_expected_ty` passes context type to Fun emit; estimates inner Fun's type from outer fn's return_ty. Verified (clang native): `let apply = fn f -> fn x -> f x in let inc = fn n -> n + 1 in apply inc 5` → 6 (curried HOF); `let twice = fn f -> fn x -> f (f x) in twice inc 5` → 7; `let make_adder = fn n -> fn x -> x + n in (make_adder 5) 10` → 15 (closure with capture). Added 4 tests + updated 1 (731 passing).
- Phase 4 #9 (Phase A): C codegen first-class functions —
Represents T1 -> T2 type function value as C struct closure_T1_T2 = { void* env; T2 (*fn)(void*, T1); }. Auto- generates env-ignoring adapter (f_closure_fn) + value const (f_as_value) for each top-level fn. c_type_of (TyArrow ...) maps to closure struct name; ty_tag also handles nesting. emit_expr Var: at value position if name is top-level fn, emit f_as_value (Codegen_error if using inner-lifted in value position). emit_expr App: known top-level Var call continues on direct call fast path; otherwise dispatches via closure ({ __auto_type __c = e; __c.fn(__c.env, arg); }). collect_arrow_types walks AST + fn signatures to gather arrow types and auto-generates closure typedefs. Verified (clang native): let inc = fn x -> x + 1 in let apply = fn f -> f 5 in apply inc → 6 (top-level fn passed as value to HOF works). Phase B (inner / anonymous fn value-ization) in separate slice. Added 6 tests (727 passing).
- Phase 4 #8: C codegen closure conversion (defunctionalization)
— Added pre-pass that lifts let h = fn x -> body in ... inside function body to top-level. free_vars helper computes free variables (excluding builtin / top-level fn names of typer's initial_env); prepends captured variables to C function's param list (defunctionalization). emit_expr Let sees Hashtbl.mem inner_lifts name and skips lifted bindings; App passes capture args at call site. Captures are int/bool/str/unit only (tuple/record/function value capture is Codegen_error). Supports multi-level nesting (h captures x and n from 2 levels). Side: changed resolve_fn_types to pull monomorphic types at call site via find_concrete_arrow for Fun.ty issue after let-poly. Verified: let outer = fn x -> let h = fn y -> x + y in h 10 in outer 5 → 15; nested 2 levels → 6. Added 4 tests + updated 1 (721 passing; replaced old "closure reject" test with "lift result verification").
- Phase 4 #7: C codegen variant + match — Compiles monomorphic
variant types (type Status = Ok | Err of str) to tagged union (typedef struct { int tag; union { const char* Err; } payload; } Status;). Constr to compound literal (((Status){.tag = 1, .payload.Err = "boom"})). Match to ternary chain in statement expression (__scrut.tag == N ? ({ binding; body; }) : ... + fallthrough abort()). Pattern subset: P_constr (nullary or P_var / P_wild sub); P_var; P_wild. Guard / polymorphic variant / nested pattern are Codegen_error. Verified (clang native): type Color = R | G | B; match G with | R -> 0 | G -> 1 | B -> 2 → 1; type Status = Ok | Err of str; match Err "boom" with | Ok -> 0 | Err msg -> str_len msg → 4. Added 9 tests (715 passing).
- Phase 4 #6: C codegen record support — Compiles
type Point
= { x: int, y: int } to typedef struct { int x; int y; } Point;. Implements Record_lit / Field_get / Record_update (Record_update uses ({ __auto_type __rupd = base; __rupd.f = v; __rupd; }) statement expression pattern). collect_record_names walks AST + fn signature to gather used record types and auto-generate typedefs. Extended compile_to_c to include top-level decl processing (same as Pipeline.type_of, skips eval; only record/variant/view/drop registration). Verified (clang native): let p = Point { x = 3, y = 4 } in p.x + p.y → 7; record update → 102; record-returning fn → 15. Polymorphic record (type 'a Box = { v: 'a }) continues to be Codegen_error. Added 7 tests (706 passing).
- Phase 4 #5: C codegen tuple support + AST type annotation
foundation — As foundation, added `mutable ty : ty option` to `Ast.expr`; `Typer.infer` now records inference results on each node. This lets codegen directly reference per-node types. Compiles `Tuple` to C struct (`typedef struct { ... } tuple_int_int;`) + C99 compound literal `((tuple_int_int){.f0 = 1, .f1 = 2})`. Compiles `fst` / `snd` builtin to `.f0` / `.f1` field access. Supports arbitrary element types (int/bool/str + nested tuple); auto-generates struct per shape (`collect_tuple_shapes` walks entire AST + fn signature). Verified (clang native): `let p = (1, 2) in fst p + snd p` → 3; `let p = ("hello", 42) in print (fst p)` → "hello"; `let split = fn s -> (s, str_len s) in print (fst (split "hello"))` → "hello". Added 6 tests (699 passing).
- Phase 4 #4: C codegen: str-taking / returning functions —
Allows lifted function param / return to also use str (const char). Added `param_ty` / `return_ty` to `fn_decl`; `lift_fn_skels` extracts skeletons → `resolve_fn_types` flows all lifted fns to typer as one let-rec group for type inference (handles self / mutual recursion) → `c_type_of` maps Ast.ty to C type (int/bool → `int`, str → `const char, unit → int). Compiles str_len builtin to C's strlen (App special case). Verified (clang native): let greet = fn n -> if n > 0 then "pos" else "neg" in print (greet 5) → "positive"; let exclaim = fn s -> s ++ "!" in print (exclaim "hello") → "hello!"; str_len "hello, world!" → 13. Added 5 tests (693 passing).
- Phase 4 #3: C codegen string support — Compiles
Str_lit
to C string literal; ++ via runtime helper __lang_str_concat (malloc-based); print builtin to puts (statement expression returning int 0). Switched let to GNU/Clang extension __auto_type so same emit works for both int/str values. Made emit_program type-aware (~main_ty); selects printf's format from main's type (int/bool → %d, str → %s, unit → printf skip). Verified: print "hello, world!" → hello, world!; "hello" ++ " " ++ "world" → hello world (all clang native). Malloc leaks (region/GC integration in future slice). Added 6 tests / restructured existing codegen tests as fragment inspection (688 passing).
- Phase 4 #2: C codegen function lifting — Lifts top-level
let f = fn x -> ... and let rec f = fn x -> ... and g = fn y -> ... as C function (with forward declaration). Compiles App (Var name, arg) form direct calls to C name(arg); both self-recursion and mutual recursion work (factorial 10 = 3628800, fibonacci 15 = 610, is_even 7 = 0 confirmed via clang native). Closure (fn ... inside function body) continues to be Codegen_error. Added 5 tests (681 passing).
- Phase 4 #1: C codegen MVP — First step from interpreter to
native. Implemented emit_program : Ast.program -> string in new lib/codegen_c.ml; converts subset of int / bool / arith / cmp / logic / Neg / If / Let (P_var only) / Var / Annot to C expression (let compiled to single C expression via GCC/Clang statement expression ({ ... })). Added -c FILE / -ce <expr> flags to CLI; outputs C source to stdout. clang OUT.c -o BIN && ./BIN for native execution. Functions / strings / record / variant / region / view etc. now Codegen_error. Added 7 tests (677 passing); manual E2E verified via clang (let a = 10 in let b = 20 in if a + b > 25 then a * b else 0 → 200).
- example: examples/pipeline.mere — Realistic example
(~75 lines) combining region / view / effect (builtin Logger / Metrics + cap passing + using sugar) / with Drop. Simple build pipeline: open/close user session with with session = open_session logger uid; process each task with region R { ... }; inside region build view Task[R] to calculate size. Output is session open/close log + per-task [task] log + [METRIC] inc / record + user log + final total. Demonstrates Mere's full feature set working consistently in a practical example.
- Phase 3.1: `with` Drop semantics —
with c = v in body
requires v's type to be a Drop type (declared drop type ...); Trivial value is type error (use let). On eval side, calls v's close: unit -> unit field at scope end (no-op if absent). Multiple with x, y in body close in LIFO order y → x. Rewrote examples/with_caps.mere based on Drop type. Implemented case (i) of design doc 12_drop_and_with.md. Added 6 tests / restructured 6 (670 passing).
- effect: builtin `Logger` / `Metrics` cap types + `mk_logger`
/ mk_metrics constructor builtins — Provides cap types as stdlib. Registered `Logger { info, warn, error: str -> unit }` and `Metrics { inc: str -> unit, record: str -> int -> unit }` in typer; added corresponding V_record constructor functions to eval. Users don't need to redefine cap types each time (overrides allowed). Rewrote examples/effects.mere with builtin usage. Added 7 tests (668 passing).
- effect: `using [cap]` syntax sugar — Desugars
fn x using
[logger] -> body to fn logger -> fn x -> body (caps are outer-most curried args). Eases partial application iteration frequent in cap-passing style (main pattern of Q-003/Q-006 solution). Type annotations allowed; multiple caps allowed; combination with regular params allowed. Implements auxiliary design of design doc 10_effect_trial_findings.md. Added 7 tests (661 passing). Rewrote examples/effects.mere in sugar form too.
- example: examples/effects.mere — Demonstration of
Capability passing pattern (about 75 lines). Declares Logger / Metrics cap types as records; demos 3 patterns: direct use in low-order function / bucket-brigade / partial application passing to high-order function. Demonstrates that design doc 05_effect_system.md's "side effects = passing capability as values" works with current Mere (HM + function args + record + curry) alone — no need for new syntax for effect system.
- region Phase 2.6:
Trivial[R]constraint — Allows declaring
Drop type with drop type Name = .... At &R v / R.alloc(v) / view field construction, walks inner type; if it includes a type registered in drop_types registry, type error "Trivial[R] violated". Function type is Trivial (closure value itself is not Drop). Syntactified case (i) of design doc 12_drop_and_with.md. with expression + Drop execution in Phase 3. Added 7 tests (654 passing).
- region Phase 2.5:
R.alloc(v)syntactic sugar — Method-call
style notation for &R v. Parser holds region_stack; inside region NAME { ... } body, desugars NAME.alloc(EXPR) to Ref (NAME, EXPR). If R is not an in-scope region, treats as regular field access; existing obj.alloc(...) patterns unaffected. Added 7 tests (647 passing).
- region Phase 2.4: type-level region tag for view values +
region propagation for field access / record update — View construction returns TyCon (name, [TyRef (target_region, TyUnit)]) to embed region in value type; Field_get / Record_update reads view name + embedded region and uses subst_region to substitute field type with actual region. View value itself becomes target of escape check (Cell[S] can't be carried out of region S). Added Name[R] notation heuristic to pp_ty. Added 5 tests (640 passing). Resolves known limitation "field access returns raw R" from Phase 2.3.
2026-06-16
- region Phase 2.3: enforces region of view construction +
region parameter substitution — View can be constructed only inside region { ... } block. At construction, view declaration's region parameter R is substituted with active region name; if field has &R T, tag aligns automatically even with different region name. Added views Hashtbl and active_regions stack to typer; push/pop at Region_block; view dispatch + subst_region at Record_lit. Ties in with §5 "view type" section of memory-model.md.
- region Phase 2.2:
view V[R] of T { fields };declaration —
Introduced view type fixed in Q-009 as syntax. Like view Node[R] of int { value: int, next: int };, takes region parameter [R] and (optional) internal type of T, declares fields with { field: ty, ... }. In Phase 2.2 treated as "region-tagged record" (region is only recorded, not enforced); Node { value = 1, next = 0 } construction and n.value access work. Strict semantics (construction only inside region; mandatory &R T fields) in future Phase. Design doc: 14_view_types.md's 3 axioms (immutable / region-scoped / structural identity) at stage of syntactifying first 2.
- region Phase 2.1:
&R vvalue expression + escape check —
&R 5 turns value into region-tagged reference type. At exit of region R { body }, checks if body's type leaks R; compile- time error if leaked. Region promoted from "type-system label" to "actual safety guarantee".
- region / `&R T` Phase 1 — First step into memory model.
region R { body } expression introduces R as region name into scope; added &R T as reference type to AST/typer/eval. Phase 1 is syntax only — escape check, Trivial constraint, view type, r.alloc(v) semantics from Phase 2 onward. Design doc: corresponds to 11_region_vs_arena.md / 14_view_types.md.
- Exhaustiveness Phase 1 (Exhaustive module) — Detects bool
and variant type exhaustiveness as warnings. match Some x with | Some n -> ... outputs "missing None" to stderr but evaluation continues. Guarded arm conservatively "not covered"; as-pattern and or-pattern transparent. lib/exhaustive.ml doesn't depend on Typer (Typer calls register_variants to populate).
- Math builtins 8 (
pi/econstants +sqrt/f_abs/f_neg/
floor/ceil/round) — Float arithmetic basics complete.
- `int_max`/`int_min` constant builtins — Mere's first
non-function builtins.
- `time : unit -> float` + `exit : int -> 'a` — Unix epoch
and process termination.
- Float comparison 4 (
f_lt/f_le/f_gt/f_ge).
- CSV parser example (~130 lines, reduced RFC 4180).
- mini_calc.mere extension: let binding + variables + env-
based eval; shadowing works.
- list_lib.mere added: 12 list utility functions written in
Mere itself (map/filter/fold_left/fold_right/length/rev/take/ drop/range/replicate/for_all/any).
- Float type introduced —
TyFloatprimitive +Float_lit
(1.5 literal) + V_float; 4 conversions (float_of_int / int_of_float / str_of_float / float_of_str) + 4 arithmetic (f_add / f_sub / f_mul / f_div). No implicit int/float conversion. Resolves known limitation "no float".
- File I/O —
read_file : str -> str/write_file : str ->
str -> unit. Can write CLI tools. Added examples/word_count .mere.
- `str_unescape` builtin — Decodes
\n\t\r\\\"
\/. Escape-string support for JSON parser.
- Character literal `'X'` — Lexer only; length 1 str.
Disambiguates with tyvar 'a (closing quote presence); match c with | 'n' -> ... for dispatch.
- List display improvement —
to_stringdisplays Cons/Nil
chain as [a, b, c]. JSON parser output dramatically more readable.
- Documentation overhaul — README rewrite + newly added
docs/{tutorial, language-reference, stdlib-reference, patterns}.md (1100+ lines).
- `divmod` — Mere's first tuple-return builtin (
int → int →
(int int)`).
- `square` / `cube` — int → int 2nd / 3rd power.
- `sum_range` — O(1) sum via Gauss formula.
- `incr` / `decr` — int → int +1 / -1.
- `iter_n` — Higher-order side-effect loop.
- Polymorphic `const` / `flip` — Mere's first 3-quantified,
higher-order polymorphic builtins. Implemented via forward-ref of apply_value_ref.
- Polymorphic `id` / `swap` / `pair` — Standard set of tuple
ops complete.
- Polymorphic `fst` / `snd` — Mere's first 2-quantified
scheme builtins.
- `try_or` — Mere's first error-handling builtin.
- `fail` / `show` — Mere's first polymorphic builtins
(scheme.quantified).
- as-pattern / or-pattern —
(a, b) as p,| 1 | 2 | 3 ->
... (typer enforces binding name/type match).
- Structural equality —
==/!=recursively compare
tuples / records / constructors.
- Type alias `type Name = T;` — Parse-time substitution;
disambiguates variant/record/alias via |/of.
- Function composition `<<` / `>>` — Right-associative;
higher precedence than |>.
- Multiple type parameters `('a, 'b) result` — Resolves known
limitation "up to 1 type parameter".
- Top-level let pattern —
let _ = ...;,let (a, b) = ...;
etc. at top-level; resolves known limitation.
- If without else —
if cond then body(body unit type).
- Match guard `| pat when expr -> body` — Resolves known
limitation "no guard".
- Block expression `{ e1; e2; eN }` — Parser sugar for
Let(P_wild) chain.
- List pattern `[a, b, ...t]` — Symmetric to literal; parser
sugar.
- Record update `{ p | x = 10 }` — Immutable update.
- Record type `type Point = { x: int, y: int }` — Nominal
records; polymorphic; partial pattern.
- Mutual recursion `let rec ... and ...` — Resolves known
limitation "no mutual recursion".
- List literal `[1, 2, 3]` — Parser sugar for Cons/Nil chain.
- Pipe `|>` / signature alias — Ergonomic improvements.
- Multi-arg typed fn —
fn (x: int, y: str) -> bodydesugars
to curry.
- Massive stdlib additions — print_int / str_of_int /
int_of_str / str_len / not / min / max / abs / pow / gcd / lcm / clamp / sign / even / odd / chr / ord / to_upper / to_lower / str_trim / str_rev / str_contains / str_count / str_replace / str_starts_with / str_ends_with / str_repeat / substring / char_at / is_digit / is_alpha / is_space / read_line / print_no_nl / print_err / assert / bool_of_str / str_compare and many more.
2026-06-15 — 06-16 (early week)
- Main extensions: operator expansion (
/ %<= >= > !=&& ||),
let pattern, with expression, polymorphic types ('a opt), tuples, sum types + pattern matching.
- Design docs: Q-008 (region/arena integration), Q-009 (view type
3 axioms), Q-010 (region-version std), Q-011 (Drop order). Mere's memory model design map complete.
2026-06-06 (start date)
- After OCaml 4-phase trial, fixed host language as OCaml (Q-001
resolved).
- In 1 day, completed minimum core "integer + let + bool + if +
function + recursion + bidirectional type check + REPL" (24 tests).
- Strings + print +
++concat + unit (slice 1); REPL (slice 2);
multiple top-level decls (slice 8).
- Hindley-Milner type inference + let-polymorphism: implemented
Algorithm W + occurs check + generalize/instantiate. Inference of annotation-less functions, polymorphic id, polymorphic compose, let-poly all work (slice 9, 29 tests).
Cumulative (as of 2026-06-16)
- Design docs: 4 (Q-008/009/010/011)
- Implementation slices: 62
- Tests: 567 (initial 35 → 567, 16×)
- Builtins: 68
- Known limitations resolved: 8 (mutual recursion / guard /
multi-type-param / top-level let pattern / list display / char literal / file I/O / float)
Not yet started (future)
- `&T` reference — borrow annotation (
&shared writeetc.)
→ core of memory model
- `region R { ... }` / `view V[R] of T` — implementation of
Q-008/009
- Effect system — capability types and effect tracking
- Native codegen — LLVM or Wasm
- Exhaustiveness check Phase 2 — precise exhaustiveness for
int/str/float/tuple/record; redundancy check
- Inline unicode / Unicode source — currently ASCII only
- Module system — file split + namespace
- Dependent types / refinement types — staged introduction
per 04_fundamental_tradeoffs.md
- Row polymorphism — no annotation needed for record update
- Multi-line REPL — REPL is single-line only