Apple Silicon deployment

128 GB of unified memory, and it still runs out.

The real problem

Apple Silicon's unified memory is what makes local inference viable: the GPU addresses the same RAM as the CPU, no copy. It is also the trap, because it is shared with everything else on the system — the browser, the IDE, the window compositor.

A 35-billion-parameter model at 4 bits takes a substantial fraction of the machine. Two resident models plus a loaded browser session, and the system starts swapping — or gets killed.

How I solve it

One explicit shared memory budget, held by a scheduler that decides which model to unload. LRU eviction, with a pinnable model that is never evicted. Applications no longer load their own weights: they ask the gateway.

The core of that scheduler is published as open source, with zero dependencies — it is useful outside MLX too.

A trap learned in the field

Two beliefs, measured then dropped.

ps rss lies. On an MLX process it understates the real footprint by a factor of roughly 30. Only phys_footprint tells the truth. Any monitoring built on rss therefore reports that all is well right up until the system dies.

Swap ratio is noise. It sits around 96% permanently because the pager resizes its own file. Using it as an alert threshold fires continuously. Likewise, swap growth of roughly 1 GiB/s while a large model loads is normal, not an incident.

A real OOM forced the final rule: measure actually-free RAM instead of trusting your own accounting. The gateway believed it had room; the system knew otherwise.

What it delivers, measured

  • 80 GB budget held on a 128 GB machine
  • Real-memory guardrail added after the incident, never removed since
  • Services restarted one at a time, never in batch — learned the hard way
  • Memory supervision daemon in production, two supervised agents

Where it runs in production

This capability is engaged in

BuildAudit

Your data cannot leave the building?

Describe the use case, the data involved and the available hardware. I answer personally within 48 business hours.