# Laya MLX
Laya MLX (`aac6fef/laya-mlx`) is a community port of [[Laya]] to [[MLX]], Apple's machine-learning framework for Apple Silicon. The model and weights are Laya's; only the runtime changes. It landed on [[HuggingFace]] on 2026-09-19, a few days after Laya became the #1 trending model there, and had 117 likes two weeks later. Apache-2.0, like the original.
## What's in the box
- A native MLX conversion of the English `convaiinnovations/laya` checkpoint, in FP16
- ModernBERT-large as the encoder, plus Laya's decision Transformer, scoring head and action head
- About 0.4B parameters, roughly 843 MB on disk
- A 512-token context, total
- The three upstream question types: `choice` (pick an option), `score` (a point on an ordinal scale) and `noul` (probability of yes)
The conversion renames parameters for MLX and leaves the FP16 weights alone: no retraining, no 4-bit or 8-bit [[AI Quantization|quantization]]. Each exported tensor was checked for exact equality with the source tensor cast to FP16. Question formatting, tokenizer, [[AI Model Calibration|calibration]] temperatures and output schema are preserved. Pass `dtype="float32"` for closer agreement with upstream FP32 math.
## Why port it
Upstream Laya already uses the Apple GPU through PyTorch's MPS backend. What the port drops is the dependency list: no PyTorch, no Transformers library. Every computation runs in MLX.
Setup is `pip install laya-mlx` on macOS 14+ with [[Python]] 3.11+. The API stays close to upstream: `laya.load("aac6fef/laya-mlx")`, then `agent.predict(state, questions)`, and you read `result["answers"]`.
You call a small decision model like this from a local script or an agent loop (route this ticket, is this a refund request, how urgent is it), so a light install matters more than raw speed. For Mac-only, inference-only use, I'd take that trade.
## Fidelity checks
The author tested it on an M3 Max (40-core GPU, 128 GB unified memory):
- FP16 picks the same top answer as upstream PyTorch MPS FP32 on 63 of 63 decision distributions, across 16 cases
- The largest gap between calibrated probabilities is 0.0054
- 100 repeated calls gave deterministic outputs, with 0 bytes of memory growth after clearing caches
Sixteen cases is a small sample. And as the model card says, these checks show the port matches upstream; they say nothing about whether the answers are correct. Laya's quality, calibration and generalization limits carry over unchanged.
## Limits
- Apple Silicon only, since it's MLX
- English only, with no language router like upstream's. For non-English text the intended choice is upstream's multilingual checkpoint, which this repo doesn't port
- The 512 tokens are shared by the input, the questions and the options, so long emails or tickets get cut. Upstream's multilingual checkpoint reads up to 8,192 tokens
- Inference only. There's no training code, so [[AI Fine-Tuning|fine-tuning]] on your own decisions still goes through upstream's PyTorch tooling
- It's an independent port, not from Convai Innovations. The card pins the exact upstream checkpoint and code commits it was built from
## References
- [aac6fef/laya-mlx on Hugging Face](https://huggingface.co/aac6fef/laya-mlx)
- [convaiinnovations/laya on Hugging Face](https://huggingface.co/convaiinnovations/laya)
- [laya-mlx on GitHub](https://github.com/mizorewww/laya-mlx)
- [laya-mlx benchmarks](https://github.com/mizorewww/laya-mlx/blob/main/BENCHMARKS.md)
- [Laya upstream code](https://github.com/NandhaKishorM/laya)
## Related
- [[Laya]]
- [[MLX]]
- [[HuggingFace]]
- [[System One Models]]
- [[AI Open Weight Models]]
- [[AI Inference]]
- [[Decision Models (DMs)]]
- [[On-Device Machine Learning]]
- [[Small Language Models (SLMs)]]
- [[Transformers]]