# Laya MLX Laya MLX (`aac6fef/laya-mlx`) is a community port of [[Laya]] to [[MLX]], Apple's machine-learning framework for Apple Silicon. The model and weights are Laya's; only the runtime changes. It landed on [[HuggingFace]] on 2026-09-19, a few days after Laya became the #1 trending model there, and had 117 likes two weeks later. Apache-2.0, like the original. ## What's in the box - A native MLX conversion of the English `convaiinnovations/laya` checkpoint, in FP16 - ModernBERT-large as the encoder, plus Laya's decision Transformer, scoring head and action head - About 0.4B parameters, roughly 843 MB on disk - A 512-token context, total - The three upstream question types: `choice` (pick an option), `score` (a point on an ordinal scale) and `noul` (probability of yes) The conversion renames parameters for MLX and leaves the FP16 weights alone: no retraining, no 4-bit or 8-bit [[AI Quantization|quantization]]. Each exported tensor was checked for exact equality with the source tensor cast to FP16. Question formatting, tokenizer, [[AI Model Calibration|calibration]] temperatures and output schema are preserved. Pass `dtype="float32"` for closer agreement with upstream FP32 math. ## Why port it Upstream Laya already uses the Apple GPU through PyTorch's MPS backend. What the port drops is the dependency list: no PyTorch, no Transformers library. Every computation runs in MLX. Setup is `pip install laya-mlx` on macOS 14+ with [[Python]] 3.11+. The API stays close to upstream: `laya.load("aac6fef/laya-mlx")`, then `agent.predict(state, questions)`, and you read `result["answers"]`. You call a small decision model like this from a local script or an agent loop (route this ticket, is this a refund request, how urgent is it), so a light install matters more than raw speed. For Mac-only, inference-only use, I'd take that trade. ## Fidelity checks The author tested it on an M3 Max (40-core GPU, 128 GB unified memory): - FP16 picks the same top answer as upstream PyTorch MPS FP32 on 63 of 63 decision distributions, across 16 cases - The largest gap between calibrated probabilities is 0.0054 - 100 repeated calls gave deterministic outputs, with 0 bytes of memory growth after clearing caches Sixteen cases is a small sample. And as the model card says, these checks show the port matches upstream; they say nothing about whether the answers are correct. Laya's quality, calibration and generalization limits carry over unchanged. ## Limits - Apple Silicon only, since it's MLX - English only, with no language router like upstream's. For non-English text the intended choice is upstream's multilingual checkpoint, which this repo doesn't port - The 512 tokens are shared by the input, the questions and the options, so long emails or tickets get cut. Upstream's multilingual checkpoint reads up to 8,192 tokens - Inference only. There's no training code, so [[AI Fine-Tuning|fine-tuning]] on your own decisions still goes through upstream's PyTorch tooling - It's an independent port, not from Convai Innovations. The card pins the exact upstream checkpoint and code commits it was built from ## References - [aac6fef/laya-mlx on Hugging Face](https://huggingface.co/aac6fef/laya-mlx) - [convaiinnovations/laya on Hugging Face](https://huggingface.co/convaiinnovations/laya) - [laya-mlx on GitHub](https://github.com/mizorewww/laya-mlx) - [laya-mlx benchmarks](https://github.com/mizorewww/laya-mlx/blob/main/BENCHMARKS.md) - [Laya upstream code](https://github.com/NandhaKishorM/laya) ## Related - [[Laya]] - [[MLX]] - [[HuggingFace]] - [[System One Models]] - [[AI Open Weight Models]] - [[AI Inference]] - [[Decision Models (DMs)]] - [[On-Device Machine Learning]] - [[Small Language Models (SLMs)]] - [[Transformers]]