// First-Principles Deep Learning Engine
mentats
A deep learning library built completely from scratch in pure Rust with zero ML external dependencies, inspired by the human computers of Frank Herbert's Dune.
What is a Mentat?
In Frank Herbert’s Dune, following the Butlerian Jihad that destroyed all artificial intelligence and banned thinking machines under the commandment “Thou shalt not make a machine in the likeness of a human mind,” society developed Mentats. These are humans conditioned from birth to be able to perform complex mental computations and huge data analysis tasks completely through raw intellect and cognitive training.
The mentats framework adopts this thinking in regard to machine learning. Modern deep learning relies heavily on towering abstraction layers like PyTorch, LibTorch bindings, CUDA runtimes or high-level crates like ndarray, burn, and candle. Mentats gets rid of all of them. Every tensor stride, matrix multiplication, backward activation derivative, Adam second moment and latent sampling equation is hand-derived and implemented natively in Rust.
// Live Evaluation: Trained CVAE Decoder
Click a digit and let the network create new handwriting styles it has never seen before.
The interactive demo below runs real-time client-side inference using the exact cvae_decoder.rmlc binary weights trained by mentats in Rust. Select a target digit class and explore the latent dimension.
Loading Mentats Model...
// Core Framework Architecture
01 / Linear Algebra Engine
Custom Tensor & Matrix Primitives
Implemented raw contiguous-memory vector and matrix buffers with custom strided indexing. Provides cache-conscious matrix multiplication (A × B), broadcasting, transpositions and elementwise activation functions (ReLU, Sigmoid, Softmax).
02 / Reverse-Mode Autodiff
Analytical Backpropagation
Forward-pass activations and input states are cached within computational layers, allowing backward passes to propagate gradients back through the network - which is what lets the model learn from its mistakes - activation layers and multi-term loss functions (Categorical Cross-Entropy, Binary Cross-Entropy, KL Divergence).
03 / Numerical Optimization
SGD & Adam with Moment Correction
Features Stochastic Gradient Descent with velocity momentum as well as the Adam optimizer built directly from the Kingma & Ba paper, maintaining running exponentially decaying first (m_t) and second (v_t) moment vectors with bias correction.
04 / Model Serialization
The .rmlc Checkpoint Format
Engineered a custom binary serialization format (.rmlc — Rust Machine Learning Checkpoint) storing layer types, weight dimensions, and raw IEEE-754 32-bit floats. Checkpoints serialize with zero external runtime overhead and can be read by both native Rust and web clients.
// Research & Evolutionary Milestones
The foundation test for Mentats. Trained on the full 60,000-sample MNIST dataset using pure SGD, learning rate 0.01, and categorical cross-entropy over 5 epochs. Proved mathematical precision of dense layer gradient updates, backprop jacobians, and activation functions.
Transitioned to deep generative modeling. Derived the Gaussian reparameterization trick (a trick that lets the network learn to generate new examples and not just recognise existing ones) (z = μ + σ ⊙ ε) and optimized the Evidence Lower Bound (ELBO). Tackled posterior collapse - a common failure mode where the model gives up on using its learned representation - by implementing per-batch β-annealing (0 to 1 over 20 epochs) and free-bits KL clamping, preventing latent units from degenerating into uninformative noise.
Conditioned both the encoder and decoder on 10D one-hot class vectors. This disentangled digit semantics (0–9) from continuous handwriting style attributes (line thickness, slant, loop size) in a 32-dimensional latent space. This is the model exported and evaluated live in the interactive demo above.
// Mathematical Foundations
01 / Latent Conditioning
Class-Guided Synthesis
By conditioning both encoder and decoder on a 10D one-hot vector, the model disentangles digit identity from handwriting style (slant, thickness, loops).
02 / Posterior Stability
Free-Bits KL Annealing
Prevents posterior collapse during training by warming up the KL weight β over initial epochs and clamping minimum information per latent dimension. This prevents the model from collapsing into producing the same boring outputs regardless of the input.
03 / Pure Rust Engine
First-Principles ML
No external ML dependencies. Custom Tensor operations, backpropagation and Adam moments implemented directly in Rust.