Skip to content
RustingBrainGitHub
02 / 04
Deep learning / CUDA / Metal

RustingBrain

Train anything from an XOR network to a 300M-parameter transformer on one desktop GPU. No Python, no C++ build step — CUDA kernels compile at startup.

300M
Parameter LMs on a single GPU
3
Backends: CPU, CUDA, Metal
16
Tutorial chapters, zero to LM
Matches or beats PyTorch on most RTX 3060 configurations.
Flash attentionMixture of ExpertsLoRA fine-tuningBF16
examples/xor.rsrust
let mut model = Network::builder()
    .input_size(2)
    .dense(8, Activation::Tanh)
    .dense(1, Activation::Sigmoid)
    .loss(Loss::BinaryCrossEntropy)
    .optimizer(Optimizer::adam(0.05))
    .build();
Tutorials

Learn by building

  1. 01

    1. Introduction - How a Neural Network Actually Works

    This chapter has no code. If you have never understood what a neural network _is_, this is the chapter that fixes that. Everything after this one is practical, and everything after this one assumes you read this one.

    30 minFoundations
  2. 02

    2. Setup — From Nothing to a Running Model

    There is no CUDA to install, no Python version to fight, no pip dependency conflicts, no 500 MB of wheels. Two commands and you're training.

    10 minFoundations
  3. 03

    3. The XOR Problem — Your First Trained Network

    XOR is the "hello world" of neural networks, and not because it's cute. It is the smallest problem that proves you need a hidden layer. We'll solve it, and then we'll deliberately fail to solve it, because the failure teaches more than the success.

    30 minFoundations
  4. 04

    4. Working With Real Data

    Chapter 3 had four hand-typed examples. Real data arrives in files, in the wrong units, with text where you need numbers. This chapter builds the pipeline that every remaining chapter uses:

    45 minFoundations