Skip to content
RustingBrainGitHub

Tutorials

16 chapters755 min total
RustingBrain pages

Tutorials

All chapters
  1. 01

    1. Introduction - How a Neural Network Actually Works

    This chapter has no code. If you have never understood what a neural network _is_, this is the chapter that fixes that. Everything after this one is practical, and everything after this one assumes you read this one.

    30 minFoundations
  2. 02

    2. Setup — From Nothing to a Running Model

    There is no CUDA to install, no Python version to fight, no pip dependency conflicts, no 500 MB of wheels. Two commands and you're training.

    10 minFoundations
  3. 03

    3. The XOR Problem — Your First Trained Network

    XOR is the "hello world" of neural networks, and not because it's cute. It is the smallest problem that proves you need a hidden layer. We'll solve it, and then we'll deliberately fail to solve it, because the failure teaches more than the success.

    30 minFoundations
  4. 04

    4. Working With Real Data

    Chapter 3 had four hand-typed examples. Real data arrives in files, in the wrong units, with text where you need numbers. This chapter builds the pipeline that every remaining chapter uses:

    45 minFoundations
  5. 05

    5. Regression — Predicting Numbers

    XOR answered yes/no. Now we predict a quantity: a house price, which could be 20 or 450 or anything between. That's regression, and it changes three things — the output activation, the loss, and (new this chapter) what you do with the target values.

    45 minThe dense toolkit
  6. 06

    6. Classification — Choosing Between Categories

    Chapter 5 predicted a number. Now we pick one option out of three: which species is this flower? Along the way we'll meet the metric that lies to more beginners than any other in machine learning.

    45 minThe dense toolkit
  7. 07

    7. The Training Loop — Taking Control

    So far fit(...) has been a black box: hand it a dataset, get a trained model. That's fine until training goes wrong — and you can't fix what you can't see.

    60 minThe dense toolkit
  8. 08

    8. Evaluation — Measuring Honestly

    You can train a model. Now the harder skill: knowing how good it actually is.

    45 minThe dense toolkit
  9. 09

    9. Saving, Loading, and Using a Model

    Everything so far has happened inside one program: train, measure, exit. The model died with the process.

    40 minThe dense toolkit
  10. 10

    10. A Complete Project

    Every chapter so far has been one main.rs that does one thing and exits. Real projects aren't shaped like that. This chapter builds the thing you'd actually ship:

    90 minThe dense toolkit
  11. 11

    11. Training on the GPU with CUDA

    Everything so far ran on your CPU, and for the models in chapters 3–10 that was the right choice. A 200-parameter XOR network on a GPU is slower than on a CPU: the work takes microseconds and the round trip to the card takes longer than the work.

    60 minGoing further
  12. 12

    12. Importing Models From TensorFlow and PyTorch

    Sooner or later you will want to run a model you did not train. A colleague trained it in Keras. You found one on Hugging Face. You prototyped in PyTorch because the plotting was easier, and now the thing has to live inside a Rust service.

    45 minGoing further
  13. 13

    13. When Things Go Wrong

    This is a reference chapter, not a lesson. Nothing here is new material; it is the list of things that actually go wrong and what each one means.

    Going further
  14. 14

    14. Tokens and a Transformer Language Model

    Everything so far predicted one thing from a fixed set of features: a price from four columns, a class from three measurements. A language model predicts the next piece of text from all the text before it, and then does it again with its own output appended.

    120 minLanguage models
  15. 15

    15. Mixture of Experts

    Chapter 14's model ran every parameter for every token. A 300M-parameter model did 300M parameters' worth of arithmetic per token, and a 3B one did ten times that. Capacity and cost move together, which is a problem when you have one GPU.

    90 minLanguage models
  16. 16

    16. Training a Language Model End to End

    Chapter 14 trained a model on one paragraph in a few seconds. This chapter is about the version that runs for a week, and everything that only becomes a problem at that length.

    Language models