Libraries like PyTorch never ask you to draw a computational graph. You write ordinary arithmetic, and every number quietly records which numbers produced it and how. Calling .backward() on the final result then walks that recorded graph in reverse, starting from and multiplying local rates: the same walk you did by hand in this chapter. Here you build that machinery for single numbers.
Task: write a class Value that wraps one number.
Value(data) stores the number in .data and starts .grad at 0.0.a + b and a * b return a new Value that remembers both operands. Either side may be a plain number instead of a Value, as in x * w + 0.5 or 3 * a; treat it as a constant.v.tanh() and v.relu() return a new Value. Take their slopes as given: if then , and ReLU's slope is when and otherwise, including at exactly .out.backward() sets out.grad to 1.0, then fills in .grad on every Value that out was built from with .Python turns operators into method calls. a + b runs a.__add__(b) and a * b runs a.__mul__(b). When the left side is a plain number, as in 3 * a, Python instead runs a.__rmul__(3), and likewise a.__radd__ for +.
A value can be used more than once: in a * a, or when an in-between result feeds two later operations. The gradient it ends up with must account for every use.
The function trace is already written for you in the starter code. Keep it exactly as it is. It turns each named input into a Value, runs a few lines of ordinary Python with them, calls backward() on the variable named out, and reports the output and each input's gradient, rounded to 4 decimal places: