Linear Algebra
A matrix is a rule for moving and combining coordinates. Start by dragging two-dimensional arrows so every symbol has a picture. Then use the same ideas to solve systems, choose bases, project onto models, compress data, inspect embeddings and decide whether a numerical answer can be trusted.
Every grid here is live. You drag an arrow and the numbers follow it; you change a number and the whole grid bends to match. Nothing is a picture of an answer worked out somewhere else. When a widget prints an area, or a stretch, or a score, it worked that number out from what you did a moment ago.
Adding, subtracting and multiplying, including with negative numbers. Halves and quarters. That is the whole list. No algebra, no geometry, no programming. Every word with a special meaning gets explained the first time it turns up, and the ones that usually trip people are given their own note you can open or ignore.
The steps
Draw an arrow, and add two of them
Start with a flat sheet ruled into squares. One point on it is special: the place where the two heavy lines cross. That point is the origin, and everything on this page is measured from it. The two heavy lines themselves are the axes: one lies flat and runs left and right, the other stands upright and runs up and down.
Now draw an arrow from the origin out to somewhere else. An arrow like that is called a vector. To say exactly which arrow you mean you only need two numbers: how far across, and how far up. Across is written first. So the vector (3, 2) is the arrow that goes three squares right and two squares up. Going left is a negative first number, and going down is a negative second number.
I have not used a grid with minus numbers on it
The two axes cross at the origin, and each one carries numbers. They get bigger going right along the flat axis and smaller going left, so one square to the left of the origin is at -1, and two squares left is at -2. The same idea works upright: above the origin the numbers count up, below it they count down into the negatives.
Nothing here needs you to be quick with negative numbers. If a sign confuses you, drag the arrow into that corner and read what the widget prints. It is showing you the answer while you move.
Why call it a vector and not just a point
A point is a place. A vector is a trip: a distance in a direction. Written down they look the same, because the trip that starts at the origin and ends at that place is described by exactly the same two numbers.
The difference starts to matter the moment you add two of them, which is the second half of this step. Adding two places makes no sense. Adding two trips does: do the first, then do the second, and see where you end up.
Two arrows can be added. Put the tail of the second arrow on the tip of the first, and the sum is the arrow that runs from the origin to where you ended up. That is the whole rule. It is called adding nose to tail.
Adding two arrows using only the numbers
You never have to draw it. Add the across numbers together, add the up numbers together, and those two answers are the sum. So (2, 1) plus (1, 3) is (3, 4), because 2 plus 1 is 3 and 1 plus 3 is 4.
The drawing and the arithmetic are the same thing. Walking two right and one up, then one right and three up, gets you to the same place as walking three right and four up in one go. The order you do the two walks in does not change where you finish.
Stretch an arrow, and turn it around
There is one more thing you can do to a single arrow: multiply it by a plain number. Multiplying (2, 1) by 3 means multiplying both of its numbers by 3, giving (6, 3). The arrow points the same way and is three times as long.
A plain number used this way has a name. It is called a scalar, because its job is to scale something. Scalars can be fractions, and they can be negative, and both of those are more interesting than they sound.
Why multiplying both numbers keeps the direction
Think of the arrow as a set of instructions: go 2 across, then 1 up. Doing that same trip three times in a row lands you 6 across and 3 up. You have walked in a perfectly straight line the whole time, because every step was in the same direction as the last one.
That is why the two numbers have to be multiplied by the same scalar. Multiplying only the across number would bend the trip. Multiplying both keeps the shape of the walk and only changes how far you go.
What a negative scalar does, and what 0 does
Multiplying by -1 flips both numbers' signs, so (2, 1) becomes (-2, -1). Same length, opposite direction. Multiplying by -2 does both at once: twice as long, and pointing backwards.
Multiplying by 0 gives (0, 0), an arrow with no length at all. It is called the zero vector, and it is the only vector that has no direction. It is not nothing. It is a real answer that a lot of the interesting questions later in this course turn out to have.
Now combine the two ideas. Take two fixed arrows, scale each one by a number of your choosing, and add the results nose to tail. This is called a combination of the two arrows, and with the right pair of scalars you can reach a surprising number of places.
A combination, written out slowly
Say u is (2, 1) and w is (-1, 2). Pick the scalars 2 and 1. Then 2 lots of u is (4, 2), and 1 lot of w is (-1, 2), and adding those nose to tail gives (3, 4).
Two dials, one landing spot. That is the whole machine, and almost everything later in this course is a question about it. Which dial settings reach a given spot? Is there ever more than one answer? Are there spots that cannot be reached at all?
Move every point at once
Until now you have moved one arrow at a time. Now move the entire sheet in one action. Every point on it slides somewhere new, all together, and the grid you have been reading off bends with it.
The thing that does the moving is a matrix: four numbers written in a box, two across and two down. Feed the sheet in, get a bent sheet back. One point never goes anywhere, and that is the origin.
The moves in the lab below have ordinary names, and one of them is worth having now. A move that leans the sheet over, sliding it sideways further and further the higher up you look while one line stays exactly where it was, is called a shear. The lab's Lean over button is a shear, and later steps use the word.
How to read a box of four numbers
The box has a top row and a bottom row, and a left column and a right column. Four slots: top left, top right, bottom left, bottom right. That is all the structure there is. Written on one line people often squash it into square brackets, but on this page it always stays as a box so you can see the rows and columns.
What the four numbers actually mean is Step 4, and it is worth not guessing. For now the honest description is that they are the settings of a machine and the picture is the machine running.
What kind of moving this is, and what it rules out
These machines are limited on purpose, and the limits are what makes four numbers enough. Straight lines stay straight. Lines that were parallel, meaning always the same distance apart and never meeting, stay parallel. Marks that were evenly spaced along a line stay evenly spaced. And the origin stays exactly where it is.
Moving that obeys those four promises is called linear, which is where the name of this whole subject comes from. It rules out bending the sheet into a curve, and it rules out sliding the whole sheet sideways, because sliding would move the origin.
Does the whole sheet really move, or only the lines I can see
The whole sheet. The lines are drawn so that you have something to look at, but every point between them moves too, and it moves in the way the lines around it suggest. If you imagine a dot halfway between two grid lines before the move, it is halfway between the same two lines after it. That is the evenly-spaced promise doing its work.
This matters because the next lab asks you to recognise a move from its picture. You are reading the fate of every point off the fate of a handful of lines, and you are allowed to, because the four promises make the rest follow.
Read a matrix by watching two arrows land
Four numbers is not many, and yet they decide where every one of infinitely many points goes. Here is the reason, and it is the single most useful idea in the subject.
Put an arrow one square to the right, at (1, 0), and call it i. Put another one square up, at (0, 1), and call it j. Those two are the basis: the two arrows everything else is built out of. The arrow (3, 2) is just 3 lots of i plus 2 lots of j.
Column and row, and the odd names i and j
In a box of numbers, a row reads across and a column reads down. The box in this course has two rows and two columns, so a column is a pair of numbers stacked one above the other. When this page says "the first column" it means the top-left number with the bottom-left number underneath it.
The names i and j are just labels that stuck, the way x and y stuck for the axes. In books they usually wear little hats to show they are one square long. Nothing depends on the names, but you will meet them everywhere, so they are worth recognising.
Now the point. When a matrix bends the sheet, i and j land somewhere. Write down where i landed as the first column of the box, and where j landed as the second column, and you have written the matrix. The four numbers are the two landing spots.
The do-nothing matrix
If i stays at (1, 0) and j stays at (0, 1), nothing has moved. That box, with 1 and 0 on the top row and 0 and 1 on the bottom row, is called the identity. Most people call it the do-nothing matrix when they are talking rather than writing.
It is the grid you see before you press Apply, and it is what a move and its undo add up to in Step 9. Recognising it on sight is worth the small effort: 1s down the diagonal from top left to bottom right, 0s in the other two slots.
Why watch i and j, and not some other pair of arrows
Any two arrows that do not lie on one line would do. You could watch (2, 1) and (-1, 2) land instead. You would still be able to work out where every other arrow goes, because every arrow is some combination of those two, exactly as in Lab 4.
The pair i and j wins because their combinations are free. The arrow (3, 2) is 3 lots of i and 2 lots of j, and those two numbers are already written on the arrow. Watching any other pair would mean working out the two dial settings first, every single time.
That also tells you how to work out where any arrow goes, without drawing anything. The arrow (x, y) was x lots of i plus y lots of j before the move. Afterwards it is x lots of where i landed plus y lots of where j landed. That sentence is the rule for applying a matrix to a vector, and there is nothing else to it.
Do two moves, and then swap the order
Bend the sheet with one matrix. Then bend the already-bent sheet with a second one. The end result is some new bent sheet, and since every bent sheet is a matrix, the pair of moves has a single matrix of its own. Working that matrix out is what people mean by multiplying two matrices.
Why "B after A" is written with B on the left
When the two are written next to each other as BA, the one nearest the arrow is the one that happens first. The arrow is imagined sitting to the right, so A touches it first and B works on the result. It reads backwards compared with English, and nearly everyone finds it awkward at the start.
This page avoids the trap by never relying on the written order. Every control says first move and second move in words, and the picture does them in that order in front of you.
Order matters, and here is the smallest example
Take a quarter turn and a flip across the flat axis. Turn first, then flip, and i ends up pointing down. Flip first, then turn, and i ends up pointing up. Same two moves, different answer, and the picture in the lab below shows both at once so you do not have to take that on trust.
Some pairs do agree. Two turns in the same plane agree, and two pure stretches along the axes agree. Agreeing is the exception, though, and assuming it is one of the most common mistakes in the subject.
Working out the combined matrix by hand uses only Step 4. Ask where i ends up after both moves, and write that down as the first column. Ask the same about j for the second column. There is no separate rule to learn.
The row-times-column recipe books teach, and where it comes from
A book will tell you to take a row of the left-hand box and a column of the right-hand one, multiply the two first numbers, multiply the two second numbers, and add. Do that for each of the four row-and-column pairings and you have the combined box.
That is not a different rule. Walking i through both moves is exactly those multiplications, in exactly that order, and the recipe is just the bookkeeping written down so you do not have to draw anything. If you ever forget it, walk i and j through by hand as in the lab below and you will rebuild it.
Measure how much bigger space got
Bending the sheet changes areas. The single square between i and j, one wide and one tall, is called the unit square, and its area starts at exactly 1. Bend the sheet and it turns into a slanted box. Measure that slanted box and you know what happened to every area on the sheet at once.
The area of a slanted box
A box whose opposite sides are parallel but whose corners are not square still has an area. You work it out the ordinary way: pick one side as the base, measure straight across to the side opposite it, and multiply the two. Leaning a box over changes neither its base nor that straight-across distance. So it does not change the area at all, which is why a heavy shear can leave the number at 1.
You do not have to measure anything by hand here. The widget shades the box and prints its area, and it works that number out from the corners it just drew.
What a factor is
A factor is a multiplier. A factor of 2 means twice as much, a factor of 1 means unchanged, and a factor of 0.5 means half. Saying that a move has an area factor of 3 means that any shape you draw first, of any size, comes out with three times the area afterwards.
The strong part of that claim is the "any shape". Measuring one small square tells you about a circle, a triangle and a drawing of a cat, because the four promises from Step 3 force every area on the sheet to change by the same multiplier.
Now the number itself. If the box reads a and b across the top and c and d across the bottom, the area factor is a×d − b×c. It is called the determinant. You have already watched it change; the formula just saves you the drawing.
Why the determinant can be negative
An area is never negative, but the determinant can be, and the minus sign carries a second piece of news: the sheet has been turned over. Before the move, going from i to j is a turn anticlockwise. If afterwards that same trip has become a turn clockwise, the sheet has been flipped, like reading a page through the back of the paper.
So the size of the determinant is the area factor, and its sign says whether the sheet is still the right way up. A determinant of -3 means areas tripled and the sheet is mirrored.
Squash space flat, and lose something forever
Push the determinant down toward zero and watch what happens to the shaded patch. It gets thinner. At exactly zero it has no thickness left, and the whole sheet has been pressed onto a single line through the origin.
What "losing information" means here
Suppose a machine turns every number you feed it into that number's distance from zero. Feed it 5 and you get 5. Feed it -5 and you also get 5. Now somebody hands you the answer 5 and asks what went in. You cannot tell them, and no amount of cleverness will help, because two different inputs really did produce the same output.
A squashed sheet is that, on a much larger scale. Whole lines of starting points get pressed onto the same landing spot. The information about which one it was is not hidden or encrypted. It is gone.
Many arrows, one landing place
When the determinant is zero, the two columns of the matrix point along the same line. Where i lands and where j lands are multiples of one another. Everything the matrix can produce is a combination of those two, so everything it can produce lies on that one line.
And once the output is a line rather than a whole sheet, there are not enough landing spots to go round. The inputs have to double up, and they do: a whole parallel family of input lines maps onto each single output point.
Here is the sharpest way to see the loss. When the sheet is squashed there is always a whole line of arrows, not just the zero vector, that land exactly on the origin. Every one of them has been wiped out.
These two families have names, if you want them
The line that everything lands on is the column space, because it is everything you can build out of the columns. The line of arrows that get wiped out to the origin is the null space, sometimes called the kernel. Both names turn up constantly in later work.
You will not need either word again in this course, and nothing later depends on it. They are here so that the ideas you have already got by dragging arrows do not feel like strangers when you meet them written down.
Ask which arrow lands on the target
Almost every use of this subject comes down to one question. You know the matrix, you know where you want to end up, and you need the arrow that gets sent there. Given the machine and the output, find the input.
Written out with letters, that question is a pair of equations. Written on the grid, it is a target ring and an arrow you drag until it lands. They are the same question, and the grid version is easier to believe.
What the pair of equations looks like written out
Suppose the box has 2 and 1 on top and 1 and 3 underneath, and the target is (5, 5). Applying the box to an arrow (x, y) gives (2x + 1y, 1x + 3y), from the recipe in Step 4. Wanting that to be the target means wanting 2x + y = 5 and x + 3y = 5 to be true at the same time.
Two statements, two unknown numbers, and one pair of values that satisfies both. That is what a system of equations is. There is no new mathematics in the phrase, only a name for the situation.
The same question in three languages
Language one, equations: find x and y that make both lines true at once. Language two, columns: find how many lots of the first column and how many of the second add up to the target, which is exactly the two-dial game from Lab 4. Language three, moving space: find the arrow that the bent grid carries onto the target.
Different books lead with different ones, and switching between them on demand is most of what being comfortable here means. The dials and the bending grid are the two you have already used.
If I can find it by dragging, why does anyone need a method
Dragging works here because the answer is a whole number and there are only two numbers to find. It stops working almost immediately. The arrow in the lab below snaps to quarter squares, so an answer of one third of a square cannot be reached by hand at all. Worse, you would not be told that, and you would stop at the nearest quarter believing you were done.
The real reason is size. A weather model has millions of unknown numbers rather than two, and nobody is going to drag those. A method that grinds out the answer without looking at a picture is the only thing that survives the jump, which is what the second lab in this step shows one doing.
A machine can compute the answer instead of you hunting for it, and it is worth watching one do it, because the interesting part is the two cases where it cannot.
Undo a move, or find out you cannot
If a matrix bends the sheet, is there a second matrix that bends it straight again? Sometimes. When there is, it is called the inverse, and doing the move and then its inverse leaves the sheet exactly as it started, which is the do-nothing matrix from Step 4.
Undoing, in ordinary numbers first
Multiplying by 4 is undone by multiplying by one quarter, because 4 times a quarter is 1, and multiplying by 1 changes nothing. Every number has a partner like that, with a single exception: nothing multiplied by 0 gives 1, so multiplying by 0 cannot be undone. Once you have multiplied by zero, the original number is gone.
Matrices work the same way, and the determinant is what plays the part of the number. A determinant of 0 is the case with no partner, for the same reason. The squash in Step 7 threw information away, and no matrix can put back what is not there.
Reading the undo off the picture
Here is a way to find the inverse without any formula. The original matrix says where i and j go. The inverse has to say where they came from. So look at the bent grid, find the two arrows on it that are now sitting at (1, 0) and (0, 1), and those are the columns of the inverse.
That is a usable trick on simple matrices, and it is also the reason the inverse of a quarter turn one way is a quarter turn the other way, with no arithmetic at all.
The short formula, for a two by two only
Start with a and b on the top row and c and d on the bottom. Swap a and d, so the top row now begins with d and the bottom row ends with a. Flip the sign of the other two, so b becomes -b and c becomes -c. Then divide all four of the results by the determinant. That box is the inverse.
The division is where the whole story sits. A determinant of 0 makes it a division by zero, which is exactly the case that has no answer. Bigger boxes have no formula this short, so leaning on this one as your understanding is a mistake; the picture is the part that survives.
Find the arrows that never turn
Bend the sheet and most arrows swing round to a new direction. A few do not. They get longer or shorter or flipped end for end, but they stay on the line they started on. Those are worth hunting for, and this step is the hunt.
What counts as "the same line"
An arrow and the same arrow pointing exactly backwards live on one line through the origin. So do an arrow and a stretched copy of it. That is the fact from Lab 3: every multiple of one arrow sits on one straight line, negatives included.
So "did not turn" means "ended up somewhere on its own line", which allows growing, shrinking and flipping. It rules out anything that ends up pointing off that line, however slightly.
Why the length of the arrow you drag does not matter
If an arrow stays on its line, so does every multiple of it. Doubling the input doubles the output, and a doubled arrow is still on the same line. So the answer is never a particular arrow, it is a whole direction and the widget draws it as a dashed line right across the grid.
That is why the hunt below only cares about the angle you drag to. Drag out to the edge if it makes the aiming easier. The reading will not change.
An arrow that stays on its own line is called an eigenvector of the matrix, and its line is an eigen-direction. The last lab of the course calls the same thing an eigen-line, because there it is drawn on the grid. The word is half German and it means something like "its own", as in the matrix's own private directions.
Some matrices have no such line at all
A pure turn moves every single direction. Turn the sheet by any amount that is not a whole half circle and nothing at all is left pointing where it started. So the hunt finds nothing and the turn bar never reaches zero. The third preset in the lab below, marked A pure turn, is exactly that case, and finding nothing there is not a bug.
A shear is the in-between case. It has one such line and no more. So the count can be two, one, or none, and the picture tells you which before any arithmetic does.
Measure the stretch along those lines
You have the special directions. Each one comes with a number: how much longer the arrow got. That number is the eigenvalue of that direction, and it is written with the Greek letter lambda, which looks like this: λ.
What the stretch number counts
If an eigen-direction's arrow comes out twice as long, its eigenvalue is 2. If it comes out at three quarters of its old length, the eigenvalue is 0.75. If it does not change at all, the eigenvalue is 1, and that direction is completely untouched by the move.
The number belongs to the direction, not to the particular arrow. Slide your arrow further out along the same dashed line and the stretch reading stays where it is, which the lab below lets you confirm by dragging.
A negative eigenvalue, and a zero one
An eigenvalue of -2 means the arrow comes out twice as long and pointing backwards. The line is still held, the arrow has just been reversed along it, which is the case from the last quiz in Step 10.
An eigenvalue of 0 means arrows on that line are sent to the origin. That is the wiped-out line from Step 7, seen from a new angle: a matrix with a determinant of 0 always has 0 as one of its eigenvalues.
A way to guess both stretch numbers without dragging anything
Two facts pin them down for a box of four. The two stretch numbers multiply to the area factor, which you already know how to work out. They also add up to the sum of the top-left and bottom-right numbers. Take the box with 3 and 1 on top and 1 and 3 underneath. Its two stretch numbers must multiply to 8 and add to 6, and only 4 and 2 do both.
You can check that pair against the lab below, which reads them off the grid instead. Guessing like this runs out quickly on unfriendly numbers, and boxes bigger than two by two need a real method, but the two facts themselves stay true however large the box gets.
One more thing these numbers are good for. Apply the same matrix over and over to any arrow you like, and the direction with the biggest stretch takes over. Whatever you start with drifts toward that line and then stays there.
Turn a character, shrink a picture, rank the pages
Three jobs that look unrelated. All three are a matrix doing what you have watched it do, and this step is where the dragging pays for itself.
Start with the easiest. A game character on screen is a list of corner points. Turning it is a matrix applied to every one of those points, and the two columns are just where i and j land after the turn.
What a picture is made of
A screen is a grid of tiny squares called pixels, each one holding a number that says how bright it is. A small grey picture eight pixels across and eight down is therefore 64 numbers. Any grid of numbers with rows and columns like that is a matrix too, just a bigger one than the boxes of four you have been dragging.
Everything in this course was stated for the two-by-two case because you can see it. The ideas do not stop there. An eight-by-eight matrix has eight columns, up to eight eigen-directions to hunt for, and the same question about what gets lost when it is squashed. Only up to eight, because a matrix can come with fewer of those lines than it has columns, exactly as the pure turn in Step 10 came with none.
Now the picture. A grey image is a big grid of brightness numbers, and storing all of them is expensive. If the image can be built out of a few simple layers instead, you store the layers.
What the word algorithm means
An algorithm is a recipe written precisely enough that a machine can follow it without judgement: do this, then that, and stop when the following is true. Long division is an algorithm. So is the repeated-applying trick from Lab 22, which is the recipe the next two labs run.
The layer-finder below really does run that recipe while you watch, with a hard limit on how many rounds it may take so it can never sit there spinning. The numbers it prints are what it worked out from the picture, not figures typed in by an author.
Last one, and it is the reason search engines work. Give every page on a small web a score, then let each page pass its score along its outgoing links, over and over. The scores settle down, and where they settle is a ranking.
Where the matrix is hiding in that
Put the pages' scores in a list. One round of passing scores along links takes that list and produces a new list, and each new score is built from the old ones by multiplying and adding. That is a matrix applied to a list, exactly as in Step 4, with one column per page instead of two.
So a round is one application of a matrix, and running many rounds is Lab 22 all over again. The ranking the scores settle on is the eigen-direction with the biggest eigenvalue, which for this kind of matrix works out to 1. The scores stop changing because they have found the direction that does not turn.
Length and angle come from the dot product
For vectors u=(u₁,u₂) and v=(v₁,v₂), the dot product is u·v=u₁v₁+u₂v₂. It is a number, not a vector. The length of v is ||v||=√(v·v). A unit vector has length 1 and records direction without scale.
The same dot product also equals ||u||||v||cos θ. A positive result means an acute angle, zero means perpendicular, and a negative result means an obtuse angle. Dividing the dot product by both lengths gives cosine similarity, which compares direction while ignoring magnitude.
Distance and similarity need a scale choice
Euclidean distance treats one unit in every coordinate equally. Standardising features, weighting coordinates or learning an inner product changes the geometry. A nearest neighbour is meaningful only after that choice, and a physical displacement may need its original length.
A basis is a coordinate measuring kit
A set of vectors is linearly independent when none can be made from the others. Their span is every linear combination they can reach. A basis is an independent spanning set: enough directions to describe the space, with no redundant one.
Coordinates belong to a basis. The physical arrow does not change when its coordinates change from standard axes to tilted axes. The number of vectors in any basis is the dimension. In a plane it is two; in RGB colour it is three; an embedding may use hundreds or thousands.
Affine coordinates add an origin
Vectors describe displacement from zero. Points also need an origin. Graphics often add a homogeneous coordinate so translation can join rotation and scaling inside one larger matrix.
Elimination exposes pivots, rank and free variables
Gaussian elimination replaces equations with equivalent ones. Swap two rows, multiply a row by a nonzero number, or add a multiple of one row to another. These moves preserve the solution set while creating zeros below pivot positions.
A pivot column adds an independent direction. The number of pivots is the rank. A non-pivot input column creates a free variable and therefore a null-space direction. Rank-nullity says rank plus nullity equals the number of input columns.
Reduced row echelon form is useful for explanation, but software often uses LU or QR factorisation and pivoting instead of forming the full reduced matrix. The result and the numerical method are separate questions.
Pivoting chooses a safer next equation
Dividing by a tiny pivot magnifies rounding error. Partial pivoting swaps in a larger available entry before elimination. The exact equations are equivalent either way, but their floating-point behaviour can differ greatly.
Projection finds the closest answer a model can express
The projection of v onto a nonzero direction u is ((v·u)/(u·u))u. The leftover error is perpendicular to u. For a subspace with an orthonormal basis, project onto each basis vector and add the pieces.
An overdetermined system Ax=b may have no exact solution. Least squares chooses x that minimises ||Ax-b||². At the optimum, the residual b-Ax is perpendicular to every column of A, giving the normal equations AᵀAx=Aᵀb.
QR is usually better than normal equations
Forming AᵀA squares the condition number and can lose accuracy. QR factorisation builds an orthonormal basis for the columns and solves the same least-squares problem more stably.
SVD separates directions, strengths and rank
Every real matrix has a singular value decomposition A=UΣVᵀ. V chooses orthonormal input directions, Σ stretches them by nonnegative singular values, and U chooses orthonormal output directions. Unlike eigenvectors, singular vectors work for rectangular matrices and always provide orthonormal bases.
Keeping the largest k singular values gives the best rank-k approximation under common matrix norms. PCA applies the same geometry to centred data and keeps directions with the most variance. The image layers in Lab 24 are a low-rank approximation.
Modern ML uses this algebra everywhere: embeddings are rows in large matrices; attention compares projected vectors; low-rank adaptation trains small factor matrices; quantisation stores approximate values. These methods save memory or computation, but evaluation must check what information was lost.
Randomised low-rank methods avoid reading every direction repeatedly
A random sketch can find an approximate important subspace before a smaller SVD is computed. The speed gain comes with probability and approximation error, so report the seed, retained energy and downstream quality rather than only matrix rank.
A correct formula can still give an unstable answer
The condition number compares the largest and smallest singular values. A large value means some input directions are stretched far more than others. Solving then amplifies measurement or rounding error along the weak direction.
Conditioning belongs to the problem; stability belongs to the algorithm. Partial pivoting, QR, SVD and iterative refinement control avoidable numerical error. Explicitly computing A⁻¹ to solve Ax=b is usually slower and less stable than solving the factorised system directly.
Report residual ||Ax-b|| and, when possible, forward error or a condition estimate. A tiny residual can coexist with an inaccurate x when the problem is ill-conditioned. Use scaled tests and higher precision to diagnose, not to hide, the geometry.
Automatic differentiation uses the same matrix products
Reverse-mode automatic differentiation propagates vector-Jacobian products backward through a computation graph. Large or tiny singular values help explain exploding and vanishing gradients. Stable training still needs measured scales, not only symbolic derivatives.
What you can do now
You started with a grid and an arrow. Here is the list, and it is worth reading slowly, because most of it is usually taught as arithmetic and you learned it as pictures.
How this maps onto a normal textbook
What this course called "where i and j land" a book calls matrix-vector multiplication."Two moves in a row" is matrix multiplication."The area factor" is the determinant,"the undo" is the inverse, "which arrow lands on the target" is solving a linear system, and "the lines that hold still" are eigenvectors with their eigenvalues.
Nothing was simplified into a shape you have to unlearn. The pictures are what the symbols mean, so a book's notation should now read as a shorthand for something you have already dragged around with your own hands.
- Read a matrix as a pair of landing spots, and sketch what it does to a grid.
- Apply a matrix to a vector without looking up a rule, by scaling the two columns and adding.
- Combine two moves, and say why the order matters.
- Work out a determinant and say what it means about area and about flipping.
- Recognise a squashed matrix on sight and say what has been lost.
- Turn a pair of equations into a question about arrows, and know when it has one answer, none, or endlessly many.
- Find eigen-directions by hunting, and their eigenvalues by measuring the stretch.
- Explain what a search engine's ranking has to do with a direction that does not turn.
- Use dot products to measure length, angle, projection and cosine similarity.
- Choose a basis, change coordinates, and connect pivots to rank and free variables.
- Fit an overdetermined system by least squares and inspect its residuals.
- Read SVD and PCA as directions plus strengths, including low-rank uses in modern ML.
- Recognise an ill-conditioned problem and distinguish sensitivity from solver stability.
Where this goes next
- More than two dimensions. Three columns instead of two, the determinant as a volume factor, and the same squash question about whether a whole plane gets flattened.
- Introduction to Algorithms. The repeated-applying trick from Lab 22 is one method among many, and comparing methods by counting their work is a subject of its own.
- Data Structures. A matrix is stored as a grid of numbers in memory, and how you lay that grid out changes how fast the arithmetic runs by a surprising margin.
- Neural Networks and Training That Works. Embeddings, attention, gradient propagation, low-rank adaptation and numerical evaluation all use the later chapters directly.
- Control Systems and Optimisation. State-space models, least squares, eigenvalues, conditioning and iterative solvers continue there.