Sound and Signals
Clap your hands and the air near your palms is squashed, then thinned, then squashed again. Draw that against time and you have a wave. Everything a computer or a phone ever does with that clap starts by throwing most of it away. It measures the wave a few thousand times a second and keeps only the numbers. This course does the throwing away by hand. You will draw a wave, add two together, and chop one into numbers. Then comes the odd part: chop too slowly and a fast wave comes back as a slow one that was never there. By the end you will be able to say exactly how much of a sound a recording keeps, and exactly what it loses.
Three small pieces of machinery sit under the whole course. One says what a wave is doing at a given moment. One asks that question at evenly spaced instants and writes the answers down. One rounds each answer to the nearest of a fixed set of levels. Every number on every page comes out of those three, worked out again each time you move a control. Where a step says a 7 Hz wave comes back as 3 Hz, something on that page searches for the wave that fits the measurements. It finds 3 without being told. Rest your pointer on any chart and it reads out what every line is doing at that instant. The small play button in a chart's corner walks a cursor across it.
Arithmetic, fractions and percentages. Nothing else, and no course before this one. There is no algebra here, no calculus, and nothing that needs a physics lesson. A sound is something you have heard and a wave is something you can draw. After that it is all counting and dividing. The words are introduced as they arrive, in bold, in the sentence that first needs them.
The steps
Draw a sound and change its three numbers
A sound is air being pushed and pulled. A drum skin moves out and squashes the air in front of it, then moves back and thins it out. That pattern of squash and thin travels to your ear at about 340 metres a second. Draw how squashed the air is against time and you have drawn the sound. The line you draw is a signal: a quantity that changes as time passes and carries something worth having.
The simplest signal that repeats is a sine wave, a smooth rise and fall that does the same thing over and over. Three numbers set one completely. The amplitude is how far it swings away from the middle line, and that is loudness. The frequency is how many complete rises and falls it fits into a second. That is how high or low the note sounds. One complete rise and fall is a cycle, and cycles per second has its own name, the hertz, written Hz. The phase says where in its own cycle the wave was at the moment you started watching.
The time one cycle takes is the period. The two are the same fact upside down. 3 Hz means three cycles a second, so one cycle takes a third of a second. That is 333 milliseconds, a millisecond being a thousandth of a second.
Why a smooth wave and not any old squiggle?
Because every other squiggle is made of these. A trumpet, a bark and a cymbal all draw complicated lines, and any of those lines can be represented by adding sine waves of different frequencies, amplitudes, and phases. Fourier analysis uses those components because linear systems respond to each sinusoid without creating a new frequency.
Step 2 does the adding with two of them, which is enough to see how it works. Working out which sine waves a given squiggle is made of is the opposite job, and it has a course of its own further along this stream.
Degrees for a wave? I thought degrees were for corners.
A cycle is a round trip, so it is measured the way a turn is measured, 360 degrees for the whole thing. Ninety degrees is a quarter of the way through a cycle, and 180 is halfway. The reason to use a fraction of a cycle rather than a number of milliseconds is that it stays true when the frequency changes. Half a cycle is half a cycle at any speed; 20 milliseconds is half a cycle only at 25 Hz.
The start point slider in the lab is that fraction. Watch the pill that reports the shift in milliseconds: the degrees stay where you put them and the milliseconds change as soon as you touch the frequency.
What do these frequencies sound like?
The wave in the lab does a few cycles a second so that you can see them, which is far too slow to hear. A young person hears from about 20 Hz to about 20,000 Hz. The A that an orchestra tunes to is 440 Hz. The lowest note on a piano is about 27 Hz, felt as much as heard, and the highest is about 4,186 Hz.
None of the arithmetic cares. A wave at 3 Hz and a wave at 3,000 Hz behave identically in everything that follows, which is why the labs use numbers you can count on the screen.
Two waves in the same air add up
Air carries more than one sound at once, and it does so by the simplest rule available. Suppose one source is pushing the air out at this instant, and another is pulling it back by a smaller amount. The air moves by the difference. The heights add, instant by instant. That is the whole rule. It has a name, superposition. It is worth having because so much later work depends on nothing else happening.
What comes out can look nothing like either part. Add a tall slow wave to a short fast one and you get a slow wave with ripples riding on it. That single line is what reaches your ear, and your ear takes it apart again.
Do two sounds get in each other's way, like two people in a doorway?
No, and that is the surprising part. Two people cannot both be in a doorway. Two waves can share the same air, and afterwards each carries on exactly as if the other had never been there. A shout across a room is not damaged by passing through a violin note on its way.
What you get at any one point is the two added together, but nothing is destroyed by the addition. That is why your ear can pick out one voice at a table where five people are talking, from a single line of pushes arriving at one eardrum.
Half a cycle behind and 180 degrees: the same thing?
Yes. A full cycle is 360 degrees, so half of one is 180. A wave shifted by 180 degrees is pulling exactly when the original is pushing, by exactly the same amount, at every instant.
The word for that in the trade is out of phase. Two waves that agree everywhere, with no shift, are in phase. Everything between the two extremes is partial, which is what the slider in Lab 3 shows.
So why do noise-cancelling headphones not cancel everything?
Because the cancelling wave has to be the exact opposite of the noise, at your eardrum, at the right instant. The headphone hears the noise with a small microphone, works out the opposite, and plays it. Everything in that chain takes time, and by the time the answer is ready the noise has moved on.
Slow rumble is easy: a drone at 100 Hz gives the electronics ten milliseconds per cycle to work in. A cymbal is hopeless, because at 8,000 Hz a cycle is over in an eighth of a millisecond and being a little late is the same as being wrong. That is why the marketing says engine noise and not conversation.
Chop a wave into a list of numbers
A machine cannot hold a curve. It holds numbers, one after another, and that is all it has ever been able to do. So it measures the wave at one instant, writes the answer down, waits a fixed time, and measures again. Each measurement is a sample. Taking them is sampling. How many you take each second is the sample rate, counted in hertz like a frequency, because it is also a number of things per second.
Between two samples, the recording says nothing. Not a small amount, not an average, nothing at all. Whatever the air did in that gap left no trace. Later processing can estimate or interpolate, but it cannot recover information that the samples never captured.
What is actually doing the measuring?
A microphone turns the pushing air into a wobbling voltage, which is an electrical pressure that follows the air exactly. Then a chip called an analogue-to-digital converter looks at that voltage at each tick of a clock and reports a number. Analogue means the smooth original; digital means the list of numbers.
The clock is the important part and it is why sampling is so regular. A crystal ticks at a fixed rate, the converter grabs a reading on each tick, and the readings arrive evenly spaced whether the sound is loud, quiet or absent.
Why is the sample rate in hertz as well? Is it a frequency?
Hertz just means per second. A wave at 300 Hz does 300 cycles a second. A sampler at 300 Hz takes 300 readings a second. They are counting different things, and keeping them apart in your head is worth the effort. From Step 6 onwards the whole subject is the relationship between the two numbers.
Where both appear at once, this course says cycles a second for the wave and measurements a second for the sampler.
Why is a compact disc 44,100 and not a round number?
The choice had two masters. It had to be a little more than twice 20,000 Hz, for the reason Step 7 works out, which rules out anything under about 40,000. And the first digital recordings were stored on video tape recorders, because in the late 1970s nothing else could take that many numbers a second.
So the rate had to fit a whole number of samples into each line of a video picture. It also had to do that on machines built to two different picture standards. 44,100 works for both. It is an engineering compromise between hearing and a tape format, wearing the disguise of a fundamental constant.
The sample rate and the gap between the dots
A list of numbers is not a sound. To hear it again, something has to push the air smoothly once more, and it has only the dots to go on. The simplest rule available is to join them with straight lines. That is close enough to what real equipment does to be worth measuring.
The gap between the joined-up line and the wave that was really there is what the sample rate cost you, and it is a number rather than an opinion. The useful way to count is not measurements a second but measurements per cycle. A wave that wiggles twice as fast needs twice as many readings to be caught equally well.
Two per cycle gives a flat line. Is that a bug?
No, and it is worth staring at. At exactly two measurements per cycle the readings land on the same two points of every cycle. With this wave those two points are both where it crosses the middle. Every reading is zero, so the recording is a row of zeros and the joined line is flat.
Shift the wave slightly and the same two points land somewhere else and you get something back, though never the right height. Two per cycle is the exact edge of what can work at all, and Step 7 is about why that edge is where it is.
Do real players really join the dots with straight lines?
No. A straight line has sharp corners at every dot, and a corner is a rough sound. Real equipment uses a smoother rule that curves between the samples. If the sample rate obeyed the limit in Step 7, there is even a rule that recovers the original wave exactly.
Straight lines are used here because you can see what they are doing and check the answer by eye. The measured gap with a better rule is smaller, but it shrinks with the sample rate in the same way, so nothing in this step changes.
A wheel under a flashing light
Leave sound alone for one step. A wheel with a single white dot painted on its rim is spinning in a dark room. The only light is a lamp that flashes 25 times a second, and between flashes you see nothing at all. Your eye is not watching the wheel. It is being handed 25 snapshots a second, which is sampling, with your eye doing the joining up.
Suppose the wheel turns 24 times a second. Between one flash and the next it gets through 24 twenty-fifths of a turn: nearly the whole way round, but not quite. Your eye has no way of knowing about the nearly-whole turn, because nothing was lit while it happened. All it sees is a dot that has ended up slightly behind where it started.
Is this why wagon wheels go backwards in films?
Exactly this. A film camera takes 24 or 25 still pictures a second and the projector shows them in order, so a cinema audience is sampling the world 24 times a second. A stagecoach wheel whose spokes come round a little slower than one spoke-space per frame appears to roll backwards. At exactly one spoke-space per frame it appears frozen, while the coach charges along.
You can catch it outside a cinema too. Some car headlights and most cheap indoor lights flicker at mains speed, 100 or 120 times a second. A fan or a wheel under them can look still, slow or backwards.
What would make it look right again?
Flash faster. Drag the flashes slider above about 48 for a wheel at 24, and the apparent speed becomes the real speed. A whole turn can no longer hide between two flashes. That is the same threshold as the sound one in Step 7, arriving here in a different costume: you need more than two flashes for every turn.
The other fix is to stop taking snapshots. Turn a steady light on and the wheel becomes a blur, which tells you less about where the dot is and never lies about which way it is going.
Predict what a 7 Hz wave comes back as
The wheel and the wave are the same problem in different clothes. Take a wave doing 7 cycles a second and measure it 10 times a second. Between one measurement and the next it gets through seven tenths of a cycle, and the recording has no way of knowing about the whole cycles it missed in between.
So there is a question worth committing to an answer on, before you see it. Join those measurements back up. What is the slowest wave that passes through every single one of them? Most people say 7 Hz, or a mess, or nothing. Press an answer in the lab first, then reveal it, because being wrong here once is what makes the rest of the course stick.
Why 3 in particular, and not some other slow number?
Because 10 take away 7 is 3. Between measurements the 7 Hz wave gets through 0.7 of a cycle. A 3 Hz wave gets through 0.3 of a cycle in the other direction, and 0.7 forwards leaves the wave in the same place as 0.3 backwards, every single time. The two waves are in the same place at every instant the sampler looks, and in different places the rest of the time, when nobody is looking.
That is the same arithmetic as the wheel: 24 turns a second under 25 flashes leaves 24 minus 25, which is one turn a second backwards. Both are the leftover after the whole cycles that nobody saw.
Can it be undone afterwards?
No, and this is the part worth being firm about. The 7 Hz wave and the 3 Hz wave produce identical lists of numbers, down to the last digit, and Lab 9 measures the difference between them as zero. Nothing that arrives later can separate two things that are the same.
So aliasing is prevented rather than repaired, and it has to be prevented before the sampler, in the electrical world where the fast waves still exist. A circuit that removes everything too fast to be sampled, before it reaches the sampler, is called an anti-alias filter. Step 7 says exactly where to draw its line.
Half the sample rate is the limit
Somewhere between 1 Hz, which came back honestly, and 7 Hz, which came back as 3, the recording started lying. The place where it starts is worth finding rather than being told. So the next lab sweeps every frequency from nothing up to two and a half times the sample rate. It plots what each one comes back as.
Along the bottom is what was really there. Up the side is what the measurements say it was. If sampling were harmless the picture would be a straight diagonal line going up for ever.
Who was Nyquist, and did he work this out with waves or with wheels?
Harry Nyquist was a Swedish-born engineer at Bell Telephone Laboratories, and in 1928 he was not thinking about sound at all. He was working out how many separate telegraph pulses a telephone line could carry each second before they smeared into each other. The answer had the same two in it.
Claude Shannon, at the same laboratory twenty years later, proved the matching statement for recordings. Suppose a signal contains nothing faster than a given frequency. Sample it at more than twice that, and the original can be rebuilt exactly. Not approximately. Exactly. That is why the limit is sometimes called Nyquist-Shannon.
Why is exactly twice not enough?
Try it in Lab 12. The call has a 6 Hz wave in it, and a rate of 12 is exactly twice that. Press Check and it fails. At exactly twice, the measurements land on the same two points of every cycle. Here those two points are where the wave crosses the middle, so every reading is zero. The lab reports the height that comes back as 0.000.
Shift the wave by a quarter of a cycle and you would get the full height back. Exactly twice sometimes works and sometimes gives nothing, depending on something you do not control. The rule is more than twice, and in practice comfortably more. A rate of 44,100 for a 20,000 Hz limit leaves 4,100 Hz of room for the anti-alias filter to work in.
Round every reading to the nearest level
Sampling chopped up time. There is a second chop, and it happens to the readings themselves. A machine cannot store a number to unlimited precision. It has a fixed set of levels spread across the range it can measure and every reading is moved to whichever level is nearest. That moving is quantisation.
The distance between neighbouring levels is the step. Nothing can ever be further than half a step from the nearest level, because if it were, some other level would be nearer. So the worst error is half a step, always, and the way to shrink it is to have more levels.
What is kept is no longer a curve but a staircase, and the difference between the two is the rounding error, which can be drawn on its own.
Is this the same thing as sampling?
No, and they are worth keeping apart. Sampling chops time: it decides when you look. Quantising chops height: it decides how precisely you can write down what you saw. A recording makes both choices. They cost different things and go wrong in different ways.
Sample too slowly and a fast wave comes back as a slow one that was never there. Quantise too coarsely and everything comes back at the right frequency with a hiss laid over it. Step 9 is where that hiss gets measured.
Where do the levels actually come from?
Inside the converter is a ladder of reference voltages, evenly spaced between the lowest and highest reading it can take. A comparator is a circuit that answers only whether one voltage is above another. A set of them works out which two rungs the incoming voltage lies between, and the answer is which rung it is nearest.
That is why the levels are evenly spaced, and why the range has a top and a bottom. Feed in something louder than the top rung and it is pinned to that rung, which is called clipping and sounds like a fuzzy crunch.
One more bit halves the error
Levels are not chosen freely. A machine stores a number as a row of yes-or-no answers, and one yes-or-no answer is a bit. One bit tells two things apart. Two bits tell four apart, because each of the two answers to the first question can be followed by either answer to the second. Every bit you add doubles the count, so the number of levels is 2 doubled as many times as you have bits. How many bits each sample gets is the bit depth.
That gives a chain: one more bit, twice as many levels, half the step, half the worst rounding error. The signal is unchanged while the error halves, so the signal stands twice as far above it as it did.
The error never goes away. Under every recording there is a permanent low hiss made entirely of rounding, and anything quieter than that hiss is lost in it. That floor has a name: the noise floor.
What is a decibel, and why does a doubling always add 6 of them?
A decibel is a way of writing a ratio so that multiplying turns into adding. Instead of saying this signal is 256 times the noise, engineers say it is 48 decibels above it. The useful property is that combining two stages means adding their decibels rather than multiplying their ratios.
The scale is fixed so that a doubling of a wave's height is about 6 decibels, whatever you start from. Twice is 6 dB, four times is 12 dB, a thousand times is about 60 dB. That is why one extra bit is always worth about 6 decibels, and the table in Lab 15 measures that column rather than quoting it.
Why is a CD 16 bits and a studio recording 24?
Sixteen bits puts the noise floor about 96 decibels below the loudest thing the disc can hold. That is quieter than the air in a silent room, so nobody hears it. For a finished recording that is plenty.
While recording, you do not yet know how loud the loudest moment will be, and going over the top clips it and ruins the take. So engineers record well below the top, on purpose, which throws away several bits of range. Starting with 24 means you can waste eight of them on safety and still finish with more than a disc needs.
Measure how much signal is left after noise
Rounding is not the only thing added to a reading. Every real measurement arrives with noise: small random amounts added to each sample by the microphone, the wiring, the warmth of the components, and everything electrical in the building. Noise is different at every reading and unrelated to the signal, which turns out to be the property that matters.
To compare a signal with the noise on top of it you need a fair size for something that spends half its time below the middle line. Adding the readings up gives roughly zero, which is no use. So: square every number, which makes them all positive, average the squares, then take the square root to get back to the original units. Engineers call the result the root mean square, said as it is spelled, backwards.
Do that to the signal, do it to the noise, and divide one by the other. What you have is the signal-to-noise ratio: how many times bigger the thing you want is than the thing you do not.
Why square them first? Could you not ignore the minus signs?
You could, and for some jobs people do. Squaring is preferred because it matches how much energy a wave carries. Two sources of noise combine in a way that the squared measure gets right and the ignore-the-minus-signs measure does not.
It also punishes large excursions more than small ones, which is usually what you want from a measure of how badly something is being messed about with.
Where does the noise come from, if everything is switched off?
From heat. The electrons in any resistor are jostling about because the resistor is warmer than absolute zero, and that jostling is a tiny random voltage across it. It cannot be designed away, only cooled away. That is why the most sensitive instruments run their first amplifier in liquid nitrogen or colder.
On top of that come the avoidable kinds. Mains hum picked up from the wiring, hiss from the amplifier, interference from a phone charger. The rounding error from Step 9 belongs on the list too: it behaves so much like noise that it is measured the same way.
Average many captures and the noise falls
Here is the property of noise that can be used against it. If the same measurement can be repeated, the signal is in the same place every time and the noise is somewhere different every time. Add up ten captures and the signal adds up ten times over, while the noise partly cancels itself, because the positive amounts in one capture meet negative amounts in another.
Divide by the number of captures and the signal is back to its original height with less noise on it. The improvement is not the number of captures, though: it is the square root of that number. The square root of a number is the value that gives it when multiplied by itself, so the square root of 16 is 4.
Four captures halve the noise. Sixteen quarter it. A hundred divide it by ten. Doing twice as well costs four times as much work, every time, for ever.
Where is this actually used?
Everywhere something is too faint to see once. An astronomer photographs the same patch of sky for an hour in short exposures and averages them. A hospital scanner repeats the same measurement many times and averages, which is part of why the machine takes so long and why moving spoils it. An oscilloscope on a workbench has an average mode for exactly this.
Sonar does it too. A faint echo from a long way off is lost in the sea's own noise on one ping, so the ping is repeated and the returns are added up.
What stops this working?
Two things. First, the signal has to be in the same place in every capture. If it drifts, the averaging blurs the signal as well as the noise. That is why a scanner asks you to keep still, and why an astronomer's mount has to track the sky accurately.
Second, the noise has to be different each time. Mains hum at a steady 50 Hz is not random. It sits in the same place in every capture, so averaging preserves it just as faithfully as it preserves the signal. Averaging removes the random and keeps everything repeatable, wanted or not.
Pressure, speed and wavelength
The wave you have been drawing since Step 1 is a drawing of air being squashed and thinned, and it is worth one step to watch that happen. Air is a crowd of molecules. A loudspeaker cone is a wall of that crowd that has started shoving. Push the cone forward and the air in front of it is crowded together: a patch of squashed air, called a compression. Pull it back and it leaves a thinned-out patch, a rarefaction. Each squashed patch shoves the air next to it, which shoves the air next to that, and the pattern runs across the room. The molecules themselves barely travel. It is a shove passing down a queue of people: the shove crosses the whole queue while each person only sways on the spot.
The pattern travels at a speed the air decides, not the sound. In air at 20 °C it is about 343 metres a second, loud or quiet, high note or low. A thunderstorm lets you measure it. The flash arrives almost at once, and the rumble plods along at 343 metres a second. Every three seconds of gap puts the lightning about another kilometre away. Warmer air passes the shove along slightly faster, so the speed creeps up with temperature.
Speed gives a wave a size you can measure with a tape. The period from Step 1 is how long one cycle takes. In that time the pattern moves forward some distance, so one whole cycle occupies a length of air, from one compression to the next. That length is the wavelength. To find it, divide the speed by the cycles a second. A 343 Hz tone in 343 metre-a-second air has a wavelength of one metre. A deep 50 Hz hum is nearly seven metres long, which is longer than most rooms.
Does sound only travel through air?
Anything that squashes and springs back will carry it. Water passes the shove along faster than air, at about 1,480 metres a second, which is part of why whale song carries so far. Steel is faster still, around 5,000 metres a second. In old films a character listens for a distant train through the rail rather than the air, and the trick is real: the rail brings the news first.
The speed changes at a boundary and the frequency does not, so the wavelength changes with it. Some of the wave also bounces back at every boundary, and those reflections are what an ultrasound scanner and a sonar are built to listen for.
Distance, reflections and rooms
Drop a stone in a pond and the ripples spread out in rings, getting lower as they travel. The ripple keeps its energy, but the ring it is spread around keeps growing, so each bit of ring gets less of it. Sound does the same in three dimensions, spreading as a growing sphere. At twice the distance the same push covers four times the area, so the pressure at your ear is half what it was. In the decibels from Step 9, a halving of height is about 6 dB, so every doubling of distance costs about 6 dB.
That clean rule holds in the open, where sound leaves and never comes back. A room is different, because the walls throw it back. One wall far away returns a separate copy, an echo. The walls of a room return thousands of copies, packed so tightly that they smear into a wash of sound. That wash is reverberation, the ringing that follows a shout in an empty stairwell.
Every copy is a wave, and Step 2 said waves add. A reflected copy arrives late, which slides it along against the direct sound. At one frequency the slide lands push on push and the note gets louder. At another it lands push on pull and the note nearly vanishes, which is the walking-past-two-speakers quiz from Step 2 happening off a wall. Which of the two you get depends on where you stand, so the same note is loud in one spot of a room and thin a step away.
Why do cushions quieten a room but do nothing for bass?
Soft porous material soaks up a wave by making the air rub through its fibres, and it works when the material is a reasonable fraction of a wavelength deep. A few centimetres of foam is a reasonable fraction of a 3,000 Hz wavelength, which is around 11 centimetres. It is nothing against a 50 Hz wave nearly seven metres long, which sails through as if the foam were not there.
Taming bass takes thick absorbers, gaps behind them, or boxes tuned to swallow one stubborn note. That is why studios measure a room before treating it: the fix depends on which frequency is misbehaving, and where.
The recording chain
Everything between the air and the list of numbers is a short chain of parts, each with one job. A microphone is a drum skin of its own: the air wiggles it, and it turns the wiggle into a wiggling voltage that copies the pressure exactly. That voltage is tiny, so the next stage is an amplifier that makes it bigger. How many times bigger is the gain, and it is a knob somebody has to set. After the amplifier come two parts you have already built the argument for. First the anti-alias filter from Step 7. Then the converter from Steps 3 and 8, which samples the voltage and rounds each sample to a level. Its trade name is the ADC, for analogue-to-digital converter.
Setting the gain is a squeeze between the two failures you have already measured. Set it too low and the signal sits just above the noise floor of Steps 9 and 10, and turning it up afterwards turns the hiss up with it. Set it too high and the loudest peak gets pinned against the top of the converter's range. That is the clipping from Step 8, and nothing later can unbend a flattened peak. The spare space left between the loudest expected peak and the top is called headroom. A careful engineer leaves several decibels of it, because the loudest moment of a take is a surprise by definition.
Why do stage microphones use those three-pin cables?
The cable carries the signal twice: once as it is, and once flipped upside down. Any interference the cable picks up on its way across a building lands on both copies the same way up. At the far end the receiver flips one copy back and adds them. The signal, flipped twice, comes through doubled. The interference, flipped once, meets itself upside down and cancels, which is the Step 2 cancellation doing honest work.
The arrangement is called balanced wiring, and it is why a hundred metres of microphone cable across a stage full of lighting equipment can arrive clean.
Anti-aliasing and reconstruction
The rule from Step 7 came with a promise attached: sample at more than twice the fastest wave, and everything survives. The room does not make that promise for you. Squeaky machinery, bats and electrical hash all put waves into the microphone that are faster than half of any rate you chose. So the recorder keeps the promise itself. The anti-alias filter's whole job is to remove everything above half the rate before the sampler sees it. It is the reason the workshop squeak in Step 6 is a story rather than an everyday fault.
A real filter cannot cut like a cliff edge. It fades, over a stretch of frequency called its transition band. Below the stretch, everything passes untouched. Above it, everything is gone. Inside it, waves are partly both. The whole stretch has to fit between the fastest wave you want to keep and half the sample rate. That is the practical reason rates sit comfortably above twice: the gap between 40,000 and 44,100 is not generosity, it is where the filter fades.
Playing back has the same problem in reverse. The output converter, the DAC, holds each number steady until the next one arrives, so what leaves it is the staircase from Step 8. The corners of a staircase are fast wiggles that were never in the music. A smoothing filter after the converter rounds the corners off and leaves the wave the samples described.
What if the clock does not tick evenly?
It never ticks perfectly evenly, and the wobble has a name, jitter. A reading taken a whisker early or late is a reading of the wave at slightly the wrong moment. Where the wave is barely moving that costs almost nothing. Where it is climbing steeply, a small error in when becomes a real error in how much.
So jitter matters most for fast, loud signals, and a jitter figure means nothing on its own. Whether a millionth-of-a-second wobble is fine or fatal depends on what is being sampled, which by now is the expected shape of every answer in this course.
Channels, bytes and clocks
A finished recording is just the list of numbers, plus an agreement about how to read it back. You have two ears and they hear different things, so most recordings keep two lists, one for each side. Each list is a channel, and two channels is stereo. Each number in each list is stored in the bits of Step 9. Bits come packed in eights: eight bits is a byte, the unit file sizes are quoted in. Storing sound this way, one plain reading after another, is called PCM, and it is what a .wav file holds.
The cost of a recording is now a multiplication you can do on paper. Numbers a second, times bits per number, times channels, gives bits per second. A compact disc keeps 44,100 numbers a second, 16 bits each, in two channels, which is 1,411,200 bits every second. Divide by eight for bytes, multiply by the length, and you have sized the file before recording a note.
The agreement about reading the list back includes the clock. The numbers say nothing about time except their order, so the player has to tick at the rate the recorder ticked. Play a 48,000-a-second list back at 44,100 ticks a second and every cycle of every wave is stretched. The whole recording comes out longer and deeper, like a singer slowed down to a growl. Play it too fast and you get the squeaky chipmunk version. Time and pitch move together, because both live in the one clock.
Why add noise on purpose before rounding?
Step 8 showed the rounding error following the wave in a repeating sawtooth. To an ear, a repeating pattern is a sound, so at low levels coarse rounding does not hiss, it buzzes along with the music, which is far more noticeable. Adding a whisper of random noise before the rounding breaks up the pattern, so the error turns into a plain steady hiss.
The trick is called dither, and it belongs at the final reduction in bit depth, before the information is rounded away. Once the pattern is baked into the list of numbers, adding noise afterwards only makes a noisy buzz.
Decibels and hearing risk
You have been reading decibels off the pills since Step 9, and here is the whole of the idea. A decibel is not a unit of loudness. It is a way of writing "times bigger" so that multiplying turns into adding. Doubling a wave's height adds about 6 dB, wherever you start from, so ten doublings, which is about a thousand times taller, is about 60 dB. There is one wrinkle to respect. For the energy a sound delivers each second, its power, a doubling adds only about 3 dB, because a wave twice as tall carries four times the energy. Height doublings go in sixes, power doublings in threes, and Lab 29 measures both.
On its own, "20 dB" means nothing. It is "ten times as big" without saying as big as what, so a decibel figure you can trust always names its reference. Sound in the air is measured in dB SPL, against roughly the quietest thing a young ear can catch. A whisper sits near 30, conversation near 60, a rock concert over 100. Recordings are measured in dBFS, against the loudest number the converter can hold. Zero is the top of the range from Step 8, and every honest level is a minus figure below it. Hearing-safety rules are written in dBA, which is dB SPL with the ear's uneven sensitivity folded in. A phone app shows a number too, but unless its microphone has been calibrated the number is a guess.
The safety rules exist because loud sound wears hearing out, quietly and permanently. The widely used workplace guideline puts the limit at 85 dBA for an eight-hour day. Every 3 dB above that doubles the sound energy arriving, so it halves the safe time. At 94 dBA, nine decibels up, three halvings leave one hour.
Why do meters show a peak and an average, and which one matters?
They answer different questions. The average, measured the square-and-root way from Step 10, tracks the energy arriving over time, which is what wears hearing down. The peak catches a single impulse, a hammer blow or a balloon burst, which can be over in a millisecond and still do damage on its own. An average taken over a minute can hide it completely.
That is why a proper measurement records the weighting and the averaging time along with the number. A bare figure with no note of how it was taken cannot be checked or repeated, which by Step 19's standard means it is barely a measurement at all.
Impulse response and correlation
Clap once in a big empty hall and listen. What comes back is not your clap. First the slap off the nearest wall, then the farther ones, then a tail that hisses away to nothing. That reply belongs to the hall, not to your hands: clap anywhere in it and the same reply comes back, because the walls have not moved. The hall's reply to one short, sharp sound is called its impulse response, and it is a complete description of what the hall does to sound.
Complete is meant literally. A recording is a row of samples, and each sample is one little push, a tiny scaled clap of its own. The hall answers every push with a copy of its reply, scaled to match, and the copies add, exactly as waves added in Step 2. So adding up shifted, scaled copies of the reply predicts what the hall does to any recording at all. That adding-up has a name, convolution, and it is how film sound puts an actor recorded in a studio into a cathedral.
The matching tool that goes with it is correlation. Take two recordings of the same sound, one delayed, and slide one along the other, scoring at every slide how well the two line up. The slide with the best score is the delay between them. That is how an echo is timed when it is too buried in noise to spot by eye, and timing echoes is range-finding, which Step 12's divide-by-two turned into distance.
When does the clap trick stop working?
The trick assumes the room treats every sound the same way, loud or quiet, now or in ten minutes. Real rooms mostly do, which is why the method is everywhere. The chain fails it first: clip the amplifier, or ride the gain during the measurement, and the reply you record belongs to that one clap only.
Movement breaks it too. Measure a hall, then open its doors and fill it with people, and the old reply no longer describes it. Careful measurers test twice at different levels and times, and if the two replies disagree they report the conditions along with the answer.
Calibration and uncertainty
Every number in this course was measured by machinery on the page, and you could pick the machinery apart. A measurement of a real room has to earn the same trust, and the test is simple to state. Another person, with your notes, should be able to repeat what you did and get the same answer. That means writing down what question you were asking, what you measured with, where the microphone stood and which way it faced, the gain and the rate. It also means keeping the raw numbers themselves.
It also means knowing how far to trust your own equipment. The way to find out is calibration: measure something whose answer is already known, and see what your chain reports. A kitchen scale is calibrated with a known weight. A microphone is calibrated with a small device that plays one exactly known tone into it. Whatever the chain then reports for that tone becomes the correction for everything else.
Being wrong comes in two flavours, and Step 11 met both without naming them. Random scatter is different on every repeat, so repeating and averaging shrinks it, at the square-root price you measured. A lean, called bias, is the same on every repeat, like a scale that always reads a kilogram heavy, and no amount of repeating touches it. Only calibration catches a lean. What remains after both is written down as an honest range of doubt, the uncertainty. Each source of doubt gets its own number, and the largest is the one worth attacking first.
Where does this stream go next?
The Frequency Domain works out which sine waves a recording is made of, which Step 1 promised was possible. Filters builds the low-pass and equalisation stages this course kept pointing at, as arithmetic on the lists of numbers you now know how to make.
Putting Data on a Wave moves information onto waves on purpose, and Software-Defined Radio points the same sampling and measuring ideas at radio instead of sound.
Your own signal, end to end
Everything in this course is in one place here, with nothing marked and nothing to get right. Build a signal out of one or two waves. Then choose how often to measure it, how many bits to keep each measurement in, how much noise to add, and how many captures to average. The chain runs in that order, which is the order a real recorder runs it in.
Two settings fight each other, and finding the fight is the point. Push the measuring rate down and a wave folds into something slower. Push the bits down and a hiss appears under everything. Push the noise up and only more captures will save you.
Why does the noise go on before the sampling?
Because that is where it happens. Most of the noise is picked up in the microphone and the wiring, before anything has been measured. The sampler measures the signal and the noise together and cannot tell them apart. The rounding error from Step 9 is the exception: that one is added by the converter itself, after the sampling, which is why it is listed separately in the panel.
The order matters for what you can do about it. Noise added before the sampler can only be reduced by better wiring, a colder amplifier or more captures. Rounding error can be reduced by asking for more bits.
What does a real recorder have that this chain does not?
Two pieces, and both are there because of steps you have just done. In front of the sampler sits the anti-alias filter from Step 7, removing everything above half the rate while it can still be removed. This panel has no filter, which is why you can make a wave fold and watch it happen.
The other is stranger. Some recorders add a tiny amount of deliberate noise before rounding, and it is called dither. The rounding error from Step 8 is not random, and a pattern is easier to hear than a hiss of the same size. Trading a small rise in noise for a nastier fault going away is a bargain worth knowing about.
What you can do now
- Draw any steady sound as a wave, and say which of amplitude, frequency and phase you would change to make it louder, higher or later.
- Explain why two loud sounds can add up to silence, and what has to be exact for it to happen.
- Work out how many numbers a recording will take, from its rate and its length.
- Say what a recording knows about the gap between two samples, which is nothing, and what that costs at a given number of measurements per cycle.
- Predict what a wave above half the sample rate will come back as, using nothing but subtraction, and check yourself against a search that fits the samples.
- Choose a sample rate for a signal whose fastest part you know, and say why exactly twice is not enough.
- Say why a wheel goes backwards on film, and connect it to a sound recording without hand-waving.
- Read a bit depth as a number of doublings, and turn it into a noise floor.
- Measure a signal-to-noise ratio rather than guessing it from a picture.
- Say how many repeated captures it takes to make a measurement ten times cleaner, and why the answer is not ten.
- Turn a frequency into a wavelength, and a travel time into a distance, using nothing but the speed of sound.
- Say why the same note can be loud in one spot of a room and thin a step away.
- Set a recording gain that clears the noise floor and still leaves headroom for the loudest peak, and say where the anti-alias filter must sit.
- Size a recording in bits and bytes before making it, and recognise a clock mismatch by what it does to time and pitch together.
- Read a decibel figure critically: which kind of ratio it is, and measured against what.
- Describe a room by its reply to a clap, and time a buried echo by sliding two recordings past each other.
- Tell random scatter from a steady lean, and say which of the two averaging can fix.
Where this goes
- Filters. The anti-alias filter that keeps coming up is one of these. A filter is arithmetic on the list of numbers you now know how to make: average three neighbouring samples and you have built one.
- The Frequency Domain. Step 1 claimed that any wave is a pile of sine waves added together. That course works out which ones, by hand, on eight samples, and then explains why the fast way of doing it changed what computers are used for.
- Putting Data on a Wave. A wave that never changes carries no information. Change its height, its frequency or its phase on purpose and you can send data on it, which is where radio starts. Everything you have measured here about aliasing, bit depth and noise applies unchanged the moment the signal is a radio signal instead of a sound.