A bird song app turns the microphone signal into a spectrogram, a picture with time running left to right and pitch running bottom to top. A model trained on many thousands of labeled recordings has learned what each species' picture looks like, much as it might learn faces. The app slices the incoming sound into short windows, scores each window against every species it knows, and shows you the birds that score above a confidence line. Some apps do this on the phone itself; others send the recording to a server and wait for the answer.
Step one: the sound becomes a picture
A microphone records one wobbling line: air pressure going up and down thousands of times a second. Nothing, human or machine, reads that directly, so the first thing every sound ID app does is redraw it as a spectrogram.
A spectrogram is a chart. Time runs along the bottom. Pitch runs up the side, low sounds at the bottom, high at the top. Loudness becomes brightness: a strong note is a bright mark, background hiss a faint smear. The computer builds it by chopping the recording into tiny overlapping slices, working out which frequencies each slice contains, and standing the results side by side. The same picture has been called a sonogram or a voiceprint, and biologists use it to study animal calls.
Once you have seen a few, you can read them by eye. A clear whistle is a thin horizontal line. A slurred note bends. A trill is a comb of short marks. A harsh chatter is a smudge that fills the whole height. Merlin shows a simple black and white version as you record, and Cornell's help page says the model is doing what you do when you look at it: recognizing shapes and patterns.
Step two: a model learns what each species looks like
The model is a neural network, the same family of software that finds faces in photos. The BirdNET team describes theirs as a deep convolutional network, a mouthful for a program that spots patterns in images. Nobody writes rules like if the line bends upward, it is a cardinal. The network is shown a huge number of spectrograms that a person has already labeled, and it adjusts itself, a little at a time, until its guesses match the labels.
The labeled recordings come from archives birders have been filling for years. BirdNET's 2022 paper credits millions of recordings from Cornell's Macaulay Library and from Xeno-canto, contributed by thousands of volunteers. Cornell says Merlin needs at least 150 recordings of a species before it will try to learn it, which is why some species are covered and others are not yet.
Three birds with very different pictures show why this works.
Three songs that draw very different pictures
Step three: scoring short windows
The model does not listen to a whole minute and deliver one verdict. It looks at short windows of sound, one after another, and gives each window a score for every species it knows: high if the picture looks like that species' training examples, low if not.
The app then applies a threshold. Species above the line are shown; species below it are not. That line is a design choice. Set it low and the app names more birds, including some that are not there. Set it high and it stays quiet more often but is right more often when it speaks. BirdNET shows a confidence score with each result and lets you mark whether it was correct. Merlin lists the birds it heard and plays a reference recording so you can compare; Cornell is clear that the app suggests and the person confirms.
Most apps also fold in where and when you are. Cornell says Merlin matches its suggestions against the species expected in your region, so a bird never seen in your county in February starts with a handicap.
Why it wants a clean phrase
All of this depends on the picture being readable. A distant bird draws faint lines that sink into the haze. A passing truck fills the bottom of the picture and hides low notes. Three birds at once means three pictures stacked. A rustling jacket is a smear on top of everything.
That is why Cornell's advice is boring and correct: get as close as you can without disturbing the bird, keep the microphone uncovered, stand still, stay quiet, and let the bird finish several full songs rather than one clipped fragment. More complete phrases mean more windows to score and a chance for the model to check its own answer.
On the phone or on a server
There are two ways to run the model. In the first, the phone sends the audio to a computer somewhere else, which runs the network and returns the answer. The BirdNET app was built this way: the 2022 paper describes users transmitting audio plus anonymized metadata to the BirdNET server, where the identification happens. You need a signal, and every recording leaves your phone.
In the second, the model is shrunk down and shipped inside the app, so the phone does all the work. Cornell says Merlin Sound ID runs on the device and works without cell service, and that saved recordings stay on your phone unless you share them. The BirdNET project now also offers embedded versions of its model for phones and low-power devices. Warble takes this approach too: it does the comparison on the phone, against 800 US species, and turns the result into a card for your album. In a canyon with no bars, that is the difference between an answer and a spinner.
The same trick works for frogs, bats and whales
Because the model only ever sees pictures, it does not care that they came from birds. A 2023 study in Scientific Reports took networks trained on bird song, BirdNET among them, and tested them on four North American bat species, 32 kinds of whales, dolphins and seals, and a set of frog calls. The bird-trained models sorted those sounds better than models trained on general audio, which suggests that learning to read bird spectrograms teaches a network something general about animal sound. The BirdNET group even lists a web tool called Ribbit for frog identification.
So when your phone names a Wood Thrush, it is using the method ecologists now use to count frogs in a marsh at night.
Questions people ask
Does the app compare my recording against a library of bird songs?
Not directly. The library is used once, during training, to teach the model what each species' spectrogram looks like. When you press record, the model recognizes a pattern it learned rather than searching files, which is why an on-device app can answer with no connection.
Why does it name birds that are not there?
The model only reports which species' picture your window most resembles. A mimic, an overlapping song or a noisy fragment can resemble the wrong species closely enough to cross the confidence line. Cornell's advice is to play the reference recording and confirm the match yourself.
Sources
- Wikipedia: Spectrogram. en.wikipedia.org/wiki/Spectrogram
- Cornell Lab Help Center: Merlin Sound ID. support.ebird.org/en/support/solutions/articles/48001185783-sound-id
- Wood et al. 2022, PLOS Biology: The machine learning-powered BirdNET app reduces barriers to global bird research. pmc.ncbi.nlm.nih.gov/articles/PMC9239458
- Ghani et al. 2023, Scientific Reports: Global birdsong embeddings enable superior transfer learning for bioacoustic classification. pmc.ncbi.nlm.nih.gov/articles/PMC10739890/
- Cornell Lab and Chemnitz University of Technology: BirdNET project site. birdnet.cornell.edu/