Warble Get the app

Warble / Guides / Apps and how they work

How accurate are bird sound identifier apps?

Good, with an asterisk. When the app is right it is impressively right. Here is what the peer-reviewed tests found, and what makes the difference between a clean answer and a wrong one.

Updated · 6 min read · by the Warble team

Good, but not infallible, and the app makers say so themselves. In a 2025 field study in Maine, Merlin Sound ID got 86 percent of its identifications right, against 92 percent for human observers, but it heard far fewer birds than the people did and named 12 species that were never there. A 2025 test of BirdNET in Munich found it could match experts once its confidence threshold was raised. The pattern is the same everywhere: close, clear, common birds are identified well; distant, noisy, overlapping, rare or mimicked ones are where the mistakes live.

What the published tests found

Sound ID apps have been around long enough to be tested properly. Three sources are worth knowing.

Merlin against humans, in Maine (2025)

Researchers at the Schoodic Institute ran 144 paired point counts in coastal spruce-fir forest from July to October 2023: an intermediate-level birder counted for three minutes while a phone running Merlin Sound ID listened alongside. Published in Ornithological Applications, the results were clear. The humans logged 382 detections, 72 percent more than Merlin's 222. When each did name a bird, the precision was close: 92 percent correct for people, 86 percent for the app. Merlin also reported 12 species the expert reviewer never found in the recordings, among them Northern Cardinal, Wood Thrush and Red-tailed Hawk.

Two more details matter in a yard. Humans did much better on distant birds, flyover calls and low-pitched voices; on American Crow the score was 36 human detections to 7 for the app. And when two phones ran Merlin side by side, they produced different species lists 57 percent of the time, which the authors put down to microphone quality and positioning.

BirdNET against experts, in Munich (2025)

A study in PLOS One ran BirdNET over 930 minutes of recordings from a European city and compared it with expert listeners. With default settings the results were mediocre, but once the confidence threshold was raised and the week-of-year filter turned on, BirdNET reached a precision of 0.90 and recall of 0.78, close to expert level. The lesson was in the misses. BirdNET flagged European Robin 2,890 times, far more than the experts heard it, and 63 percent of those hits had confidence scores of 0.3 or lower. Having a person check the infrequent detections improved the results noticeably.

The BirdNET paper itself (2022)

The 2022 paper in PLOS Biology describing the BirdNET app reported more than 1.1 million users and 31 million submissions in 2020 alone, of which 5.8 million observations survived the quality filters. The authors are frank about what shapes accuracy: how vocal each species is by season, ambient noise that differs between habitats, and recognition that is better for some species than others. That is why the app shows the user a confidence score.

Why accuracy depends on the moment

Every one of these apps turns sound into a spectrogram, a picture of pitch over time, and looks for shapes it learned from labeled recordings. Cornell says Merlin needs at least 150 recordings of a species before it will try to recognize it. Anything that blurs the picture lowers the odds.

  • Distance. A far bird is quiet, and quiet means the song sits down in the background hiss. Cornell's own advice is to get closer without disturbing the bird; halving the distance roughly doubles how loud the bird is compared with everything else. The Maine study's biggest human advantage was on distant and flyover birds.
  • Background noise. Road rumble, wind on the microphone, a creek or a mower. The Maine authors note that any sound overlapping a bird in time and pitch can mask the features the model relies on.
  • Overlapping birds. A dawn chorus is several spectrograms stacked on top of each other. Merlin can list more than one species at a time, but the loudest voice wins.
  • Mimics. A mockingbird doing a cardinal produces a spectrogram that looks a lot like a cardinal. The false positives in Maine included several birds that are commonly imitated.
  • Dialects and odd songs. Cornell's FAQ lists an unusual song or call type as a reason Merlin may miss a bird. Songs vary by region and by individual; a mockingbird's repertoire, for example, ranges from 43 to 203 song types depending where it lives.
  • Rare species. Merlin weights its suggestions with eBird sighting data for your location and date, so a bird that is rare where you stand is less likely to be offered even if it is singing. Rare birds also have fewer training recordings, which cuts the other way too: the model knows common birds best.

Three birds that fool apps

If your app keeps announcing a bird you cannot find, one of these three is often nearby.

How to get a better answer from any app

The studies and Cornell's help pages agree on the habits.

  1. Give it a full phrase, then another. Cornell recommends recording for at least 30 seconds, longer if the bird cooperates, so the app hears several complete songs rather than one clipped fragment. Keep it under about ten minutes.
  2. Hold still. Rustling clothes and a moving hand on the phone are loud at the microphone. Stand still, stay quiet, and keep the mic uncovered.
  3. Get closer. Every step toward the bird is worth more than any external microphone, according to Cornell. Just stop before the bird changes what it is doing.
  4. Pick a quiet moment. Wait for the truck to pass. Step away from the fountain. Put your back to the wind so your body shields the phone.
  5. Compare before you believe. Every app shows example recordings. Play them. If the match sounds nothing like your bird, it is not your bird. Cornell's advice is to verify every suggestion yourself.
  6. Read the confidence. A faint bar or a low score means the model is guessing. The Munich study found most bad BirdNET hits sat at the low end of the confidence scale.

Where Warble fits

Warble works the same way in principle: it listens on the phone and compares the sound against 800 US species, no signal required. The one design choice worth mentioning is that when the sound is too faint or too messy, it says it is not sure instead of naming something. A guess that becomes a card in your album is worse than no catch, so on a windy afternoon it will ask you to try again from closer. The habits above help it too.

Get Warble free

Questions people ask

Is Merlin Sound ID more accurate than BirdNET?

No published study has tested both on the same recordings, so nobody can say. In separate 2025 tests, Merlin reached 86 percent precision in a Maine forest and a well-configured BirdNET reached 0.90 in Munich. Both makers say the app suggests and the person confirms.

Why does my bird app keep saying a hawk is nearby?

Often it is a Blue Jay. Jays imitate Red-tailed and Red-shouldered Hawks closely enough that people are fooled too, and Red-tailed Hawk was among the species Merlin reported in the Maine study that were never actually present.

Does an external microphone help?

Less than you would think. Cornell's help page says good technique, meaning getting closer, standing still and recording a full song, makes more difference than a small clip-on microphone. The Maine study did find that different phones heard different things.

Sources

  1. Goodman et al. 2025, Ornithological Applications: Humans outperform Merlin Sound ID in field-based point-count surveys. academic.oup.com/condor/article/127/4/1/8222742
  2. Fairbairn et al. 2025, PLOS One: BirdNET can be as good as experts for acoustic bird monitoring in a European city. pmc.ncbi.nlm.nih.gov/articles/PMC12425287/
  3. Wood et al. 2022, PLOS Biology: The machine learning-powered BirdNET app reduces barriers to global bird research. pmc.ncbi.nlm.nih.gov/articles/PMC9239458
  4. Cornell Lab Help Center: Merlin Sound ID. support.ebird.org/en/support/solutions/articles/48001185783-sound-id
  5. Cornell Lab Help Center: Merlin Sound ID Best Practices. support.ebird.org/en/support/solutions/articles/48001214056-merlin-sound-id-best-practices

Warble is a game, not a field guide. Where this page states a fact about a bird, it comes from the sources above.