Test date: 7 September 2026
Project: Alex Kryve Music & AI Initiative — The Lab
Research independence. This article is not advertising and was not paid for by the developers of any software mentioned. The assessments are based on our own tests with real musical material, not on manufacturers’ demonstrations.
In brief
Can a modern AI transcription tool turn a finished song into a publishable score?
The short answer is: not yet. But it can already shorten the route to a workable notation draft considerably—provided that the musician knows which errors to look for and does not mistake an automatically generated file for a finished result.
For this experiment, we used an excerpt from the chorus of Alex Kryve’s released song “When the Soul Aches” and tested two different approaches:
- converting an isolated vocal to MIDI locally with Spotify Basic Pitch 0.4.0;
- cloud-based transcription in Klangio Transcription Studio, tested separately on the isolated vocal and the full rock mix.
The results were revealing. Basic Pitch offered better control at the level of individual notes; Klangio produced a more readable representation of the vocal phrase and lyrics; and Rock mode estimated the overall metre and tempo much more accurately. Yet every tool made errors serious enough to render its automatic export unsuitable for sale or performance without manual review.
The question behind the experiment
We were not testing an abstract ability to “recognise music.” We wanted to know whether the software could support a specific professional workflow:
audio → editable MIDI or MusicXML → manual musical review → engraved score → PDF.
We examined seven parameters:
- pitch;
- note onsets and durations;
- melismas and vibrato;
- tempo and metre;
- instrument separation;
- lyric recognition and alignment;
- the amount of manual work required after export.
Material and test conditions
The test used the chorus of “When the Soul Aches.” For local analysis, we prepared a 30-second isolated-vocal excerpt beginning at approximately 00:58.9. Klangio’s free demonstration is limited to 20 seconds per transcription, so we used 20-second excerpts from the same section of the song there.
A vocal MIDI containing 69 notes across the range A3–D5 served as the control line. The percentages below therefore describe comparative agreement between the outputs rather than an absolute “truth” of the performance.
The song had already been officially released. Only two short test excerpts—the isolated vocal and the full mix—were sent to the cloud service. Neither the complete master nor the full set of individual source tracks was uploaded.
Test 1. Spotify Basic Pitch: three settings for the same vocal
Basic Pitch is an open-source automatic music transcription system developed by Spotify’s Audio Intelligence Lab. It creates MIDI, supports pitch bend and can process polyphonic sources, although its developers note that it works best on one instrument at a time.
In our test, Basic Pitch 0.4.0 ran locally through ONNX. The audio file was not sent to an external service.
We compared three profiles.
| Profile | Notes detected | Range | F1 at 100 ms | F1 at 200 ms |
|---|---|---|---|---|
| Default | 84 | A3–D♯6 | 48.4% | 66.7% |
| Restricted vocal range | 93 | G♯3–E5 | 51.9% | 69.1% |
| Tuned vocal profile | 62 | A3–E5 | 53.4% | 71.8% |
| Control line | 69 | A3–D5 | — | — |
Here, F1 is not a promotional “accuracy score.” It is our comparative measure of how closely the detected notes matched the control line when note onsets were allowed a tolerance of 100 or 200 milliseconds.
What happened with the default settings
The default profile created 84 notes instead of 69 and extended the vocal range to D♯6. That is unrealistic for this performance. The algorithm interpreted some overtones, breath transitions and vocal vibrato as separate high notes.
In other words, the system heard not only the melody but also the acoustic consequences of performing it. That is understandable for a spectral model; in musical notation, it is an error.
Why restricting the range did not solve the problem
Once the obviously excessive upper register was removed, the result moved closer to the actual range—but the number of events increased to 93. The model stopped creating extremely high notes, yet fragmented sustained vocal sounds even more aggressively into short attacks.
This is a useful example of why a “correct range” does not necessarily mean a “correct part.” The pitches may look plausible while the musical phrase still falls apart.
The best profile
The most stable result came from these settings:
- frequency range: 110–700 Hz;
- minimum note length: 120 ms;
- onset threshold: 0.60;
- frame threshold: 0.35;
- project tempo: 81 BPM.
The number of events fell to 62, false notes in the extreme upper register disappeared, and comparative agreement rose to 71.8% at the 200-millisecond tolerance.
Even the best version retained the familiar problems of automatic vocal transcription:
- unnecessary triplets;
- short rests inside continuous phrases;
- repeated attacks on a single syllable;
- occasional simultaneous notes in a monophonic vocal;
- literal transcription of vibrato rather than musical interpretation of it.
Basic Pitch verdict: a useful source of editable MIDI and a good pitch map, but not a finished vocal score. Its strengths are local processing, transparent parameters and the ability to reduce particular categories of error deliberately.
Test 2. Klangio Transcription Studio: isolated vocal
In Klangio Transcription Studio, we selected Single-Instrument → Vocals → Include Lyrics.
Visually, the result was more convincing than raw MusicXML derived from Basic Pitch. The service created a single vocal staff, placed lyrics beneath it and broadly preserved the recognisable direction of the melody. Most of the chorus lyrics were identified and connected to notes well enough to form a basis for further editing.
Then a structural error became apparent: Klangio interpreted the excerpt as 3/4 at 56 BPM.
The source song is organised in 4/4 at approximately 81 BPM. The problem therefore affected not one duration but the entire rhythmic grid. Notes were placed in the wrong bars, rests acquired the wrong functions, and the natural movement of the phrase was reinterpreted through an alien metre.
This matters far more than a handful of incorrect pitches. A wrong note can be replaced quickly. A tempo and metre error forces the editor to rebuild almost the entire passage: moving barlines, reconnecting durations and rechecking pickups and syllable alignment.
What Klangio did better
- presented the vocal line more cleanly;
- created a readable draft layout;
- recognised most of the English lyrics;
- connected the lyrics to the general sequence of vocal events;
- avoided turning every element of vibrato into as conspicuous a stream of false notes as default Basic Pitch did.
What still requires manual correction
- tempo;
- metre;
- barlines;
- durations and rests;
- the placement of some syllables;
- melismas and sustained phrase endings.
Klangio Vocals verdict: easier to read than unprocessed MIDI, but the error in the metrical foundation prevents it from being considered a more accurate score. It is a good visual draft—and a poor candidate for automatic publication without a music editor.
Test 3. Klangio Rock: full mix
For the second cloud test, the same section was uploaded as a 20-second full mix. We selected the experimental Rock (Beta) mode with automatic instrument detection.
Here, the system behaved differently.
| Parameter | Klangio Rock result | Control |
|---|---|---|
| Tempo | 79 BPM | approximately 81 BPM |
| Metre | 4/4 | 4/4 |
| Parts detected | Guitar, Guitar 2, Bass, Drums | the mix also contains vocals and keyboards |
| Guitar tuning | one whole step down | requires verification against the source tracks |
| Bass tuning | C–G–C–F | requires verification against the source bass part |
Compared with the vocal mode, the metrical estimate was considerably more accurate: a two-BPM discrepancy is small for automatic analysis of a full rock mix, and the 4/4 metre was identified correctly.
The service also produced a plausible harmonic outline for the excerpt:
Am – F – Dm – G – C – F – Dm – E.
This can already serve as a starting point for a simplified accompaniment—for example, a voice-and-acoustic-guitar version.
Instrument separation was more limited. Klangio did not create a vocal or keyboard part, while the two guitar parts appeared nearly identical in several places. This may partly reflect genuine guitar doubling in the arrangement, but exact correspondence also suggests a risk that the same spectral material was duplicated across two output parts.
The bass and drums were reduced to a fairly simple rhythmic skeleton. That is useful for orientation within the form, but insufficient for reconstructing the nuances of a live part: ghost notes, fills, dynamics, articulation and interaction with the vocal phrase.
Rock mode verdict: it understood the overall metrical and harmonic structure better than the other tests, but did not recover the full arrangement and did not demonstrate reliable separation of similar timbres.
One source, three representations of the music
The images below show the opening passage of the same 20-second chorus excerpt. Basic Pitch and Klangio Vocals received the isolated vocal; Klangio Rock received the full mix because its task was to separate the band arrangement. The difference in visual density is therefore part of the result, not a difference in the musical passage being compared.
The notation is shown only in part and is used solely as an illustrative example of the experiment’s results.
| Task | Basic Pitch, tuned profile | Klangio Vocals | Klangio Rock |
|---|---|---|---|
| Vocal pitch contour | Most controllable result | Broadly recognisable | Vocal not separated |
| Rhythmic readability | Fragmented, with extra events | Visually cleaner, but wrong metre | Correct 4/4 and close tempo |
| Lyrics | None | Most lyrics recognised | No vocal part |
| Instruments | One source at a time | Vocal only | 2 guitars, bass, drums |
| Editable result | MIDI; MusicXML can follow | Formats available after unlocking | Formats available after unlocking |
| Test privacy | Fully local processing | Short excerpt sent to the cloud | Short excerpt sent to the cloud |
| Best use | Pitch map and raw MIDI | Vocal draft with lyrics | Tempo, metre, harmony and rhythm-section skeleton |
The central observation is that these tools do not simply differ in “quality.” They create different interpretations of the same audio.
Basic Pitch primarily answers: “Which pitch events are present in the signal?”
Klangio Vocals attempts to answer: “How might this be shown as a vocal staff with lyrics?”
Klangio Rock addresses a third question: “How can the overall spectrum be divided into the typical functions of a rock band?”
None of these is the same as the musician’s question: “How was this part conceived, how is it phrased, and how should it be performed?”
Why one metre error is more dangerous than ten wrong notes
Automatic transcription is often evaluated by the percentage of pitches it identifies correctly. That is not enough for real notation work.
If an algorithm is one semitone out, the editor sees a local problem. If it mistakes an overtone for an extra note, that event can be deleted. But the wrong metre changes the structure of the entire document:
- the phrase is divided in the wrong places;
- strong beats no longer align with the semantic stresses of the lyric;
- syncopations become unnecessarily complex durations;
- melismas cross arbitrary barlines;
- accompaniment and vocal become difficult to align.
For professional work, tempo, metre, pickup and form must therefore be checked before individual notes are corrected.
Protecting the material
Local Basic Pitch processing keeps the source audio on the working computer. For unreleased songs, training materials and individual stems, this is the most cautious option.
Klangio requires audio to be uploaded to its server. According to the current Klangio terms, files are stored temporarily to create the transcription, are not used to train models and are not shared with third parties. Anonymous demo transcriptions are limited to 20 seconds and are scheduled for deletion after 30 days; users may also request deletion of uploaded material.
Even under those conditions, we followed the principle of data minimisation: only short excerpts from an already released work were used. For unpublished material, it is sensible to begin with local processing and use a cloud service only when its additional capability is genuinely required.
There is another important limitation. The free demonstration is intended for testing. It does not provide full editing, while Klangio links commercial use to a complete, unlocked transcription of an original composition. Export options appeared in our interface, but attempting to retrieve a file required registration; we therefore did not create an account or purchase an unlock for this test.
A workflow that proved reasonable
The experiment did not identify a single winner. It revealed a more dependable combination of tools:
- Prepare separate stems. The fewer competing timbres a model hears, the easier its output is to verify.
- Establish the musical foundation manually: tempo, metre, key, pickup and section boundaries.
- Process the isolated vocal through local Basic Pitch or NeuralNote to obtain editable MIDI.
- Filter out physical features of the voice that are not separate notes: vibrato, breath attacks, consonant noise and overtones.
- Use Klangio as a second interpretation, especially for lyric alignment, the harmonic outline and checks on individual instruments.
- Assemble the parts in MuseScore, Dorico or Sibelius via MIDI/MusicXML.
- Carry out manual musical editing: check the melody by ear, restore phrasing, and add breathing, articulation and dynamics.
- Only then engrave a commercial PDF and create a demonstration playback.
The detected chord sequence provides a useful harmonic framework, but on its own it does not convey the picking pattern, bass movement, dynamics or interaction with the voice.
What AI did—and what remains human
In this experiment, AI:
- detected probable pitches and onsets;
- proposed durations;
- recognised some of the lyrics;
- estimated tempo and metre;
- classified instruments;
- built a draft harmonic and notational framework.
The human participants:
- selected the material and control excerpt;
- prepared the isolated vocal and mix;
- defined the comparison criteria;
- created and tuned the processing profiles;
- established the correct tempo and metre;
- compared the outputs with the control MIDI line and source audio;
- identified false notes, fragmentation and duplicated parts;
- judged the usefulness of each result;
- determined the subsequent arranging and editorial work.
The distinction matters. Automatic transcription does not create an authored score with one click. It produces a hypothesis; musical meaning emerges through selection, verification and editing.
Conclusion
Our preliminary ranking changed after the very first real-world test.
Klangio Transcription Studio remains a strong candidate because of its handling of lyrics, multiple instruments and export to professional formats. But its 3/4-versus-4/4 error on the isolated vocal shows that a polished appearance does not guarantee correct musical structure.
Basic Pitch was less visually impressive, but more transparent and controllable. After its range and thresholds were tuned, it produced the best quantitative result for the vocal pitch line while keeping the entire process local. Its weakness is the conversion of a living phrase into an excessively literal set of events.
Klangio Rock mode produced the best estimate of tempo, metre and harmonic direction, but recovered only part of the ensemble and duplicated guitar material in places.
The Lab’s principal conclusion is therefore:
Modern AI can already hear enough to accelerate transcription. But it still cannot decide which elements of what it hears are musically significant.
An automatic system can find notes. A score begins when a human explains why those notes exist, how they relate to one another and how they should sound.
Test material: “When the Soul Aches” — Alex Kryve
Music & Lyrics: Alex Kryve / Oksana Kalinkina
Production: Alex Kryve
Copyright © 2026 Alex Kryve / Oksana Kalinkina