Research by Alex Kryve Music & AI Initiative · The Lab
A note from the author
In this article, the technical characteristics of generative audio are intentionally discussed only at the level of general principles.
I will not describe specific methods for removing, altering, masking, or circumventing digital fingerprints, watermarks, metadata, or other mechanisms used to establish the provenance of a file. Nor will this article provide practical instructions for making AI-generated or AI-processed material appear to a detector as material of a different origin.
The reason is simple.
Studying identification mechanisms is necessary if we want to understand the technology. Turning such research into instructions for bypassing those mechanisms is an entirely different matter.
We are interested in the opposite question:
To what extent is it actually possible today to determine how music was created from its sound alone — and where does auditory analysis end and verifiable digital provenance begin?
1. What exactly are we trying to identify?
When someone says:
“I can hear that this music was made by AI,”
the first technical question should be:
What exactly are they claiming to hear?
A piece of music generated entirely by a model?
An AI-generated composition?
A generated vocal?
A generated instrument?
A reconstructed instrumental performance?
AI-assisted production?
AI processing applied to a live performance?
Stem separation?
Pitch correction?
Mastering?
These are fundamentally different processes.
A modern recording can simultaneously contain a human composition, live vocals, live instruments, MIDI, samples, amp simulation, drum replacement, Melodyne or Auto-Tune, software reverbs, generative elements, AI-assisted processing, and automated mastering.
The binary question:
“AI or not AI?”
is therefore becoming increasingly uninformative.
A far more precise question is:
What exactly did AI do?
2. What technically distinguishes live performance from generative audio?
There is no single universal characteristic.
There is, however, a collection of characteristics that can be studied statistically.
| Characteristic | Live performance | Possible characteristics of generative audio |
|---|---|---|
| Microtiming | Small, musically meaningful deviations from the timing grid | May be unusually regular or contain statistically unusual deviations |
| Dynamics | Attack intensity follows phrasing and the physics of performance | May sometimes display excessive averaging or unusual dynamic structures |
| Transients | Attack corresponds to a specific physical event | Attacks may sometimes appear smoothed or reconstructed |
| Harmonics | Governed by the physics of the instrument and method of excitation | Unusual spectral relationships may sometimes occur |
| Pitch | Natural microvariations in frequency | May exhibit excessive stability or unusual micro-instability |
| Noise floor | Noise has physical causes: room, amplifier, breathing, fingers | Background structure may sometimes change without an obvious physical cause |
| Stereo / Phase | Spatial relationships correspond to sources, microphones, and acoustics | Statistically unusual spatial or phase relationships may occur |
| Reverb | Reflections correspond to a particular acoustic or simulated environment | Spatial characteristics may sometimes change with the content |
| Timbre over time | Connected to one physical instrument and recording chain | Gradual timbral drift may sometimes occur |
| Instrument interaction | Microtiming and dynamics between musicians are interrelated | Correlations may have a different statistical structure |
But this is where the most important point begins:
None of these characteristics, by itself, proves the use of AI.
3. Microtiming: human “inaccuracy” that is actually a system
Imagine a drummer playing steady quarter notes.
A mathematical grid might look like this:
0 / 500 / 1000 / 1500 ms
A human musician might play:
+8 / −4 / +11 / +3 ms
But what makes human performance interesting is not simply the existence of error.
A good musician does not deviate randomly.
A drummer may subtly push the music forward before a chorus.
A bass player may sit slightly behind the kick.
A guitarist may attack slightly ahead of the snare.
These microscopic relationships create groove.
For technical analysis, therefore, the interesting question is not:
“Are there deviations from the grid?”
but:
“How are those deviations correlated with one another?”
Even here, however, there is no absolute answer.
Modern generative systems can model microtiming.
And a live drummer can be completely quantized.
This produces a paradox:
A human can technically sound more “machine-like” than a machine.
4. Performance dynamics
A person rarely plays a sequence of notes with exactly the same energy.
A simplified representation of a musical phrase might look like:
72 — 76 — 81 — 88 | 69 — 74 — 79 — 91
The changing energy follows the direction of the phrase.
A generative system can also reproduce dynamics. But the statistical structure of those changes may sometimes differ from a live performance.
The problem is that compression, limiting, transient shaping, and mastering can transform the original dynamics so extensively that those distinctions become difficult to observe.
Good LUFS, RMS, or crest-factor values therefore tell us very little, by themselves, about the origin of a performance.
5. Transients: the moment a sound is born
A drumstick striking a drum produces a physical sequence:
stick → membrane → shell → snare → room → microphone.
A guitar-string attack involves:
pick → string → fret → pickup → amplifier → cabinet → microphone.
Every stage leaves temporal and spectral information.
A generative system creates a plausible representation of the result of this process.
At sufficiently detailed levels of analysis, differences in attack and decay structure may sometimes be observed.
But once again, modern studio production complicates the picture.
Drum replacement, transient designers, amp simulation, saturation, compression, and mastering can radically alter a genuine physical attack.
Therefore:
An unusual transient is something worth investigating — not proof of provenance.
6. Spectrum and harmonics
A physical musical instrument obeys acoustic laws.
Its fundamental frequency and harmonics are related.
For a hypothetical 110 Hz note, we might observe components around:
110 → 220 → 330 → 440 → 550 Hz…
But frequency alone is not the interesting part.
What matters is how the amplitudes of those harmonics evolve over time.
In a real string, harmonics change because of physical processes: decay, interaction with the instrument body, pickups, amplification, microphones, and the room.
Generative audio may sometimes exhibit spectral smearing, unusual harmonic evolution, or instability in high-frequency components.
But generative models continue to improve rapidly.
A spectrogram can therefore provide material for probabilistic analysis.
It cannot tell us:
“This frequency is AI.”
There is no universal AI frequency.
7. Space, stereo, and phase
A physical recording has geometry.
There is a sound source.
There is a room.
There is a microphone.
There is distance.
The left and right channels therefore exist in particular temporal and phase relationships.
Generative audio can sometimes contain spatial relationships that are difficult to explain through a physical arrangement of sources.
Cymbals, reverbs, distorted guitars, backing vocals, and ambience can be particularly interesting areas for analysis.
But again, any conclusion has limits.
Stereo widening, artificial reverb, chorus, delay, Mid/Side processing, and modern mastering can produce extremely complex phase behaviour from a recording that began entirely with human musicians.
8. Timbral stability over time
This is one of the more interesting areas of analysis.
If a musician records a Stratocaster through a particular amplifier and recording chain, forty seconds later that Stratocaster is physically still the same instrument.
In generative audio, very small changes in the spectral identity of a source may sometimes occur.
A listener may continue to perceive it as the same guitar.
Technical analysis over time may reveal gradual timbral drift.
Similar effects may occur with vocals: relationships between formants, harmonic-to-noise characteristics, sibilants, or breathing structures can change.
Once again:
This is a diagnostic characteristic, not proof.
9. Why a combination of characteristics matters more than any single one
This is why serious detection systems cannot reasonably depend on one parameter.
Their task is to analyse many characteristics simultaneously and estimate the probability that material belongs to a particular class.
This is fundamentally different from a human saying:
“I can just hear it.”
By June 2026, Deezer publicly reported 99.8% detection accuracy for fully AI-generated music using its detection technology.
The wording matters:
100% AI-generated music.
A fully generated track and a hybrid production are not the same technical category.
10. How well can humans identify it?
This is where the results become especially interesting.
In the Deezer/Ipsos blind-listening study, 97% of participants were unable to correctly distinguish fully AI-generated music from human-made music.
An even more interesting study was published in Frontiers in Psychology on August 26, 2026.
The experiment involved 71 participants with musical training, evaluating 18 excerpts from three categories:
- AI-generated;
- human–AI collaboration;
- human-composed.
This produced 1,278 individual provenance judgments.
The results were remarkable.
For AI-generated material:
31.2% classified it as AI-generated; 36.4% classified it as human–AI collaboration; 32.4% classified it as human composition.
For human–AI collaborative music:
32.4% — AI; 36.6% — collaboration; 31.0% — human.
Even fully human compositions were correctly identified as human in only 45.3% of judgments.
That means 54.7% of genuinely human-composed excerpts were incorrectly classified as either AI-generated or human–AI collaborative.
The statistical association between actual provenance and listener classification existed, but it was very weak: Cramér's V = 0.095.
That alone should make us considerably more cautious about statements such as:
“I can hear AI with absolute certainty.”
11. The stranger effect: we may begin hearing AI after being told that it is there
There is another side to the problem that may be even more interesting.
In two preregistered studies involving a combined 399 participants, researchers examined whether perceived authorship changes the way people experience music.
In one experiment, musical excerpts were presented with either a Human or AI label.
Participants were less likely to imagine a story and evaluated the internal narrative they experienced as less rich when the music was labelled as AI — regardless of its actual origin.
That is an extraordinarily important result.
The label did not change the WAV.
It did not change the composition.
It did not change the performance.
It changed the context of perception.
12. Does that make technical detection meaningless?
No.
It means something very different.
We need to distinguish at least four levels.
1. Auditory perception
“I think this sounds like AI.”
A subjective judgment.
2. Acoustic analysis
“The signal contains technical characteristics statistically associated with a particular production method.”
A technical observation.
3. Detector classification
“A particular model estimates an X% probability that the material belongs to a certain category.”
The output of a specific model.
4. Provenance
“Verifiable information exists concerning the origin and history of the digital object.”
This is an entirely different category of evidence.
13. Metadata, fingerprints, and watermarks are different things
These terms are frequently confused.
Metadata is information associated with a digital object.
A fingerprint is a signature derived from or associated with content that can help identify it or related versions.
A watermark can embed information directly into digital content in a form designed to function as an identifier.
Modern provenance systems can combine several layers.
C2PA, for example, distinguishes cryptographic hard bindings from soft bindings, which can include mechanisms such as fingerprinting and invisible watermarking.
The distinction is fundamental:
A detector tries to infer provenance.
A provenance system tries to preserve information about provenance.
These are different problems.
14. Why I consider the official Download function critically important
For creators working with generative platforms, I would formulate one very simple practical rule:
The result of your work with a platform should be obtained through the platform's officially provided Download function, and the originally downloaded file should be preserved.
Suno provides a particularly interesting example.
Under its Terms scheduled to take effect on September 3, 2026, Suno explicitly connects permitted commercial use of qualifying Output with obtaining a permitted Download through an approved Suno download channel.
The Terms also prohibit obtaining copies through alternative means such as stream ripping or recording playback.
More importantly, Suno reserves the right to append to Output a:
fingerprint, watermark, or metadata
indicating the applicable service tier and whether the Output was a permitted Download.
The Terms also prohibit removing, obscuring, altering, or circumventing such information for the purpose of concealing or misrepresenting provenance, service tier, or Output status.
The Download button therefore should no longer be thought of merely as:
“Save WAV.”
It can form part of the provenance and legitimate acquisition chain of the file.
15. What exactly is stored inside the file?
Here we need to be particularly careful.
We can state only what the available documentation supports.
Suno explicitly refers to the possibility of indicating:
- provenance;
- service tier;
- permitted Download status.
But its publicly available Terms do not justify claiming that every downloaded audio file necessarily contains encrypted versions of the user's name, address, payment information, or other personal information.
That distinction matters.
Modern provenance does not necessarily require publishing a person's real-world identity inside the media file.
C2PA specifically addresses privacy and supports provenance systems that do not require disclosure of a creator's personal identity.
A more defensible conclusion is therefore:
A platform may possess considerably more information about the creation and acquisition history of an asset than a user can see by inspecting the ordinary metadata of a WAV or MP3 file.
The file itself, server-side records, the user account, fingerprints, watermarks, and Download history can represent different layers of the same provenance ecosystem.
16. Why would platforms need this information?
There are several reasons.
Provenance protection
A digital object can be connected to a particular creation process or system.
Rights disputes
Creation history, account records, and permitted Output acquisition may become relevant evidence in future disputes.
Prevention of large-scale abuse
Platforms have an interest in distinguishing normal creative workflows from automated mass extraction or misuse.
Usage statistics
Platforms need to understand which tools are being used, how they are used, and at what scale.
Product improvement
Usage data can help platforms study real workflows and improve models, interfaces, and production tools.
Trust
This is the broader problem addressed by systems such as Content Credentials: rather than trying to guess provenance from appearance or sound, establish a verifiable history of the digital object.
17. The hardest category: hybrid production
Imagine a recording where:
lyrics — human; composition — human; vocals — live; guitar — live; bass — live; drums — live;
but the material subsequently passes through AI-assisted arrangement, reconstruction, processing, or mastering.
What percentage of that song is AI?
10%?
40%?
60%?
The question itself may be wrong.
There is no universal physical unit called:
“one percent of artificial intelligence.”
A percentage returned by a public detector is not necessarily the percentage of creative contribution made by AI.
It may simply represent the model's classification of characteristics found in the final signal.
Those are fundamentally different things.
18. My own unfinished experiment
This is exactly why I became interested in testing the problem in practice.
I have a song whose original performance was entirely sung by me and played by human musicians.
The recording later underwent processing involving AI tools.
I deliberately disclosed that involvement.
But I became curious about taking the experiment further.
I submitted the material to one publicly accessible AI detector and received an estimate of approximately 50–60% AI.
To me, this does not mean:
“Half of the song was created by artificial intelligence.”
I know the history of that recording.
What the result tells me is something different:
The detector found characteristics in the final signal that its model associates with AI.
The experiment is not finished.
For that reason, I do not want to draw final conclusions from it yet.
I want to see where it leads.
19. The next experiment
Now I am tempted to conduct an even simpler experiment.
Write a song entirely myself.
Play everything myself.
Sing it myself.
Document the complete creation process.
And then place an AI-related label on the finished recording.
Not in order to permanently conceal its origin.
Quite the opposite.
The true provenance would be documented in advance and disclosed after the experiment.
But for a period of time, the audience would know only one thing:
AI was used.
And then the truly interesting part begins.
20. What will society hear?
Will people find AI where it is not?
Will someone say:
“You can immediately hear that these drums are AI,”
while I have the recording documenting that I played the part myself?
Will somebody hear:
“That characteristic AI guitar,”
when it was played by my own hands?
Will someone call my own voice synthetic?
Will technically experienced listeners explain exactly which AI characteristics they believe they are hearing?
Will attitudes toward the song itself change?
Will the comments change?
Will emotional responses decrease?
Will evaluations of quality change?
And perhaps most interestingly:
Will people change their opinion of the music after learning the truth?
21. And then reveal the provenance
The second half of the experiment would be more important than the first.
Show the original recordings.
Show the performance process.
Show the provenance of the material.
And say:
Here is the vocal.
Here are the drums.
Here is the guitar.
Here is the musician.
Then we can observe an entirely different reaction.
What happens to the person who, five minutes earlier, was explaining with absolute certainty why the drums were AI?
Do they reconsider?
Do they acknowledge the mistake?
Do they search for another explanation?
Do they relocate the “AI characteristic” to another element?
Or do they insist that AI must somehow still be present?
22. At this point, the experiment stops being about artificial intelligence
The object of the research changes.
Now the object is:
us.
Our perception.
Our expectations.
Our prejudices.
Confirmation bias.
Our trust in labels.
And the remarkable human ability to form a belief first and then discover evidence for that belief in what we perceive.
Existing research suggests this is not merely a philosophical possibility.
In the studies involving 399 participants, an AI label itself reduced the tendency to experience human narrative in music, regardless of the music's actual provenance.
Meanwhile, research involving musically trained listeners demonstrates the opposite side of the same problem: even without a label, humans are remarkably unreliable at determining the true origin of musical material.
Put those findings together and an extremely interesting hypothesis emerges:
We are becoming less capable of identifying AI from the music itself, while information telling us that AI is present can still change what we believe we hear in that music.
23. What can actually be established today?
We should therefore distinguish four very different statements.
| LevelWhat we actually know | |
|---|---|
| “I hear AI” | A listener's subjective impression |
| Acoustic analysis | Certain technical characteristics exist in the signal |
| AI detector: X% | A specific model classified the signal with a particular probability |
| Provenance | Verifiable information exists concerning the origin and history of the object |
The first does not replace the second.
The second does not replace the third.
And the third does not replace the fourth.
Most importantly:
A probability returned by a detector is not a percentage of AI authorship.
24. Why the phrase “I can hear AI” amuses me
Not because human hearing is useless.
Quite the opposite.
An experienced musician or audio engineer can hear extraordinarily subtle details of performance and production.
What amuses me is something else:
the confidence with which a subjective impression is sometimes transformed into a statement about the provenance of a work.
Someone can hear a bad transient.
But they do not know whether it appeared during generation, compression, or mastering.
They can hear perfectly aligned drums.
But they do not know whether those drums were generated or whether a human drummer was quantized.
They can hear an unusual guitar.
But they do not know whether it was generated or played by a human and subsequently passed through amp simulation and processing.
They can hear an unnaturally clean vocal.
But they do not know whether they are hearing voice generation or a real singer after Melodyne.
We hear the result.
We do not hear the history of the file.
25. Someone who understands the tool can move in both directions
This may be one of the most important conclusions.
The more deeply a musician understands a generative tool, the less AI becomes a button that says:
“Make me a song.”
It becomes another instrument inside the production environment.
And it can be used in completely different directions.
You can preserve human irregularity.
You can destroy it.
You can make generated material more organic.
You can intentionally make a live performance more mechanical.
You can preserve a live vocal almost untouched.
Or you can process a real vocal until a listener becomes convinced that it must be synthetic.
Someone who understands the principles behind contemporary tools can therefore move in either direction, depending on the creative objective.
And at that point, attempting to establish the provenance of a recording purely by listening begins to lose its meaning.
26. Perhaps we are asking the wrong question
The debate around AI music has spent too much time asking:
“Was AI used?”
Contemporary production increasingly makes that question insufficient.
More useful questions are:
Who wrote the work?
Who performed it?
Which elements were generated?
Which elements were transformed?
What did the human create?
What did the algorithm do?
What verifiable provenance exists?
And finally:
What exactly did AI do?
Conclusion
Perhaps we are already approaching a point where the idea of a universal “AI sound” simply does not exist as a reliable acoustic marker.
There are statistical characteristics.
There are artifacts.
There are detectors.
There are fingerprints.
There are watermarks.
There is metadata.
And there are provenance systems.
But these are different layers of information.
Blind-listening research demonstrates just how unreliable human provenance judgments can be: in a recent study, even musically trained participants misclassified 54.7% of fully human compositions, while AI-generated and human–AI collaborative material was classified at rates approaching chance.
At the same time, research suggests that the simple presence of an AI label can alter our perception regardless of the true origin of the music.
Perhaps the next stage of the discussion about AI and music should therefore begin not with an attempt to prove:
“I can hear AI.”
But with a much more serious question:
“What do I actually know about the provenance of this recording — and what did I merely decide to hear after somebody told me how it was made?”
That is exactly why I want to conduct this experiment.
Not to prove that artificial intelligence has defeated the human musician.
And not to prove the opposite.
But to observe what happens to our perception of music when what we know about its origin and what we believe about its origin stop being the same thing.
Perhaps the result of this experiment will tell us much less about artificial intelligence than we expect.
But considerably more about ourselves.
Alex Kryve Music & AI Initiative The Lab · August 2026