Musicians are hunting AI grifters by ear, and the evidence is thinner than the confidence
A small group of EDM producers has started naming names. They think certain rising acts are passing off Suno output as their own work, and they’re saying so in public, on video and on Threads.
The Verge profiled two of them: Max Harris, who records as H4RRIS, and the Italian producer Nihil Young. Young’s posts about Suno are what pushed Harris to start making callout videos in the first place.
I went looking for what they’re actually using as proof. The answer is their ears, and nothing else.
What H4RRIS and Nihil Young use as evidence


Max Harris, who records as H4RRIS (left), and Nihil Young, real name Mattia Marotta (right). Young’s Threads posts about Suno are what pushed Harris to start making callout videos.
Harris describes a specific tell. Aside from the similar vocals, you can usually hear a sharp hissing throughout the track,
he told The Verge, and he has a theory for where it comes from: models that start with a block of white noise and guess their way toward a waveform.
He has been posting the breakdowns publicly on his Instagram, playing suspect tracks and pointing at what he hears.
He also points to composition. Arrangement choices that strike him as things a human just wouldn’t make,
whole sections treated like one big instrument.
Young’s test is simpler. It becomes obvious that you’re listening to AI-generated music when you put on a good pair of headphones,
he said.
One piece of their evidence does not depend on ears at all. Harris cites the AI persona Lionsaddle, where the character’s fingers vanish in and out of existence on video. Anyone can check that by looking.
Everything else is listening.
Young is not a bystander to this question either. He is Mattia Marotta, 15 years in, past 30 million Spotify streams as an artist and 100 million as a writer and producer, and he released his debut album ‘Compass’ on Zerothree on August 14, 2026. The record is built out of meeting his wife, the birth of his daughter, and the death of his father.
Asked by PLAYY Magazine about the technology, he put the objection in terms that have nothing to do with spectrograms.
AI is the scary part only if we start believing the imitation is enough. Most generated tracks still sound goofy and full of artifacts. They have no stakes.
Hold on to the middle clause. Full of artifacts
is a claim about audio in August 2026, and the rest of this piece is about how long that stays true.
Josh Fawaz and Fenix Flexin show the callouts sometimes land
This is not a witch hunt, and writing it up as one would be dishonest.
Josh Fawaz’s cover of Madonna’s “Like a Prayer” now carries AI credits, added after a wave of online backlash. That same backlash fed into ARIA’s chart ban in Australia.
Rapper Fenix Flexin denied using AI on “Rubberz” until scrutiny forced the admission. Tyga admitted AI on the Starface album. Timbaland never denied it at all.


Two tracks that ended up disclosed. Josh Fawaz added AI vocal and AI drum credits to “Like a Prayer” (left) after backlash. Fenix Flexin’s “Rubberz” (right) carries artwork multiple image detectors flagged as AI-generated.
Both of those covers are worth a second look, because in each case the picture was part of what tipped people off before the audio settled anything.
So the pressure works, sometimes. The problem is what happens when it doesn’t.
Harris named MANSA’s “Midnight on My Mind” and Danny and Ian Asher’s “Take Me (To The Moon)” in a national outlet. None of those artists has publicly commented on whether they used generative tools. No detector result exists. No admission, no denial, no retraction.
Here is “Midnight on My Mind”. Listen to it before you decide what you think, because the whole argument rests on whether an ear can settle this.
I have listened to it several times and I could not tell you. That’s the honest answer, and I produce music.
Hiss was a production choice long before it was an AI tell
Harris’s hiss theory has a real mechanism behind it. It also has a long list of innocent explanations.
Tape hiss and vinyl crackle are load-bearing parts of lo-fi hip hop. Producers sampling old jazz records kept the crackle instead of cleaning it out, and the genre now adds it deliberately, mixed roughly 18 to 24 dB under the master so it reads as room tone rather than an effect. Shoegaze, bedroom pop and plenty of house records lean on noise floors on purpose.
An elevated noise floor also falls out of heavy limiting, lossy codecs, aggressive stem separation and analog-emulation plugins.
The comparison I keep coming back to is the em dash in written English. It became an AI tell in 2025, and it has been a legitimate piece of punctuation for two hundred years. Plenty of writers got accused of using ChatGPT because they punctuate well.
So hiss on its own does not prove anything.
Now the part I’m not going to pretend about, because it would be dishonest to leave it at “could be lo-fi” and walk away. When you hear that hiss sitting on the vocal, the odds it came out of a generator are high. Not certain. High. A lo-fi producer puts noise under the whole mix as a bed. A generator leaves it welded to the voice, because the voice is the part the model synthesized. Those are different artifacts in different places, and a producer who has been listening closely can hear which one is which.
Two years of Suno and Udio output has trained a lot of producers to catch it. That’s real skill, honestly earned, and I’m not going to talk anybody out of trusting their own ears about it.
Where it goes wrong is when the ear is treated as the verdict instead of the trigger. The right sequence runs one way: you hear something, then you verify it with tooling built for the job, then you decide whether to say anything in public. A producer’s ear is the smoke alarm. It is excellent at telling you to go and look, and it was never built to tell you what’s burning.
I keep a running directory of AI music detection tools for exactly this handoff, with what each one checks and what it costs. Which sets up the uncomfortable part. If the ear passes the job to a detector, the detector has to be worth passing it to.
IRCAM Amplify flagged 4.7% of human tracks as AI in an independent test
The detectors are not a safe fallback either, and there are measured numbers for this.
Researchers at KTH tested the commercial detector IRCAM Amplify against 30,000 tracks in The AI Music Arms Race, published in TISMIR. On human-made music from the Million Song Dataset, IRCAM correctly identified 95.3% of it. The other 4.7% it called AI.
IRCAM markets the same product on roughly 99% accuracy and under 1% false positives. The authors spell out what the gap means at catalog scale: a rate like that could result in millions of errors when scaled to music catalogs exceeding 22 million items,
and platforms purging AI content could inadvertently censor human art.
Then there’s how easy the number moves. The same team found that high-pass filtering above 8 kHz was enough to push IRCAM into false classifications, and that a human recording could be read as AI-generated after low-pass filtering or a high-pass around 10 kHz.
Filtering is not an attack. It’s Tuesday in a mixing session.
The generalization results are worse for anyone hoping detection settles this. Classifiers scored above 0.93 F1 when trained and tested on the same generator, and dropped to 0.629 going from Suno to Udio. One open-source detector mislabeled 75.3% of Udio tracks as human, because it had learned Suno.
The paper’s own word for the situation is the one in its title: an arms race.
Which brings the problem back to EDM specifically. If routine filtering moves the needle, then a genre built on synthetic sources, hard limiting, tight quantization and aggressive high-end processing is starting closer to the line than a live folk recording ever will. That’s the genre H4RRIS and Nihil Young are policing, and it’s the genre where both the ear and the tooling have the least margin.
The Fourier paper explains the artifact and dates its expiry
The strongest technical work here comes from Deezer’s research team. In A Fourier Explanation of AI-music Artifacts, presented at ISMIR 2025, Darius Afchar and colleagues proved mathematically that the deconvolution modules inside generative models throw off systematic spectral spikes. A frequency-based detector reading only those spikes passed 99% accuracy in several scenarios.
Read the finding carefully, because the important clause is the one about where the artifact comes from. The spikes are inherent to a chosen model architecture rather than a consequence of training data or model weights.
The tell is a property of how today’s models are built. Change the architecture and the tell moves, or goes.
The companion finding is blunter. In The AI Music Arms Race, researchers built a 30,000-track dataset (10,000 human recordings from the Million Song Dataset, 20,000 from Suno and Udio) and tested detectors against it. A commercial baseline system was, in their words, easily fooled by simply resampling audio to 22.05 kHz.
That’s an export setting. Anyone determined to hide would clear that bar on the first try, which means detection currently catches the careless and misses the deliberate.
None of that makes detection useless. It makes detection a second opinion you run yourself, quietly, before you decide what you actually know.
AI vocals are detectable today, and that is a moving target
Right now, Suno vocals carry processing artifacts a trained EDM producer can plausibly hear. Nihil Young is describing something real when he says generated tracks are full of them. I just don’t think it lasts.
Eleven Music trains on licensed material through deals with Merlin and Kobalt, and it comes out of a company whose core competence was speech synthesis before it was music. The vocals carry less obvious processing than Suno’s. By ear alone, they’re harder to place.
The trajectory elsewhere makes the point better than any audio example. In March 2023, the benchmark for AI video was a clip of Will Smith eating spaghetti, and it was a melting horror show with a stock watermark across it. By May 2025 Google’s Veo 3 rendered the same prompt with working physics and a face that chewed.
Two years, from unwatchable to unremarkable.
Which is why the sharper half of Young’s argument is the half that never mentions audio quality at all. Generated tracks, he says, have no stakes.
His own album is about his father dying. No amount of model progress touches that claim, and it’s the one worth building a career on.
Frequently asked questions
How do H4RRIS and Nihil Young identify AI-generated tracks?
By listening. Max Harris cites a sharp hiss across the track, similar-sounding vocals, and arrangement choices he says a human producer would not make. Nihil Young says AI becomes obvious on a good pair of headphones. Neither reported running a detection tool on the tracks they named.
Has the anti-AI callout culture in EDM ever been proven right?
Yes. Josh Fawaz added AI credits to his Madonna "Like a Prayer" cover after a wave of online backlash, and rapper Fenix Flexin denied using AI on "Rubberz" until public scrutiny forced him to admit it. The callouts have a real hit rate.
Does hiss on a vocal mean a track is AI-generated?
It is a strong signal rather than proof. Lo-fi and shoegaze add noise deliberately, but that noise sits under the whole mix as a bed, while a generator tends to leave it welded to the synthesized vocal. Hearing it is a good reason to verify with a detector, not a reason to accuse someone in public.
How accurate is the IRCAM Amplify AI music detector?
In independent testing published in TISMIR, IRCAM Amplify correctly identified 95.3% of human-made recordings and misclassified the remaining 4.7% as AI-generated, against a marketed figure of under 1% false positives. The same study fooled a commercial detector entirely by resampling audio to 22.05 kHz, which is one export setting rather than a sophisticated attack.
What did the ISMIR Fourier paper prove about AI music artifacts?
Researchers at Deezer showed mathematically that deconvolution layers in generative models produce systematic spectral spikes, and that these come from the model's architecture rather than its training data or weights. A simple frequency-based detector using them passed 99% accuracy on several scenarios.

