Back to MasterForge MasterForge Blog

Suno V6 Runs Out of Steam:
Why the Last Minute Falls Apart

Petri Korhonen  ·  September 2026  ·  18 min read  ·  Analysis

Everyone who has generated a full-length song on Suno V6 has heard it. The first minute is fine, sometimes better than fine. Then the mix starts to shrink. By the last chorus the stereo image has folded into the middle, the top end is gone, and the whole thing sounds like it is playing from a plastic bucket. The words go soft. The ending arrives as if the model no longer knew how to finish a song.

The forums call it running out of tokens. Suno's answer is Max Mode, which costs double credits and is recommended for anything over two minutes. Since 3 September a Pro account also gets twenty downloads a month, so a usable three-minute track now costs two generations out of a budget that was already small. That makes the question worth answering properly: is the drift real, is it V6, and does Max Mode fix it?

We measured it. Twenty-six full tracks, every one split into stems, every one analysed ten seconds at a time from the first bar to the last. Half of them are V6. The other half are v5 and v5.5 tracks generated from the same prompts back when those models still existed, which gives us the one thing a complaint thread never has: a control.

The short version

The drift is real and it is V6-specific. Over the length of a track V6 narrows, pushes its bass to mono and darkens, its neural artifacts pile up toward the end, and in heavy genres the lyrics fall apart. v5 and v5.5 tracks from the same prompts hold or even open up. Max Mode does not fix it, by meter or by blind listening, and a one-minute generation does not escape it either.

Line chart over the position in the track: how far the music has drifted from its own first minute, as heard by the MERT music model, for Suno v5, v5.5 and V6. The V6 line separates from the older models at about 40 percent of the track and ends twice as far away. The 10 to 40 percent region is marked as usable
The whole article in one picture. Each line is how far a song has drifted from its own first minute, as heard by a music model, averaged over every track of that model generation. Up to about 40 percent of the way through, V6 (cyan) behaves like v5 and v5.5. From there it drifts away and never comes back. The usable part of a V6 song is roughly the 10 to 40 percent stretch. Click to enlarge

Hear it first: Kings of the Northern Star, Max Mode off

Epic rock, V6. Thirty seconds from the first minute against thirty seconds from the last, loudness matched. No mastering, no processing of any kind: this is the file Suno delivered.

This is the track that started the whole investigation. Same song, same take, one minute apart in either direction. Listen for the width of the guitars and the air above the vocal, then for what is left of both at the end.

Warning: this is an example of V6 failing, not a demo of good sound

The "Last minute" clip is meant to sound bad. It is the raw, untouched Suno V6 export, and it is what this article is about. Nothing has been done to it and nothing here is trying to sell it. Compare it with the "First minute" clip of the same song: same take, same singer, one minute apart. The collapse in width, air and clarity between the two is the whole story.

What we measured, and against what

Two sets of tracks. The first is five new V6 songs in five genres, each generated twice with the same prompt, lyrics and settings, once with Max Mode on and once with it off: a metal track, a rock track, a Finnish pop ballad, an epic rock track and a pop song. The second is the five prompts from our V6 launch article, which exist as v5, v5.5 and V6 versions of the same song, plus a v5.5 version of the metal track. Twenty-six tracks in total, fifteen of them V6.

Every track went through the same pipeline. It was split into six stems with a neural separator, so the vocal, the drums, the bass and the accompaniment could be measured on their own. The mix and every stem were then analysed in ten-second windows, five seconds apart, from start to finish. For each measurement we asked one question: does it trend over the length of the track, and does the trend differ between V6 and the older models?

That last part is the control, and it matters more than any single number. Every song gets louder and denser toward its final chorus, on any model, because that is how songs are built. A measurement that only says the ending is different from the beginning proves nothing. A measurement that says V6 endings are different in a way v5.5 endings are not is a finding.

Nerd box: the machine that did the listening

All twenty-six tracks were separated, transcribed and embedded on a workstation we call Somnus: a 24-core AMD system with two AMD GPUs, a Radeon AI PRO R9700 with 32 GB and a Radeon RX 7900 XT with 20 GB, running ROCm 7.2. Stem separation with BS-RoFormer took 43 to 90 seconds per track. The same job on the CPU of our production analysis server takes twenty-two minutes, so the GPU is roughly twenty times faster, which is the difference between an afternoon and a fortnight. Whisper transcribed a vocal stem in seven seconds, MERT embedded a whole track in just over one. Consumer AMD hardware is now entirely adequate for this kind of audio science, and nobody had to ask NVIDIA for permission.

If you enjoy this sort of thing, the same machine keeps a blog about AI science and mathematics at sf3d.fi. Warning: extremely dry mathematical content, occasionally moistened by a joke.

Finding 1: the mix narrows and darkens as the song goes on

Start with stereo width, measured as one minus the correlation between the left and right channels, so that zero is mono and one is fully wide. For each track we asked whether width trends downward over time. On V6 it does: eleven of fifteen V6 tracks narrow steadily over their length, against two of eleven older tracks. Averaged over the group, the trend on V6 is clearly negative and on v5 and v5.5 it is zero. The difference is statistically solid for a sample this size.

The size of the effect is easier to grasp per track. On the metal track the width fell from 0.28 in the first minute to 0.16 in the last with Max Mode off, and from 0.25 to 0.12 with it on. Kings of the Northern Star went from 0.42 to 0.31. The pop song went from 0.22 to 0.08, which is most of the way to mono. The v5.5 version of the same metal prompt went from 0.38 to 0.37.

Two horizontal bar charts, one per track: stereo width and air in the last minute as a percentage of the first minute, coloured by model. The bottom of both charts is V6 tracks, the top is v5 and v5.5 tracks
Every track in the study. Left: stereo width in the last minute as a percentage of the first. Right: the same for air. The bottom of both lists is almost entirely V6; the top is almost entirely v5 and v5.5. Click to enlarge

The high end tells the same story from the other side. Air, the share of energy between 8 and 16 kHz, falls over the length of a V6 track and rises over the length of a v5.5 track. Not stays flat: rises. On the same thrash metal prompt the v5.5 take ends with 187 percent of the air it started with; the V6 take ends with 109 percent, and it started lower. The spectral centre of gravity moves the same way, down on V6, up on the older models. And the bass, which V6 keeps commendably in the centre at the start of a song, goes fully mono by the end, while on v5.5 it stays where it began.

Two line charts over the position in the track from start to end: stereo width and air relative to each track's own first minute, for v5, v5.5 and V6. The V6 line falls below 100 in the second half while v5 and v5.5 hold or rise
Each track measured against its own first minute, then the median of every model generation. The V6 line drops below 100 from about the halfway point and keeps dropping. The older models hold or open up. The shaded bands are the middle half of the tracks. Click to enlarge

Two things make this more than a curiosity. First, it happens with Max Mode on and off alike; the two V6 lines sit on top of each other. Second, the stems show which layer is responsible. The vocal does not narrow at all, on any model. The bass narrows only on V6. The wide elements of the arrangement collapse: on the metal track the rhythm guitars went from a width of 0.90 to 0.53, on the pop song the pads from 1.30 to 0.45. What the vocal loses instead is its top end: on Kings the vocal's air share fell from 4.5 percent to 1.1 over the song, while on v5.5 tracks the vocal got brighter as it went.

Finding 2: the artifacts pile up at the end

Our SpectralForge analysis engine scores a track for thirteen classes of neural-audio defect. We ran it on the first minute and the last minute of every track, loudness matched, and looked at what changed.

On V6 two readings rise sharply toward the end. The shimmer and high-frequency noise detector rose on eleven of fifteen tracks, by fourteen points on average; on the pop song it went from 44 to 100. The tonal artifact bed, the faint layer of sustained tones that neural codecs leave under a mix, rose on ten of fifteen, again by fourteen points; on the metal track from 0 to 47. Pumping rose on nine of fifteen. The overall health score dropped on most tracks.

On v5 and v5.5 the same detectors stay flat or fall. That is the control doing its job: the last minute of an older track has fewer artifacts than its first, and the last minute of a V6 track has more.

Five bar panels comparing the first minute and the last minute per model generation: stereo width, air, shimmer severity, tonal artifact bed severity and lyric intelligibility. V6 gets narrower, darker, noisier and less intelligible; v5.5 holds
First minute (faded) against last minute (solid), averaged per model. Width and air fall on V6 only. Shimmer and the tonal artifact bed rise on V6 only. Lyric intelligibility, from the next section, falls on V6 only. Click to enlarge

Finding 3: in heavy genres the words fall apart

The complaint that comes up most after the plastic-bucket sound is that the singer stops making sense. We tested that with a speech recogniser. Whisper transcribed every isolated vocal stem, the language was detected once per track and locked, and for every ten-second window we recorded how confident the model was about the words it heard. A vocal that a speech model cannot follow is a vocal a listener cannot follow either.

The result splits cleanly by genre, and the control makes it unambiguous because these are the same prompts on two models. On the thrash metal prompt, Whisper's average word confidence on V6 fell from 0.75 in the first minute to 0.54 in the last, and the share of words it was unsure about doubled from 19 to 40 percent. On the v5.5 take of the same prompt confidence went the other way, from 0.83 to 0.86. On the industrial track, V6 fell from 0.85 to 0.65 while v5.5 held at 0.93. On the new metal track it fell from 0.82 to 0.58, with unsure words going from 9 to 34 percent.

Pop and ballads keep their words to the end. The pop song stayed at 0.90, the Finnish ballad got clearer, and so did the pop prompts from the launch article. Kings of the Northern Star, the track you heard at the top, also keeps its words: its ending fails in the accompaniment and the tone, not the lyric. Different tracks come apart at different seams. Metal loses the words, epic rock loses the backing, ballads lose the width and the air.

Bar chart per song of the change in Whisper word confidence from the first minute to the last, v5.5 against V6 on the same prompt. The three heavy tracks drop 20 to 24 points on V6 and hold on v5.5; the pop tracks hold on both
Change in word confidence from the first minute to the last, same prompt on v5.5 and V6. The heavy genres drop 20 to 24 points on V6 and hold on v5.5. Pop holds on both. Click to enlarge

Iron Doctrine: the same last minute on v5.5 and V6

Thrash metal, 168 BPM, same prompt and lyrics. The final thirty seconds before the fade of each take, loudness matched.

Listen for the consonants. On v5.5 the words are still there at the end, and so is the width of the guitars. On V6 the vocal has turned into a texture, and the guitars have moved into the middle. This is the track where the numbers and the ears agree most loudly.

Warning: the V6 clip is the failure, the v5.5 clip is the control

Both are raw Suno exports of the same prompt, the final thirty seconds of each. The V6 side is supposed to sound wrong: that is the finding. The v5.5 side is the older model doing the same job a few months earlier. Neither has been mastered or processed in any way.

Finding 4: a music model hears the drift

Spectral measurements describe what changes. They do not say whether the change is the kind a listener notices. So we asked a model built to understand music. MERT is an open model trained on music at 75 frames per second; its middle and upper layers encode pitch, harmony and timbre rather than genre labels. We embedded every ten-second window of every track and measured how far each window sits from the same track's own first minute.

On the older models the music stays close to where it started and drifts a little toward the final chorus, as any song does. On V6 it moves twice as far: by the last minute a V6 track is, on average, 0.13 units from its own opening against 0.05 for v5.5, and the gap opens from about the halfway point. Ranked per track, the largest drift in the whole study is Kings of the Northern Star with Max Mode off, at 0.30. The metal track sits at 0.23 on V6 and at 0.04 on v5.5. The pop prompts from the launch article sit at 0.04 to 0.08 on every model.

Line chart over the position in the track: distance from the first minute in MERT embedding space, for v5, v5.5 and V6. The V6 line separates from the others at about 40 percent and rises to double their value by the end
The chart from the top of the article, in its proper place. How far the music has moved from its own first minute, as heard by MERT. V6 separates from the older models at about the 40 percent mark and ends twice as far away. Click to enlarge

We then ran a sterner test. We trained a small classifier to tell a first-minute window from a last-minute window using nothing but the MERT features, and tested it on tracks it had never seen. On v5 and v5.5 it managed 60 to 67 percent, barely above a coin toss, which is what you expect when endings differ from openings only in the ordinary ways. On V6 it managed 86 to 89 percent. A V6 ending is systematically a different place in music space from a V6 opening, in a way that generalises from song to song.

Max Mode does not fix it

Five songs, each generated with the same prompt with Max Mode on and off, gave us the comparison directly. In the last minute the Max Mode take is not wider, brighter or cleaner than the standard take in any consistent way. On the metal track it is narrower and darker. On Kings it is a little wider and a little darker. Across every mix and stem measurement we have, no metric moves the same way on more than four of the five songs, which is what you get from noise.

Two bar charts per song comparing Max Mode off and on in the last minute: stereo width and air. No consistent direction across the five songs
The last minute with Max Mode off and on, five songs. No consistent direction on width or air. Click to enlarge

Then we took the ears out of the equation, or rather put them under controlled conditions. Thirteen pairs of thirty-second clips, each pair the same song and the same position with Max Mode on and off, loudness matched, labelled X and Y at random. One listener, who knew the songs well, picked the better one of each pair without knowing which was which. Max Mode won two of the five endings, three of the five openings and all three of the one-minute generations: eight of thirteen overall, which is chance. On Kings, the track that motivated all of this, the blind pick for the ending was the standard take.

There is something Max Mode buys. On the three pairs where the two takes were structurally almost identical, the listener chose Max Mode every time and described the difference as tightness in the low end. That is a real but subtle benefit, and it does not touch the drift. You pay double for a firmer bass, and the last minute still collapses.

A shorter song does not escape it

If the model has a fixed budget per generation, a one-minute song should get the whole budget and sound like the first minute of a long one. We generated six: three new songs in three genres, each with Max Mode on and off, all sixty to eighty seconds long.

They did not sound like first minutes. All six are narrower than the average last minute of a full V6 track: widths of 0.02 to 0.11 against 0.17 for a full track's ending and 0.24 for its opening. The techno track had a stereo width of 0.02, which is mono. Two of the three songs also darkened inside their own sixty seconds: the techno track's 95-percent roll-off fell from 2.8 kHz in its first half to 0.9 kHz in its second. Whatever the drift is tied to, it is not the absolute length of the generation. It looks like the position in the song.

Bar chart of stereo width for six one-minute V6 generations with two reference lines: the average first minute and the average last minute of full V6 tracks. All six bars sit below both lines
Six one-minute generations against the average opening and ending of the full tracks. Every one is narrower than a full track's last minute. Click to enlarge

What it is not

We looked for several other explanations and did not find them, which is worth saying because they are the ones people reach for first.

Why this might be happening

One fact and several guesses, kept apart on purpose.

The fact: in November 2025 Warner Music Group settled its lawsuit with Suno. The public terms were that Suno would phase out its existing models, introduce new models trained on licensed material during 2026, and restrict downloads for paying users to a monthly cap. V6 shipped on 9 September with every older model removed, a week after the download cap arrived. V6 is the licensed generation those terms described.

The guesses. A licensed training set is a different and probably smaller set than whatever v5.5 learned from, and that alone changes the sound; our launch article measured those static changes. It does not by itself explain a drift over the length of a track. A drift that follows the position in the song rather than its absolute length, that appears in a one-minute generation as readily as a four-minute one, and that touches width, air, artifacts and diction at once, looks like a rendering stage whose quality decays as it goes, whatever is feeding it. Max Mode costing exactly double, and buying a subtle firmness rather than a fix, looks like the same engine given more compute. Beyond that we would be inventing.

The one thing we would rule out is the idea that the labels asked for worse audio. Releases are limited by the download cap, which is explicit and visible. A muffled last minute serves nobody, least of all Suno. The likelier reading is a new engine shipped in a hurry with long-context quality unfinished. It is also a testable reading: if the drift shrinks in the next V6 update, it was unfinished. If it does not, it is structural.

What you can do about it today

Method and limits

Twenty-six tracks: five V6 songs generated with Max Mode on and off from identical prompts and settings, one v5.5 version of the metal song, and five prompts generated on v5, v5.5 and V6 for our launch article. Stems by BS-RoFormer. Windows of ten seconds with a five-second hop, measured on the mix and on each stem. Trend per track is the Spearman correlation of a measurement with time; groups are compared with a Mann-Whitney test. First-minute and last-minute clips were loudness matched to -16 LUFS integrated. Artifact severities from SpectralForge's analysis pass. Transcription by Whisper medium with the language fixed per track. Embeddings from MERT-v1-95M, layers 6, 9 and 12; the position classifier is a regularised logistic regression evaluated leave-one-track-out.

The limits are real. One take per track for the older models. For the Max Mode pairs the author generated two to four takes per mode and kept the one closest to the other mode, which is why we do not draw conclusions from how similar the pairs are in structure. The stems come from a separator and carry its fingerprint, on every model equally. Fifteen V6 tracks against eleven older ones is enough for the direction and rough size of the effects reported here; it is not enough to rank genres. The blind test had one listener. Everything else is in the numbers, and the numbers are reproducible.

See how far your track drifted

Step 1. Drop the V6 export into the MasterForge Audio Analyzer for a free health score and artifact breakdown. No account needed.

Step 2. Run Preparation to remove the shimmer and the tonal bed that build up toward the end, then finish it in Pro Master, where the stereo and EQ tools bring the width and the air back.

masterforge.app
Prefer it done for you?

If you would rather not touch a single control, our hand-finished mastering service takes your track and returns a release-ready master, finished by a person, not a preset.

Before You Master · Guide Series

12 · Suno V6 Runs Out of Steam: Why the Last Minute Falls Apart You are here
More guides on the way Coming soon