Suno V6 Runs Out of Steam:
Why the Last Minute Falls Apart
Everyone who has generated a full-length song on Suno V6 has heard it. The first minute is fine, sometimes better than fine. Then the mix starts to shrink. By the last chorus the stereo image has folded into the middle, the top end is gone, and the whole thing sounds like it is playing from a plastic bucket. The words go soft. The ending arrives as if the model no longer knew how to finish a song.
The forums call it running out of tokens. Suno's answer is Max Mode, which costs double credits and is recommended for anything over two minutes. Since 3 September a Pro account also gets twenty downloads a month, so a usable three-minute track now costs two generations out of a budget that was already small. That makes the question worth answering properly: is the drift real, is it V6, and does Max Mode fix it?
We measured it. Twenty-six full tracks, every one split into stems, every one analysed ten seconds at a time from the first bar to the last. Half of them are V6. The other half are v5 and v5.5 tracks generated from the same prompts back when those models still existed, which gives us the one thing a complaint thread never has: a control.
The drift is real and it is V6-specific. Over the length of a track V6 narrows, pushes its bass to mono and darkens, its neural artifacts pile up toward the end, and in heavy genres the lyrics fall apart. v5 and v5.5 tracks from the same prompts hold or even open up. Max Mode does not fix it, by meter or by blind listening, and a one-minute generation does not escape it either.
Hear it first: Kings of the Northern Star, Max Mode off
This is the track that started the whole investigation. Same song, same take, one minute apart in either direction. Listen for the width of the guitars and the air above the vocal, then for what is left of both at the end.
The "Last minute" clip is meant to sound bad. It is the raw, untouched Suno V6 export, and it is what this article is about. Nothing has been done to it and nothing here is trying to sell it. Compare it with the "First minute" clip of the same song: same take, same singer, one minute apart. The collapse in width, air and clarity between the two is the whole story.
What we measured, and against what
Two sets of tracks. The first is five new V6 songs in five genres, each generated twice with the same prompt, lyrics and settings, once with Max Mode on and once with it off: a metal track, a rock track, a Finnish pop ballad, an epic rock track and a pop song. The second is the five prompts from our V6 launch article, which exist as v5, v5.5 and V6 versions of the same song, plus a v5.5 version of the metal track. Twenty-six tracks in total, fifteen of them V6.
Every track went through the same pipeline. It was split into six stems with a neural separator, so the vocal, the drums, the bass and the accompaniment could be measured on their own. The mix and every stem were then analysed in ten-second windows, five seconds apart, from start to finish. For each measurement we asked one question: does it trend over the length of the track, and does the trend differ between V6 and the older models?
That last part is the control, and it matters more than any single number. Every song gets louder and denser toward its final chorus, on any model, because that is how songs are built. A measurement that only says the ending is different from the beginning proves nothing. A measurement that says V6 endings are different in a way v5.5 endings are not is a finding.
All twenty-six tracks were separated, transcribed and embedded on a workstation we call Somnus: a 24-core AMD system with two AMD GPUs, a Radeon AI PRO R9700 with 32 GB and a Radeon RX 7900 XT with 20 GB, running ROCm 7.2. Stem separation with BS-RoFormer took 43 to 90 seconds per track. The same job on the CPU of our production analysis server takes twenty-two minutes, so the GPU is roughly twenty times faster, which is the difference between an afternoon and a fortnight. Whisper transcribed a vocal stem in seven seconds, MERT embedded a whole track in just over one. Consumer AMD hardware is now entirely adequate for this kind of audio science, and nobody had to ask NVIDIA for permission.
If you enjoy this sort of thing, the same machine keeps a blog about AI science and mathematics at sf3d.fi. Warning: extremely dry mathematical content, occasionally moistened by a joke.
Finding 1: the mix narrows and darkens as the song goes on
Start with stereo width, measured as one minus the correlation between the left and right channels, so that zero is mono and one is fully wide. For each track we asked whether width trends downward over time. On V6 it does: eleven of fifteen V6 tracks narrow steadily over their length, against two of eleven older tracks. Averaged over the group, the trend on V6 is clearly negative and on v5 and v5.5 it is zero. The difference is statistically solid for a sample this size.
The size of the effect is easier to grasp per track. On the metal track the width fell from 0.28 in the first minute to 0.16 in the last with Max Mode off, and from 0.25 to 0.12 with it on. Kings of the Northern Star went from 0.42 to 0.31. The pop song went from 0.22 to 0.08, which is most of the way to mono. The v5.5 version of the same metal prompt went from 0.38 to 0.37.
The high end tells the same story from the other side. Air, the share of energy between 8 and 16 kHz, falls over the length of a V6 track and rises over the length of a v5.5 track. Not stays flat: rises. On the same thrash metal prompt the v5.5 take ends with 187 percent of the air it started with; the V6 take ends with 109 percent, and it started lower. The spectral centre of gravity moves the same way, down on V6, up on the older models. And the bass, which V6 keeps commendably in the centre at the start of a song, goes fully mono by the end, while on v5.5 it stays where it began.
Two things make this more than a curiosity. First, it happens with Max Mode on and off alike; the two V6 lines sit on top of each other. Second, the stems show which layer is responsible. The vocal does not narrow at all, on any model. The bass narrows only on V6. The wide elements of the arrangement collapse: on the metal track the rhythm guitars went from a width of 0.90 to 0.53, on the pop song the pads from 1.30 to 0.45. What the vocal loses instead is its top end: on Kings the vocal's air share fell from 4.5 percent to 1.1 over the song, while on v5.5 tracks the vocal got brighter as it went.
Finding 2: the artifacts pile up at the end
Our SpectralForge analysis engine scores a track for thirteen classes of neural-audio defect. We ran it on the first minute and the last minute of every track, loudness matched, and looked at what changed.
On V6 two readings rise sharply toward the end. The shimmer and high-frequency noise detector rose on eleven of fifteen tracks, by fourteen points on average; on the pop song it went from 44 to 100. The tonal artifact bed, the faint layer of sustained tones that neural codecs leave under a mix, rose on ten of fifteen, again by fourteen points; on the metal track from 0 to 47. Pumping rose on nine of fifteen. The overall health score dropped on most tracks.
On v5 and v5.5 the same detectors stay flat or fall. That is the control doing its job: the last minute of an older track has fewer artifacts than its first, and the last minute of a V6 track has more.
Finding 3: in heavy genres the words fall apart
The complaint that comes up most after the plastic-bucket sound is that the singer stops making sense. We tested that with a speech recogniser. Whisper transcribed every isolated vocal stem, the language was detected once per track and locked, and for every ten-second window we recorded how confident the model was about the words it heard. A vocal that a speech model cannot follow is a vocal a listener cannot follow either.
The result splits cleanly by genre, and the control makes it unambiguous because these are the same prompts on two models. On the thrash metal prompt, Whisper's average word confidence on V6 fell from 0.75 in the first minute to 0.54 in the last, and the share of words it was unsure about doubled from 19 to 40 percent. On the v5.5 take of the same prompt confidence went the other way, from 0.83 to 0.86. On the industrial track, V6 fell from 0.85 to 0.65 while v5.5 held at 0.93. On the new metal track it fell from 0.82 to 0.58, with unsure words going from 9 to 34 percent.
Pop and ballads keep their words to the end. The pop song stayed at 0.90, the Finnish ballad got clearer, and so did the pop prompts from the launch article. Kings of the Northern Star, the track you heard at the top, also keeps its words: its ending fails in the accompaniment and the tone, not the lyric. Different tracks come apart at different seams. Metal loses the words, epic rock loses the backing, ballads lose the width and the air.
Iron Doctrine: the same last minute on v5.5 and V6
Listen for the consonants. On v5.5 the words are still there at the end, and so is the width of the guitars. On V6 the vocal has turned into a texture, and the guitars have moved into the middle. This is the track where the numbers and the ears agree most loudly.
Both are raw Suno exports of the same prompt, the final thirty seconds of each. The V6 side is supposed to sound wrong: that is the finding. The v5.5 side is the older model doing the same job a few months earlier. Neither has been mastered or processed in any way.
Finding 4: a music model hears the drift
Spectral measurements describe what changes. They do not say whether the change is the kind a listener notices. So we asked a model built to understand music. MERT is an open model trained on music at 75 frames per second; its middle and upper layers encode pitch, harmony and timbre rather than genre labels. We embedded every ten-second window of every track and measured how far each window sits from the same track's own first minute.
On the older models the music stays close to where it started and drifts a little toward the final chorus, as any song does. On V6 it moves twice as far: by the last minute a V6 track is, on average, 0.13 units from its own opening against 0.05 for v5.5, and the gap opens from about the halfway point. Ranked per track, the largest drift in the whole study is Kings of the Northern Star with Max Mode off, at 0.30. The metal track sits at 0.23 on V6 and at 0.04 on v5.5. The pop prompts from the launch article sit at 0.04 to 0.08 on every model.
We then ran a sterner test. We trained a small classifier to tell a first-minute window from a last-minute window using nothing but the MERT features, and tested it on tracks it had never seen. On v5 and v5.5 it managed 60 to 67 percent, barely above a coin toss, which is what you expect when endings differ from openings only in the ordinary ways. On V6 it managed 86 to 89 percent. A V6 ending is systematically a different place in music space from a V6 opening, in a way that generalises from song to song.
Max Mode does not fix it
Five songs, each generated with the same prompt with Max Mode on and off, gave us the comparison directly. In the last minute the Max Mode take is not wider, brighter or cleaner than the standard take in any consistent way. On the metal track it is narrower and darker. On Kings it is a little wider and a little darker. Across every mix and stem measurement we have, no metric moves the same way on more than four of the five songs, which is what you get from noise.
Then we took the ears out of the equation, or rather put them under controlled conditions. Thirteen pairs of thirty-second clips, each pair the same song and the same position with Max Mode on and off, loudness matched, labelled X and Y at random. One listener, who knew the songs well, picked the better one of each pair without knowing which was which. Max Mode won two of the five endings, three of the five openings and all three of the one-minute generations: eight of thirteen overall, which is chance. On Kings, the track that motivated all of this, the blind pick for the ending was the standard take.
There is something Max Mode buys. On the three pairs where the two takes were structurally almost identical, the listener chose Max Mode every time and described the difference as tightness in the low end. That is a real but subtle benefit, and it does not touch the drift. You pay double for a firmer bass, and the last minute still collapses.
A shorter song does not escape it
If the model has a fixed budget per generation, a one-minute song should get the whole budget and sound like the first minute of a long one. We generated six: three new songs in three genres, each with Max Mode on and off, all sixty to eighty seconds long.
They did not sound like first minutes. All six are narrower than the average last minute of a full V6 track: widths of 0.02 to 0.11 against 0.17 for a full track's ending and 0.24 for its opening. The techno track had a stereo width of 0.02, which is mono. Two of the three songs also darkened inside their own sixty seconds: the techno track's 95-percent roll-off fell from 2.8 kHz in its first half to 0.9 kHz in its second. Whatever the drift is tied to, it is not the absolute length of the generation. It looks like the position in the song.
What it is not
We looked for several other explanations and did not find them, which is worth saying because they are the ones people reach for first.
- It is not less music. Chord changes per second, vocal notes per second, drum hits per second and the number of distinct chords do not trend differently on V6 than on the older models. The ending is not emptier in what it plays. It is emptier in how it sounds.
- It is not a loop. How much a window repeats material from earlier in the song is the same on every model, at every point in the track.
- It is not the vocal tuning. Earlier we suspected V6's metallic vocals came from tighter pitch correction. With the same prompt on both models, v5.5 and V6 vocals are equally quantised, equally coherent, equally reverberant. Those are properties of the song, not the model.
- It is not one bad take. Eleven of fifteen V6 tracks narrow, ten of fifteen gain artifacts, and the classifier that recognises a V6 ending was tested on songs it had not seen.
Why this might be happening
One fact and several guesses, kept apart on purpose.
The fact: in November 2025 Warner Music Group settled its lawsuit with Suno. The public terms were that Suno would phase out its existing models, introduce new models trained on licensed material during 2026, and restrict downloads for paying users to a monthly cap. V6 shipped on 9 September with every older model removed, a week after the download cap arrived. V6 is the licensed generation those terms described.
The guesses. A licensed training set is a different and probably smaller set than whatever v5.5 learned from, and that alone changes the sound; our launch article measured those static changes. It does not by itself explain a drift over the length of a track. A drift that follows the position in the song rather than its absolute length, that appears in a one-minute generation as readily as a four-minute one, and that touches width, air, artifacts and diction at once, looks like a rendering stage whose quality decays as it goes, whatever is feeding it. Max Mode costing exactly double, and buying a subtle firmness rather than a fix, looks like the same engine given more compute. Beyond that we would be inventing.
The one thing we would rule out is the idea that the labels asked for worse audio. Releases are limited by the download cap, which is explicit and visible. A muffled last minute serves nobody, least of all Suno. The likelier reading is a new engine shipped in a hurry with long-context quality unfinished. It is also a testable reading: if the drift shrinks in the next V6 update, it was unfinished. If it does not, it is structural.
What you can do about it today
- Write the ending short. The collapse follows the position in the song, so a track that ends at 2:30 keeps more of its opening than one that ends at 3:30. Put the last chorus earlier and let the outro be brief.
- Do not pay double expecting a fix. Max Mode buys a firmer low end. It does not buy a wider or cleaner ending. If credits are tight, spend them on more takes and pick the one that holds together longest.
- Mastering recovers some of it, not all of it. Width and air can be pushed back with a mid-side widener and a high shelf applied more strongly toward the end of the track; the Pro Master stereo and EQ tools do that. What no mastering chain restores is a vocal that has turned into a texture. That is the part to check in the raw generation before you spend anything on it.
- Clean the artifacts before you master. The shimmer and the tonal bed that build up in the last minute are exactly what Preparation is built to remove. Run it first, then master, and the ending gets closer to the opening than either step alone.
Method and limits
Twenty-six tracks: five V6 songs generated with Max Mode on and off from identical prompts and settings, one v5.5 version of the metal song, and five prompts generated on v5, v5.5 and V6 for our launch article. Stems by BS-RoFormer. Windows of ten seconds with a five-second hop, measured on the mix and on each stem. Trend per track is the Spearman correlation of a measurement with time; groups are compared with a Mann-Whitney test. First-minute and last-minute clips were loudness matched to -16 LUFS integrated. Artifact severities from SpectralForge's analysis pass. Transcription by Whisper medium with the language fixed per track. Embeddings from MERT-v1-95M, layers 6, 9 and 12; the position classifier is a regularised logistic regression evaluated leave-one-track-out.
The limits are real. One take per track for the older models. For the Max Mode pairs the author generated two to four takes per mode and kept the one closest to the other mode, which is why we do not draw conclusions from how similar the pairs are in structure. The stems come from a separator and carry its fingerprint, on every model equally. Fifteen V6 tracks against eleven older ones is enough for the direction and rough size of the effects reported here; it is not enough to rank genres. The blind test had one listener. Everything else is in the numbers, and the numbers are reproducible.
See how far your track drifted
Step 1. Drop the V6 export into the MasterForge Audio Analyzer for a free health score and artifact breakdown. No account needed.
Step 2. Run Preparation to remove the shimmer and the tonal bed that build up toward the end, then finish it in Pro Master, where the stereo and EQ tools bring the width and the air back.
If you would rather not touch a single control, our hand-finished mastering service takes your track and returns a release-ready master, finished by a person, not a preset.