PT TechnologiesFor People

/research · R-001

Data · CSVCode · measure.py

[PT-R001][REV1.0][N780][SRCAIME]

PT Technologies

Spectral profiling at scale

Cross-dataset synthesis of high-frequency roll-off and phase disruption in commercial AI music

Band edgeset by the decoding representation and the filenot by the paradigm

Commercial files14.5 kHz to 20.0 kHz, medianoften a steep edge at one frequency

Human referenceMP3 at 320 kbpsmedian 18.2 kHz

StereoSuno and Udio more correlated than human musicno phase disruption

85 % roll-off0.57 kHz to 1.72 kHz for every sourcenot a band-edge measure

Abstract

Two hypotheses about the audio of text-to-music systems follow from how they are built: that autoregressive systems, decoding through neural audio codecs, lose the top of the spectrum near 16–18 kHz while latent diffusion systems keep the full band, and that systems which quantise their channels separately deliver a disrupted stereo image. We test both against the public record. We collate the documented decoding path of twelve generative systems from their technical reports and documentation, and we measure 780 tracks from the AIME dataset (Grötschla et al., 2025) - 60 from each of twelve generators, among them Suno v3 and v3.5, Udio and Stable Audio 1 and 2, and 60 human-made tracks from MTG-Jamendo - for spectral roll-off, the bandwidth of the long-term spectrum and the steepness of its upper edge, spectral flatness, and inter-channel correlation, with a pipeline validated on synthetic signals and published with the paper.

Bandwidth ended where the decoding representation or the file ended, whatever the paradigm: at 8.0 kHz for AudioLDM 2 and Mustango, 10.0 kHz for Riffusion, and near 15.4 kHz for MusicGen. The commercial systems' files ended between 14.5 kHz and 20.0 kHz (medians), most often at one steep edge per system, a shape shared with lossy encoding; the human reference, itself distributed as 320 kbps MP3, has the same kind of edge near 20 kHz. No commercial system showed phase disruption: Suno v3.5 (median ρ = 0.93), Suno v3 (0.84) and Udio (0.83) were more strongly correlated than the human reference (0.55) and lost less level when summed to mono. The 85 % spectral roll-off of every source lay between 0.57 kHz and 1.72 kHz, far from the band edge it is often quoted for. Audits of generated audio should separate the decoder from the delivery codec before attributing a band limit to a model.

Keywords AI music · spectral roll-off · bandwidth · stereo correlation · neural audio codecs · latent diffusion

1 Introduction

Text-to-music generation moved, between 2023 and 2025, from research systems that rendered short mono clips to commercial services that deliver full songs in stereo. Two architectural families account for most of it. Autoregressive models predict sequences of discrete tokens produced by a neural audio codec and decode them back to a waveform (Agostinelli et al., 2023; Copet et al., 2023; Yuan et al., 2025). Latent diffusion models denoise a continuous latent representation and decode it through an autoencoder (Evans, Carr, et al., 2024; H. Liu, Yuan, et al., 2024; Melechovsky et al., 2024). Two widely used consumer services, Suno and Udio, have not published their architectures.

Two hypotheses about the audio these systems deliver follow from how they are built. The first is that the paradigm fixes the bandwidth: a codec with a limited codebook discards the top of the spectrum, so autoregressive systems stop at 16–18 kHz, while diffusion systems decoded at 44.1 kHz keep the full band to 22.05 kHz. The second is that systems which code each channel separately lose the coherence between them, and so deliver a smeared, phase-unstable stereo image that collapses when summed to mono - a design that exists, in the stereo variants of MusicGen, which encode left and right separately and interleave their codebooks (Meta AI, n.d.).

This paper tests both with measurements on public data, and asks three questions.

  1. RQ1. What sets the upper edge of a generated track's spectrum: the paradigm, the representation the audio is decoded from, or the file it is delivered in?
  2. RQ2. How do the stereo images of commercial systems compare with human-produced music in inter-channel correlation and mono compatibility?
  3. RQ3. Which signal-level measures can rank generated audio, and which cannot?

It contributes a collation of the documented decoding path of twelve systems (Table 1), measurements of 780 tracks with a pipeline that is validated on signals with known answers and published in full, and the corrections to the common account that the evidence supports.

2 Background

2.1 Decoders and what they can hold

A sampled signal can represent nothing above half its sampling rate (Shannon, 1949). A generative system therefore has a ceiling set by the rate of whatever it decodes from, and that rate is documented for most open systems. MusicLM generates 24 kHz audio through SoundStream tokens (Agostinelli et al., 2023; Zeghidour et al., 2022). MusicGen uses “a 32kHz EnCodec tokenizer with 4 codebooks sampled at 50 Hz” (Meta AI, n.d.; Copet et al., 2023; Défossez et al., 2023). YuE tokenises with X-Codec at 16 kHz and 50 Hz, then upsamples the decoded 16 kHz audio to 44.1 kHz with a lightweight Vocos-based vocoder (Yuan et al., 2025). The Stable Audio models are latent diffusion models that generate stereo at 44.1 kHz (Evans, Carr, et al., 2024; Evans, Parker, et al., 2024, 2025); AudioLDM 2 and Mustango decode at 16 kHz (H. Liu, Yuan, et al., 2024; Melechovsky et al., 2024). A codec designed for full-band audio exists - the Descript Audio Codec compresses 44.1 kHz audio to 8 kbps (Kumar et al., 2023) - so a codec as such does not imply a low ceiling.

A file's rate, though, says only what it could hold, not what it holds. YuE's 44.1 kHz output carries, above 8 kHz, whatever its vocoder synthesises. Riffusion renders spectrogram images at 44.1 kHz, but its spectrograms stop at a maximum frequency of 10 kHz (Forsgren & Martiros, 2022).

Table 1The documented decoding path of twelve text-to-music systems
SystemParadigmDecoded fromOutputCeilingSource
MusicLMAutoregressiveSoundStream tokens24 kHz12 kHzAgostinelli et al. (2023)
MusicGenAutoregressiveEnCodec, 4 codebooks at 50 Hz32 kHz mono or stereo16 kHzCopet et al. (2023); Meta AI (n.d.)
YuEAutoregressiveX-Codec at 16 kHz, 50 Hz44.1 kHz via vocoder8 kHz decodedYuan et al. (2025)
Stable Audio 1Latent diffusionAutoencoder latent44.1 kHz stereo22.05 kHzEvans, Carr, et al. (2024)
Stable Audio 2Latent diffusionLatent at 21.5 Hz44.1 kHz stereo22.05 kHzEvans, Parker, et al. (2024)
Stable Audio OpenLatent diffusionAutoencoder latent44.1 kHz stereo22.05 kHzEvans, Parker, et al. (2025)
AudioLDM 2Latent diffusionMel-spectrogram latent16 kHz8 kHzH. Liu, Yuan, et al. (2024)
MustangoLatent diffusionMel-spectrogram latent16 kHz8 kHzMelechovsky et al. (2024)
RiffusionImage diffusionSpectrogram to 10 kHz44.1 kHz10 kHzForsgren & Martiros (2022)
Lyria 2Not publishedNot published48 kHz WAV24 kHzGoogle Cloud (n.d.)
SunoNot publishedNot publishedNot documented--
UdioNot publishedNot publishedNot documented--

Note. “Ceiling” is half the rate of the stage named, the most the output can hold from that stage. A ceiling says nothing about content below it. Lyria's 48 kHz is the container its API returns; its architecture is unpublished.

2.2 Artifacts the literature has already found

Upsampling layers in neural decoders leave tonal and filtering artifacts and spectral replicas (Pons et al., 2021), the audio counterpart of the checkerboard patterns of transposed convolution in images (Odena et al., 2016). Afchar, Meseguer-Brocal, Akesbi, and Hennequin (2025) prove that such deconvolution modules imprint small, periodic spectral peaks determined by the architecture rather than the weights, show them in open systems and in Suno and Udio, and detect generated music from them at above 99 % accuracy in several scenarios. Large corpora of commercial output now exist - SONICS holds over 49,000 synthetic songs from Suno and Udio (Rahman et al., 2025) - but detection built on such traces is fragile under manipulation and on unseen generators (Afchar, Meseguer-Brocal, & Hennequin, 2025; Oh, 2026a), and benchmarks that compare detectors are themselves hard to keep consistent: an audit of ArtifactBench found three incompatible manifest lineages and audio that could no longer be fully reconstructed (Oh, 2026b).

2.3 The delivery codec as a confound

Lossy encoders remove the top of the spectrum by design, and the result looks like a model's band limit. The detector of Oh (2026a) initially labelled real music as generated after MP3 encoding, at a false-positive rate of 98.7 %, and had to be retrained across codecs. Reference collections are not exempt: MTG-Jamendo, a common source of human-made music in this field, distributes all its audio as 320 kbps MP3 (Bogdanov et al., 2019).

2.4 Measures

Spectral roll-off, the frequency below which a set share of a frame's energy lies, was introduced as a feature for telling speech from music (Scheirer & Slaney, 1997); it describes where the energy is, not where the spectrum ends. Spectral flatness, the ratio of the geometric to the arithmetic mean of the power spectrum, is 1 for white noise and falls towards 0 for a few strong partials (Gray & Markel, 1974). Inter-channel correlation measures how alike two channels are, and is a basic cue of perceived width (Faller & Baumgarte, 2003). The Fréchet Audio Distance compares the distribution of embeddings of generated audio with that of a reference set (Kilgour et al., 2019); its value depends on the embedding model and the reference, so scores computed with different embeddings or reference sets cannot be compared (Gui et al., 2024). On AIME's own human ratings, FAD computed with a music-trained CLAP embedding tracked human judgements of quality best of the metrics tested (Grötschla et al., 2025).

3 Method

3.1 Data

AIME holds 6,000 tracks from twelve generators current in July 2024, and 500 human-made tracks from MTG-Jamendo, released for research with 15,600 human pairwise judgements (Grötschla et al., 2025); the generated audio is licensed CC BY 4.0. We drew 60 tracks per source - the first 60 in dataset order, taking each source's own shards first and filtering shared shards by source - for 780 tracks in all. No audio was kept after measurement.

Table 2 lists what the files are. Every stereo source is stored as 48 kHz WAV; MusicGen as 32 kHz mono; Riffusion as 44.1 kHz mono; AudioLDM 2 and Mustango as 16 kHz mono. AIME does not say how the commercial systems' audio was obtained or whether it passed through a lossy format on the way; the measurements below are of the files as published.

Table 2The sample, by source, as stored in AIME
SourceKindTracksStored asMedian length
MTG-JamendoHuman (MP3 320 kbps)6048 kHz stereo165.9 s
Suno v3Commercial6048 kHz stereo120.0 s
Suno v3.5Commercial6048 kHz stereo204.3 s
UdioCommercial6048 kHz stereo32.8 s
Stable Audio v1Commercial6048 kHz stereo10.0 s
Stable Audio v2Commercial6048 kHz stereo10.0 s
MusicGen SmallOpen6032 kHz mono10.2 s
MusicGen MediumOpen6032 kHz mono10.2 s
MusicGen LargeOpen6032 kHz mono10.2 s
RiffusionOpen6044.1 kHz mono10.2 s
AudioLDM 2 LargeOpen6016 kHz mono10.0 s
AudioLDM 2 MusicOpen6016 kHz mono10.0 s
MustangoOpen6016 kHz mono10.2 s

3.2 Processing

Each file was decoded at its own sample rate and channel count. From each, the 10-second window with the highest RMS level was taken - the rule AIME used for its listening clips. No loudness normalisation was applied, because every measure used is unchanged by gain. Nothing was resampled: all results are in absolute frequency, and a resampling filter short enough to be practical leaves images a few tens of decibels down, which a bandwidth measure reads as content the source never had. Short-time spectra used a Hann window of 2,048 points at 32, 44.1 and 48 kHz and 1,024 at 16 kHz (43–64 ms), with a hop of a quarter window; frames more than 60 dB below the loudest were dropped.

3.3 Measures

With Xt[k] the spectrum of frame t at bin k, fk the bin frequency and K the number of bins up to the Nyquist frequency:

froll,γ(t)=min⁡j{fj  :  ∑k=0j∣Xt[k]∣2  ≥  γ∑k=0K∣Xt[k]∣2}f_{\mathrm{roll},\gamma}(t) = \min_{j} \left\{ f_j \;:\; \sum_{k=0}^{j} \lvert X_t[k] \rvert^2 \;\ge\; \gamma \sum_{k=0}^{K} \lvert X_t[k] \rvert^2 \right\}
(1)

with γ = 0.85 and γ = 0.99, averaged over frames. The long-term average spectrum L[k], the mean of |Xt[k]|² over frames in decibels and smoothed over about 100 Hz, gives the bandwidth

B60=max⁡k{fk  :  L[k]>max⁡k′L[k′]−60 dB}B_{60} = \max_{k} \left\{ f_k \;:\; L[k] > \max_{k'} L[k'] - 60\ \mathrm{dB} \right\}
(2)

with B80 defined the same way. Where B60 is not simply the file's Nyquist frequency, the steepness of the edge is the mean level 250–750 Hz below it minus the mean level 250–750 Hz above it, in dB per kHz; an edge of 40 dB per kHz or more is called steep. The share of energy above 16 kHz is the ratio of L summed above 16 kHz to L summed over all bins. Spectral flatness is

SF(t)=exp⁡ ⁣(1N∑kln⁡∣Xt[k]∣2)1N∑k∣Xt[k]∣2\mathrm{SF}(t) = \frac{\exp\!\left( \dfrac{1}{N} \displaystyle\sum_{k} \ln \lvert X_t[k] \rvert^2 \right)}{\dfrac{1}{N} \displaystyle\sum_{k} \lvert X_t[k] \rvert^2}
(3)

over the N bins from 50 Hz to 7.5 kHz, a band every source holds, per channel and then averaged. For the two channels xL and xR of a window of M samples, the zero-lag correlation is

ρ=∑nxL[n] xR[n]∑nxL2[n] ∑nxR2[n]\rho = \frac{\displaystyle\sum_{n} x_L[n]\, x_R[n]}{\sqrt{\displaystyle\sum_{n} x_L^2[n] \, \sum_{n} x_R^2[n]}}
(4)

computed over the whole window and over frames of 2,048 samples (43 ms at 48 kHz), and the change in level when the two are summed to mono is

ΔM=10log⁡10mean⁡ ⁣[(xL[n]+xR[n]2) ⁣2]12(mean⁡ ⁣[xL2[n]]+mean⁡ ⁣[xR2[n]]) dB\Delta M = 10 \log_{10} \frac{\operatorname{mean}\!\left[ \left( \dfrac{x_L[n] + x_R[n]}{2} \right)^{\!2} \right]}{\dfrac{1}{2} \left( \operatorname{mean}\!\left[ x_L^2[n] \right] + \operatorname{mean}\!\left[ x_R^2[n] \right] \right)} \ \mathrm{dB}
(5)

which is 0 dB for identical channels, −3 dB for unrelated ones and falls without limit as they approach inversion. A file whose two channels were identical was treated as mono.

3.4 Validation

Before any real audio was measured, the pipeline was run on signals with known answers.

  • White noise at 16, 24, 32, 44.1, 48 kHzbandwidth 8.0, 12.0, 16.0, 22.05, 24.0 kHzflatness 0.56
  • Identical channelstreated as monopassed
  • Unrelated channelsρ = 0.00, ΔM = −3.0 dBpassed
  • Right = −leftρ = −1.00, ΔM → −∞passed
  • LAME 3.100 MP3, 128 · 192 · 320 kbpsedge at 17.0 · 19.0 · 20.3 kHz72–78 dB/kHz

Two faults were found and fixed this way: a smoothing step that biased the top bins upwards, and a resampling step whose images read as bandwidth; the second is why nothing is resampled.

3.5 Statistics

Measures are reported as medians with interquartile ranges, since their distributions are skewed. Each generator was compared with the human reference by a two-sided Mann-Whitney U test, with the rank-biserial correlation r as effect size (positive when the generator's values are higher), and p-values Holm-corrected within each measure. α = .05. The human reference is itself lossy; section 5.4 returns to what that allows.

4 Results

4.1 Where the spectrum ends

No source in the sample filled a full-band file (Figure 1, Table 3). The open systems ended at their decoding ceiling or just under it: AudioLDM 2 and Mustango at the 8.0 kHz Nyquist frequency of their 16 kHz files; MusicGen at a median of 15.3 kHz–15.4 kHz, under the 16 kHz of its 32 kHz codec; and Riffusion at 10.01 kHz (interquartile range 9.99 kHz–10.03 kHz), although its files run at 44.1 kHz - the maximum frequency of its spectrograms, not the rate of its files.

The commercial systems' files ended higher and more tightly: Stable Audio 1 at a median of 16.6 kHz (16.2 kHz–16.6 kHz), Suno v3 at 17.7 kHz, Suno v3.5 at 18.5 kHz, and Udio at 20.0 kHz. Stable Audio 2 ended lower and more variably (14.5 kHz; 10.3 kHz–16.3 kHz), as a gradual roll-off of its content rather than one edge. The human reference reached a median of 18.2 kHz (13.7 kHz–20.0 kHz).

04812162024bandwidth at −60 dB (kHz)MTG-JamendoSuno v3Suno v3.5UdioStable Audio v1Stable Audio v2MusicGen SmallMusicGen MediumMusicGen LargeRiffusionAudioLDM 2 LargeAudioLDM 2 MusicMustangohuman referencecommercial systemopen systemmiddle 50%medianNyquist of the file
drag to turn · ctrl + wheel to zoom · point at a mark to read it
Figure 1. Bandwidth of every track at −60 dB, by source In three dimensions each source has a lane, and a bar's height is the number of its tracks that end in that half-kilohertz; the flat view, which is also what prints, shows every track as a dot.

4.2 The shape of the edge

Where the edge falls matters less than its shape. In 46 of 60 Stable Audio 1 tracks the spectrum ended in a steep edge, and those edges stood at one frequency - a median of 16.59 kHz. The same held for Suno v3 (26 tracks, at 18.84 kHz), Suno v3.5 (25, at 18.75 kHz) and Udio (34, at 20.02 kHz). Within each system these edges stood in a narrow band - the middle half of Stable Audio 1's spanned 94 Hz, Suno v3.5's 93 Hz and Udio's 64 Hz - across tracks of different genres. Music does not end at one frequency; a filter does. The human reference, known to be 320 kbps MP3, shows the same kind of edge in 15 tracks, at 20.04 kHz. Figure 3 sets these edges beside those LAME leaves on noise at 128, 192 and 320 kbps, which fall a few hundred hertz above each cluster. Steepness itself depends on the signal - music carries less energy just below its edge than noise does, so its edges measure less steep - and can flag a filter but not name an encoder.

A steep edge at one frequency is what a lossy encoder leaves, but also what a fixed low-pass stage inside a system would leave, so the data cannot say which produced these. They can say that the edges are not the Nyquist frequency of the files - none of the stereo files is Nyquist-limited - and not a property of the architectures as documented: Stable Audio decodes at 44.1 kHz (Evans, Carr, et al., 2024), yet its files end at 16.6 kHz.

81012141618202224020406080100steep edge ≥ 40 dB/kHzfrequency of the −60 dB edge (kHz)edge steepness (dB per kHz)MP3 128kMP3 192kMP3 320k
drag to turn · ctrl + wheel to zoom · point at a mark to read it
Figure 3. Steepness of each track's upper edge against its frequency, with the edges MP3 encoding leaves Tracks whose edge is the file's own Nyquist frequency are left out; they have no edge to measure. The MP3 marks were measured on white noise, which measures steeper than music: compare their frequencies, not their heights. In three dimensions each source has a lane of its own.

4.3 The 85 % roll-off

The 85 % roll-off of every source lay between 0.57 kHz and 1.72 kHz (Table 3): most of a full mix's energy sits in the bass and low mids, so the frequency below which 85 % of it lies is low, whatever happens at the top. Even the 99 % roll-off stayed under 6 kHz for every source. The measure locates the body of the spectrum and cannot detect a cut at 16 kHz; values of 14–19 kHz reported as “85 % roll-off” for full mixes describe something else.

Table 3Spectral measures by source: medians, with interquartile ranges
SourceRoll-off 85 %Roll-off 99 %Bandwidth B60Steep edges (at)> 16 kHzFlatness
MTG-Jamendo0.80 kHz2.43 kHz18.2 kHz (13.7–20.0)15/60 (20.0 kHz)-50.0 dB0.006
Suno v30.57 kHz1.53 kHz17.7 kHz (6.7–18.8)26/60 (18.8 kHz)-54.1 dB0.003
Suno v3.50.63 kHz2.83 kHz18.5 kHz (15.5–18.7)25/60 (18.8 kHz)-47.2 dB0.020
Udio1.58 kHz4.59 kHz20.0 kHz (18.1–20.0)34/60 (20.0 kHz)-42.6 dB0.038
Stable Audio v10.92 kHz4.00 kHz16.6 kHz (16.2–16.6)46/60 (16.6 kHz)-49.5 dB0.027
Stable Audio v20.98 kHz2.29 kHz14.5 kHz (10.3–16.3)12/60 (16.6 kHz)-72.5 dB0.004
MusicGen Small1.72 kHz5.55 kHz15.3 kHz (14.4–15.7)0/60-90.9 dB0.048
MusicGen Medium1.13 kHz3.97 kHz15.4 kHz (14.9–15.7)0/60-91.0 dB0.016
MusicGen Large0.85 kHz4.06 kHz15.3 kHz (12.7–15.6)0/60-94.6 dB0.025
Riffusion0.98 kHz3.75 kHz10.0 kHz (10.0–10.0)7/60 (10.0 kHz)-66.6 dB0.011
AudioLDM 2 Large0.99 kHz3.41 kHz8.0 kHz (8.0–8.0)0/60none0.033
AudioLDM 2 Music0.74 kHz2.64 kHz8.0 kHz (8.0–8.0)0/60none0.019
Mustango0.79 kHz3.04 kHz8.0 kHz (8.0–8.0)0/60none0.033

Note. Bandwidth at −60 dB relative to the spectrum's peak; IQR in kHz. Steep edges: 40 dB per kHz or more, with their median frequency. “> 16 kHz” is the share of energy above 16 kHz; “none” where the file cannot hold any. Flatness over 50 Hz–7.5 kHz.

4.4 Flatness

Within the band every source holds, several generators were significantly flatter - more noise-like - than the human reference: Udio (median 0.038 against 0.006; p < .001, r = 0.57), Stable Audio 1 (p < .001, r = 0.47), Suno v3.5 (p = .008, r = 0.34), MusicGen Small and Large, AudioLDM 2 Large and Mustango. Suno v3 and Stable Audio 2 did not differ from it. Flatness also rises with genre and arrangement - a snare-heavy mix is flatter than a solo piano - and AIME's prompts differ in genre between sources, so these differences describe the samples; they are not a ranking of quality.

4.5 Stereo

Six sources are stereo in AIME: the human reference, Suno v3 and v3.5, Udio, and Stable Audio 1 and 2 (Figure 2, Table 4). None of the commercial systems showed the disruption the second hypothesis predicts. Suno v3.5 was the most strongly correlated source in the sample (median ρ = 0.93; 0.84–0.98), well above the human reference (ρ = 0.55; p < .001, r = 0.61); so were Udio (ρ = 0.83; p < .001) and Suno v3 (ρ = 0.84; p = .005). Their channels moved in opposite directions in a median of 0 % of frames, against 3.1 % for the human reference, and summing them to mono cost less: -0.17 dB for Suno v3.5 against -1.11 dB. The Stable Audio models did not differ from the human reference (ρ = 0.75 and 0.78; p = .142 and p = .077), though Stable Audio 2 held the most tracks with a negative overall correlation (7 of 60, against 4 for the human reference).

MTG-Jamendo · human

60 tracks · 48 kHz stereoMedian inter-channel correlationρ 0.55

Suno v3.5

60 tracks · 48 kHz stereoMedian inter-channel correlationρ 0.93

Suno v3

60 tracks · 48 kHz stereoMedian inter-channel correlationρ 0.84

Udio

60 tracks · 48 kHz stereoMedian inter-channel correlationρ 0.83

Stable Audio v1

60 tracks · 48 kHz stereoMedian inter-channel correlationρ 0.75

Stable Audio v2

60 tracks · 48 kHz stereoMedian inter-channel correlationρ 0.78

-1.0-0.50.00.51.0zero-lag inter-channel correlation ρ (−1 phase-inverted · 0 unrelated · +1 mono)mono sum, medianMTG-Jamendo-1.11 dBSuno v3-0.36 dBSuno v3.5-0.17 dBUdio-0.39 dBStable Audio v1-0.59 dBStable Audio v2-0.51 dB
drag to turn · ctrl + wheel to zoom · point at a mark to read it
Figure 2. Zero-lag correlation between the channels of every stereo track, and the median level change on summing to mono The shaded half is out of phase. In three dimensions each track stands as high as the level it loses summed to mono; in the flat view each source's median loss is printed at the right. The hatched box is the middle half of ρ; the heavy bar, its median.
Table 4Stereo measures by source, against the human reference
Sourceρ median (IQR)Tracks ρ < 0Frames ρ < 0Mono sum ΔMp (Holm)r
MTG-Jamendo0.55 (0.23–0.83)4/603.1 %-1.11 dBreference-
Suno v30.84 (0.49–0.95)2/600.0 %-0.36 dB= .0050.33
Suno v3.50.93 (0.84–0.98)0/600.0 %-0.17 dB< .0010.61
Udio0.83 (0.73–0.92)0/600.0 %-0.39 dB< .0010.40
Stable Audio v10.75 (0.45–0.83)1/601.8 %-0.59 dB= .1420.16
Stable Audio v20.78 (0.38–0.95)7/602.0 %-0.51 dB= .0770.22

Note. Mann-Whitney U against MTG-Jamendo on ρ, Holm-corrected over the five comparisons; r is the rank-biserial correlation, positive when the source's ρ is higher.

5 Discussion

5.1 Paradigm does not set the band edge

The first hypothesis does not survive the comparison. An autoregressive system (MusicGen) ended at 15.4 kHz, and a diffusion system (AudioLDM 2) at 8.0 kHz; a diffusion system decoded at 44.1 kHz (Riffusion) ended at 10.0 kHz; and diffusion systems documented at 44.1 kHz (Stable Audio) delivered files that end at 16.6 kHz (Stable Audio 1) or lower. In every open system the edge is the ceiling of the representation the audio was decoded from - the codec's rate, the diffusion model's mel rate or the spectrogram's range (Table 1) - and not the paradigm. A codec's number of codebooks governs how much detail survives within its band (Défossez et al., 2023; Kumar et al., 2023), but the band itself is set by the rate the codec runs at, and codecs exist at 44.1 kHz.

5.2 The delivery codec comes before the model

The commercial systems' files carry steep edges at one frequency per system, of the kind a lossy encoder leaves, and the human reference - known to be MP3 - carries the same kind. Until the provenance of such files is known, a band limit measured in them cannot be attributed to the model. Three practices follow. Studies should obtain lossless exports where a service offers them and state the format they measured. They should report the steepness of the upper edge as well as its frequency, since a steep edge at one frequency points to a filter. And human references should be lossless, or matched in codec and bitrate to the material they are compared with.

5.3 The stereo image is narrow, not broken

The second hypothesis is not supported for the commercial systems measured. Suno and Udio delivered images more strongly correlated than the human reference, with less level lost on a mono sum; in production terms they are narrow and mono-safe, not phase-disrupted. The design the hypothesis rests on - channels coded separately - is documented for MusicGen's stereo variants (Meta AI, n.d.), which AIME does not contain; it remains a fair question for those models, not a finding about commercial music generation.

5.4 In post-production

For an engineer receiving generated audio, the evidence suggests an order of work. First, look at the top of the long-term spectrum: a steep edge at 16–20 kHz marks a band-limited file, and no equaliser will bring back what is not there. Content above the edge can only be synthesised - by an exciter that derives harmonics from the band below, or by neural bandwidth extension such as AudioSR, which extends inputs band-limited to between 2 and 16 kHz up to 24 kHz at a 48 kHz rate (H. Liu, Chen, et al., 2024) - and either adds plausible content, not the original. Second, read the correlation before touching the stereo image: the measured images call for widening, if anything, rather than repair, and the usual mastering check of mono compatibility applies as for any mix (Katz, 2015). Third, normalise loudness to the target of the destination, measured to ITU-R BS.1770-5: −23 LUFS for broadcast under EBU R 128, and the platform's own reference for streaming - Spotify, for one, normalises to −14 LUFS (European Broadcasting Union, 2023; International Telecommunication Union, 2023; Spotify, n.d.).

5.5 Limitations

  • AIME was made in July 2024. Suno, Udio and Stable Audio have released later versions since, and Lyria and YuE are not in it; none of these results describes them.
  • AIME does not state how the commercial audio was obtained, so a lossy stage cannot be confirmed or ruled out.
  • The human reference is 320 kbps MP3; a lossless human reference would sharpen every comparison of bandwidth, and MP3's joint-stereo coding may also have changed how its channels relate.
  • Sixty tracks per source, taken in dataset order, from 10-second windows; the sources also differ in genre and length, which move flatness and width.
  • The measures are signal-level. They describe the files; they do not measure what listeners hear or prefer.
  • FAD was not recomputed; this paper relies on the published evaluation of AIME for it (Grötschla et al., 2025).

6 Conclusion

Measured on public data, the common account of AI music's acoustic faults needs two corrections. The upper edge of a generated track is set by the representation it is decoded from and by the file it is delivered in, not by whether the model is autoregressive or diffusion; and the commercial systems measured deliver narrow, strongly correlated stereo, not disrupted phase. Many of their files end in steep edges between 16.6 and 20.0 kHz, each system's at one frequency, of the kind a lossy encoder or a fixed filter leaves. Deciding whether those edges belong to the models or to their delivery needs lossless exports and a lossless human reference, and the pipeline published with this paper is ready to measure them.

A Data and code

Every number in this paper comes from aime-measurements.csv (780 rows: one per track, identified by its AIME id, with every measure) or from the sources cited. The measures are computed by measure.py (Python, NumPy 2.4.6, SciPy 1.17.1). The audio itself is AIME's and is not redistributed here.

Cite this paper (APA 7)

PT Technologies Research. (2026). Spectral Profiling at Scale: Cross-Dataset Synthesis of High-Frequency Roll-Off and Phase Disruption in Commercial AI Music (Research Paper No. R-001, Version 1.0). PT Technologies. https://techatpt.com/research/spectral-profiling-ai-music

R References

  1. Afchar, D., Meseguer-Brocal, G., Akesbi, K., & Hennequin, R. (2025). A Fourier explanation of AI-music artifacts [Paper presentation]. 26th International Society for Music Information Retrieval Conference (ISMIR 2025). https://doi.org/10.48550/arXiv.2506.19108
  2. Afchar, D., Meseguer-Brocal, G., & Hennequin, R. (2025). AI-generated music detection and its challenges. In ICASSP 2025 - 2025 IEEE International Conference on Acoustics, Speech and Signal Processing (pp. 1–5). IEEE. Preprint: https://doi.org/10.48550/arXiv.2405.04181
  3. Agostinelli, A., Denk, T. I., Borsos, Z., Engel, J., Verzetti, M., Caillon, A., Huang, Q., Jansen, A., Roberts, A., Tagliasacchi, M., Sharifi, M., Zeghidour, N., & Frank, C. (2023). MusicLM: Generating music from text [Preprint]. arXiv. https://doi.org/10.48550/arXiv.2301.11325
  4. Bogdanov, D., Won, M., Tovstogan, P., Porter, A., & Serra, X. (2019). The MTG-Jamendo dataset for automatic music tagging [Paper presentation]. Machine Learning for Music Discovery Workshop, International Conference on Machine Learning (ICML 2019), Long Beach, CA, United States. https://github.com/MTG/mtg-jamendo-dataset
  5. Copet, J., Kreuk, F., Gat, I., Remez, T., Kant, D., Synnaeve, G., Adi, Y., & Défossez, A. (2023). Simple and controllable music generation. In Advances in Neural Information Processing Systems 36 (NeurIPS 2023). https://doi.org/10.48550/arXiv.2306.05284
  6. Défossez, A., Copet, J., Synnaeve, G., & Adi, Y. (2023). High fidelity neural audio compression. Transactions on Machine Learning Research. https://openreview.net/forum?id=ivCd8z8zR2
  7. European Broadcasting Union. (2023). R 128: Loudness normalisation and permitted maximum level of audio signals (Version 5.0). https://tech.ebu.ch/publications/r128
  8. Evans, Z., Carr, C. J., Taylor, J., Hawley, S. H., & Pons, J. (2024). Fast timing-conditioned latent audio diffusion. In Proceedings of the 41st International Conference on Machine Learning (Vol. 235, pp. 12652–12665). PMLR. https://proceedings.mlr.press/v235/evans24a.html
  9. Evans, Z., Parker, J. D., Carr, C. J., Zukowski, Z., Taylor, J., & Pons, J. (2024). Long-form music generation with latent diffusion. In Proceedings of the 25th International Society for Music Information Retrieval Conference (pp. 429–437). https://doi.org/10.48550/arXiv.2404.10301
  10. Evans, Z., Parker, J. D., Carr, C. J., Zukowski, Z., Taylor, J., & Pons, J. (2025). Stable Audio Open. In ICASSP 2025 - 2025 IEEE International Conference on Acoustics, Speech and Signal Processing (pp. 1–5). IEEE. https://doi.org/10.1109/ICASSP49660.2025.10888461
  11. Faller, C., & Baumgarte, F. (2003). Binaural cue coding - Part II: Schemes and applications. IEEE Transactions on Speech and Audio Processing, 11(6), 520–531. https://doi.org/10.1109/TSA.2003.818108
  12. Forsgren, S., & Martiros, H. (2022). Riffusion [Computer software]. GitHub. https://github.com/riffusion/riffusion-hobby
  13. Google Cloud. (n.d.). Lyria 2 (lyria-002) [Technical documentation]. Vertex AI. Retrieved September 30, 2026, from https://docs.cloud.google.com/vertex-ai/generative-ai/docs/models/lyria/lyria-002
  14. Gray, A., & Markel, J. (1974). A spectral-flatness measure for studying the autocorrelation method of linear prediction of speech analysis. IEEE Transactions on Acoustics, Speech, and Signal Processing, 22(3), 207–217. https://doi.org/10.1109/TASSP.1974.1162572
  15. Grötschla, F., Solak, A., Lanzendörfer, L. A., & Wattenhofer, R. (2025). Benchmarking music generation models and metrics via human preference studies. In ICASSP 2025 - 2025 IEEE International Conference on Acoustics, Speech and Signal Processing (pp. 1–5). IEEE. Data set: https://huggingface.co/datasets/disco-eth/AIME. Preprint: https://doi.org/10.48550/arXiv.2506.19085
  16. Gui, A., Gamper, H., Braun, S., & Emmanouilidou, D. (2024). Adapting Frechet audio distance for generative music evaluation. In ICASSP 2024 - 2024 IEEE International Conference on Acoustics, Speech and Signal Processing (pp. 1331–1335). IEEE. https://doi.org/10.48550/arXiv.2311.01616
  17. International Telecommunication Union. (2023). Algorithms to measure audio programme loudness and true-peak audio level (Recommendation ITU-R BS.1770-5). https://www.itu.int/rec/R-REC-BS.1770
  18. Katz, B. (2015). Mastering audio: The art and the science (3rd ed.). Focal Press.
  19. Kilgour, K., Zuluaga, M., Roblek, D., & Sharifi, M. (2019). Fréchet audio distance: A reference-free metric for evaluating music enhancement algorithms. In Proceedings of Interspeech 2019 (pp. 2350–2354). https://doi.org/10.21437/Interspeech.2019-2219
  20. Kumar, R., Seetharaman, P., Luebs, A., Kumar, I., & Kumar, K. (2023). High-fidelity audio compression with improved RVQGAN. In Advances in Neural Information Processing Systems 36 (NeurIPS 2023). https://doi.org/10.48550/arXiv.2306.06546
  21. Liu, C., Wang, H., Zhao, J., Zhao, S., Bu, H., Xu, X., Zhou, J., Sun, H., & Qin, Y. (2025). MusicEval: A generative music dataset with expert ratings for automatic text-to-music evaluation. In ICASSP 2025 - 2025 IEEE International Conference on Acoustics, Speech and Signal Processing (pp. 1–5). IEEE. https://doi.org/10.48550/arXiv.2501.10811
  22. Liu, H., Chen, K., Tian, Q., Wang, W., & Plumbley, M. D. (2024). AudioSR: Versatile audio super-resolution at scale. In ICASSP 2024 - 2024 IEEE International Conference on Acoustics, Speech and Signal Processing. IEEE. https://doi.org/10.1109/ICASSP48485.2024.10447246
  23. Liu, H., Yuan, Y., Liu, X., Mei, X., Kong, Q., Tian, Q., Wang, Y., Wang, W., Wang, Y., & Plumbley, M. D. (2024). AudioLDM 2: Learning holistic audio generation with self-supervised pretraining. IEEE/ACM Transactions on Audio, Speech, and Language Processing, 32, 2871–2883. https://doi.org/10.1109/TASLP.2024.3399607
  24. Melechovsky, J., Guo, Z., Ghosal, D., Majumder, N., Herremans, D., & Poria, S. (2024). Mustango: Toward controllable text-to-music generation. In Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers) (pp. 8293–8316). Association for Computational Linguistics. https://doi.org/10.18653/v1/2024.naacl-long.459
  25. Meta AI. (n.d.). MusicGen: Simple and controllable music generation [Documentation]. AudioCraft, GitHub. Retrieved September 30, 2026, from https://github.com/facebookresearch/audiocraft/blob/main/docs/MUSICGEN.md
  26. Odena, A., Dumoulin, V., & Olah, C. (2016). Deconvolution and checkerboard artifacts. Distill. https://doi.org/10.23915/distill.00003
  27. Oh, H. (2026a). ArtifactNet: Detecting AI-generated music via forensic residual physics [Preprint]. arXiv. https://doi.org/10.48550/arXiv.2604.16254
  28. Oh, H. (2026b). ArtifactBench: Lineage-aware evaluation of AI-generated music detectors under distribution shift [Preprint]. arXiv. https://doi.org/10.48550/arXiv.2609.23550
  29. Pons, J., Pascual, S., Cengarle, G., & Serrà, J. (2021). Upsampling artifacts in neural audio synthesis. In ICASSP 2021 - 2021 IEEE International Conference on Acoustics, Speech and Signal Processing (pp. 3005–3009). IEEE. https://doi.org/10.1109/ICASSP39728.2021.9414913
  30. Rahman, M. A., Hakim, Z. I. A., Sarker, N. H., Paul, B., & Fattah, S. A. (2025). SONICS: Synthetic or not - Identifying counterfeit songs. In International Conference on Learning Representations (ICLR 2025). https://openreview.net/forum?id=PY7KSh29Z8
  31. Scheirer, E., & Slaney, M. (1997). Construction and evaluation of a robust multifeature speech/music discriminator. In 1997 IEEE International Conference on Acoustics, Speech, and Signal Processing (Vol. 2, pp. 1331–1334). IEEE. https://doi.org/10.1109/ICASSP.1997.596192
  32. Shannon, C. E. (1949). Communication in the presence of noise. Proceedings of the IRE, 37(1), 10–21. https://doi.org/10.1109/JRPROC.1949.232969
  33. Spotify. (n.d.). Loudness normalization on Spotify. Spotify for Artists. Retrieved September 30, 2026, from https://support.spotify.com/us/artists/article/loudness-normalization/
  34. Yuan, R., Lin, H., Guo, S., Zhang, G., Pan, J., Zang, Y., Liu, H., Liang, Y., Ma, W., Du, X., Du, X., Ye, Z., Zheng, T., Jiang, Z., Ma, Y., Liu, M., Tian, Z., Zhou, Z., Xue, L., . . . Guo, Y. (2025). YuE: Scaling open foundation models for long-form music generation [Preprint]. arXiv. https://doi.org/10.48550/arXiv.2503.08638
  35. Zeghidour, N., Luebs, A., Omran, A., Skoglund, J., & Tagliasacchi, M. (2022). SoundStream: An end-to-end neural audio codec. IEEE/ACM Transactions on Audio, Speech, and Language Processing, 30, 495–507. https://doi.org/10.1109/TASLP.2021.3129994

[DATE2026-09-30][LABPT-R]

[TRACKS780][REFS35]