docs/source/en/api/models/minimax_music3_transformer.md
The 2.4B flow-matching Diffusion Transformer of MiniMax Music 3. It denoises 128-channel Flow-VAE audio latents conditioned on the per-frame hidden states of the model's autoregressive language-model stage, prepending the flow-matching timestep as an extra sequence token (a Stable-Audio-lineage continuous transformer with partial rotary attention and GLU feedforwards).
[[autodoc]] MiniMaxMusic3Transformer1DModel