- YuE2 Concept Sliders v2
- Choose an adapter
- What changed in v2
- Controls and downloads
- Listen to v2
- How the sliders learn
- Why train with particles?
- Historical evidence: why we pursued this recipe
- Ordinary LoRA and the routed-particle adapter
- Inference: a nonlinear correction in AR attention
- Seedbank supervision
- Normalize the paired edit
- Paired-error game and noise
- Global-mix critic
- Objectives, gradient cap and moving average
- Ordinary LoRA distillation and ComfyUI
- Reproduction and previous versions
- Choose an adapter
YuE2 Concept Sliders v2
Sixteen voice and genre controls for YuE2-3B. Version 2 publishes the final 1,600-update EMA gmix teachers, newly distilled rank-8 ordinary LoRAs, ComfyUI exports and matched listening comparisons.
Download v2 Β· How it works and math Β· Distillation method and measurements Β· Native usage Β· Live Space β v1 particles
Choose an adapter
| V2 option | Downloads | ComfyUI loading |
|---|---|---|
| Native routed particles | All 16 teachers | YuE2 Concept Slider custom node |
| Ordinary rank-8 distills | ComfyUI files Β· Native files | Standard Load LoRA, MODEL 0, CLIP 1 |
Standard LoRA workflow Β· Native particle workflow Β· Download all ComfyUI distills
Place files ending in _comfyui.safetensors in ComfyUI/models/loras/.
Strength 0 is Off; 1 is the trained positive endpoint; 0.5 is intermediate.
Negative strengths are unsupported. The ordinary students approximate the
nonlinear teachers. QKV is fused to rank 24; O remains rank 8.
What changed in v2
The inference branch still routes through 128 learned four-dimensional particles shared across 112 AR attention projections. Training now uses 512 seedbank sources per control, 32-token shared histories, paired-edit normalization and an eight-token, width-48 global-mix critic with four attention heads. Each critic token mixes the complete hidden state. The bounded critic score, paired shared noise, lazy gradient cap and particle variance/covariance penalty define the adversarial game.
The catalog records the actual schedules: Metal retains the longer anneal; Pop and Hip-Hop hold noise at 1; the remaining controls use the per-run 1.3Γ edit-RMS hold. The final noise does not reach 0.03. Female and Male use the retained expanded-cue h13 runs. All 16 are final EMA checkpoints at update 1600. This is a fixed-budget release; no listening-quality selection is claimed. See the full equations and training traces.
Controls and downloads
| Control | Sound | Particles | Ordinary LoRA |
|---|---|---|---|
| Female | One adult female lead with clear melodic phrasing | Native | ComfyUI Β· Native |
| Male | One adult male lead with clear melodic phrasing | Native | ComfyUI Β· Native |
| Pop | Clear hooks, crisp drums and a polished chorus | Native | ComfyUI Β· Native |
| Hip-Hop | Rapped verses, deep sub bass and nimble hats | Native | ComfyUI Β· Native |
| R&B | Warm keys, deep pocket and fluid vocal phrasing | Native | ComfyUI Β· Native |
| Indie Rock | Chiming guitars, moving bass and a human drum kit | Native | ComfyUI Β· Native |
| Pop Punk | Palm-muted power chords and driving chorus drums | Native | ComfyUI Β· Native |
| Metal | Heavy guitar riffs, tight kicks and big melodic choruses | Native | ComfyUI Β· Native |
| Country | Acoustic strum, twangy fills and an easy backbeat | Native | ComfyUI Β· Native |
| Acoustic Folk | Fingerpicked strings and a warm small-room performance | Native | ComfyUI Β· Native |
| House | Steady club kick, offbeat hats and a rolling bass line | Native | ComfyUI Β· Native |
| Disco Funk | Elastic bass, clipped guitar and bright dance-floor strings | Native | ComfyUI Β· Native |
| K-pop | Sharp synth hooks, tight edits and a big chorus lift | Native | ComfyUI Β· Native |
| Reggaeton | Dembow drums, rounded sub bass and clipped melodic hooks | Native | ComfyUI Β· Native |
| Afrobeats | Interlocking percussion, melodic bass and buoyant guitar | Native | ComfyUI Β· Native |
| Lo-fi | Soft swung drums, mellow keys and gentle tape warmth | Native | ComfyUI Β· Native |
Listen to v2
Same caption, lyrics and seed 1709 within each group. Off is the base; Particles is the v2 teacher; Distill is the ordinary LoRA applied during composition. Female and Male use an opposite-voice caption at Off, then apply the target control to that same caption at On.
All clips below use the second held-out prompt, a 500-token diagnostic guard (about 20 seconds), and 16 acoustic steps. Native GPU recordings are provided as original FLAC and MP3 listening copies. The sample folders retain both held-out prompts, the original diagnostic recordings and caption/lyric sidecars. Full-song reliability and full ComfyUI GPU audio generation are unvalidated. Actual ComfyUI CPU loader, weight and particle-integration checks pass for the released formats.
Female
Off
Particles
Distill
Male
Off
Particles
Distill
Pop
Off
Particles
Distill
Hip-Hop
Off
Particles
Distill
R&B
Off
Particles
Distill
Indie Rock
Off
Particles
Distill
Pop Punk
Off
Particles
Distill
Metal
Off
Particles
Distill
Country
Off
Particles
Distill
Acoustic Folk
Off
Particles
Distill
House
Off
Particles
Distill
Disco Funk
Off
Particles
Distill
K-pop
Off
Particles
Distill
Reggaeton
Off
Particles
Distill
Afrobeats
Off
Particles
Distill
Lo-fi
Off
Particles
Distill
How the sliders learn
Version 2 contains the final EMA checkpoint at 1,600 updates for all 16
controls. The inference architecture remains the routed-particle adapter. The
training formulation changes its training sources, error normalization, critic
and noise schedule. The release also creates new ordinary rank-8 students from
these exact v2 teachers. The archived v1 math describes the earlier release.
The embedded adapter-format identifier still ends in ar-v1: its tensor
layout and inference architecture are unchanged. It is separate from the
release version and the training formulation.
Why train with particles?
The aim is to learn a useful change in YuE2's behavior while keeping the base model frozen. G is the trainable slider attached to YuE2; D is a critic that learns to recognize the remaining error against a positive-caption hidden target. G learns the correction through that adversarial signal.
The particle cloud gives the adapter a shared, trainable set of features to draw on. Each projection routes its current input through those features, then combines them with its own low-rank features using a nonlinear network. The routing can change with the musical context. The variance/covariance penalty encourages the cloud to retain spread, while the critic gradient cap discourages overly steep critic responses. Together, these are the design rationale for a more flexible adapter and a manageable training game. Their separate effects require controlled experiments.
The recipe draws on ParticleGAN at revision 441fdf42:
learned particles, relativistic adversarial losses, a gradient cap and
variance/covariance regularization. Our YuE2 adapter uses input-dependent
routing through the cloud inside transformer projections. The critic is
needed during training only; the cloud and routing remain in the native
adapter at inference. Distillation learns an ordinary LoRA that approximates
this correction for standard loaders.
Historical evidence: why we pursued this recipe
The following graph is restored from v1. It compares earlier Metal, seed 7 experiments using the same frozen base, four prompt/lyric pairs and rank-8 attention targets. It motivated further work on the particle recipe. The v2 checkpoints use the revised formulation below and finish at 1,600 updates; they are not the runs plotted here.
Full-size graph Β· Measurements Β· Recipes and provenance
| Historical recipe | Updates | Peak G adversarial loss | Update at peak |
|---|---|---|---|
| Ordinary LoRA, original GAN recipe | 600 | 549.76 | 373 |
| Routed particles, v1 recipe | 1,200 | 35.66 | 945 |
Every logged update is shown without smoothing. The upper panel plots the generator's adversarial term, excluding particle regularization, on a log scale. The lower panel measures alignment with the hidden target. The earlier ordinary-LoRA run has a large loss spike and loses alignment; the particle run stays close to the target through the released v1 Metal checkpoint. The dashed line marks the end of the 600-update run.
This is evidence about the complete historical recipes. Their critics, normalization, noise, batching, optimizer settings and regularization also differ, so the graph cannot attribute the change to particles alone or rank audio quality from loss values. The ordinary-LoRA curve is an earlier GAN experiment; the current Distill adapters instead learn from particle teachers through regression and hidden-state matching.
Ordinary LoRA and the routed-particle adapter
Full-size architecture diagram
Both paths add a strength-scaled correction to the frozen projection:
An ordinary LoRA compresses the input with A and expands it with B. Their product is a fixed matrix, so the update can be merged as Wβ + sBA. The particle adapter also routes the compressed input through its cloud to obtain z(x). Its nonlinear bridge combines these features before U. The extra capability is nonlinear computation inside the adapter itself; the full base model remains nonlinear in both cases.
Distill learns separate matrices A and B to approximate the teacher's whole correction on representative activations. It trades that nonlinear inference path for the simple two-matrix update supported by ordinary LoRA loaders. The architecture diagram applies to both v1 and v2.
Inference: a nonlinear correction in AR attention
YuE2 has 28 autoregressive layers. Each slider modifies their Q, K, V and O projections: 112 branches, each with rank/alpha 8/8. NAR attention, both MLP paths, embeddings, normalization, output head and VAE remain frozen. During native generation the adapter runs in semantic composition, and its hooks are removed before acoustic synthesis.
For input column vector x, a branch uses its own down projection V, router R, nonlinear bridge Ο and up projection U. The learned cloud P contains 128 four-dimensional vectors shared across this slider's branches:
The router has three width-16 hidden layers; the bridge has three width-48 hidden layers. Both use LeakyReLU with slope 0.2. The cloud is learned once per slider and stays fixed at inference; its mixing weights depend on each input. Strength zero bypasses the branch exactly. Strength one is the trained positive endpoint. Intermediate strengths scale the correction; negative strengths are unsupported. This nonlinear adapter cannot be merged into a fixed base-weight update. The critic is used only in training.
Seedbank supervision
Each of four sound-only caption/lyric templates supplies 128 distinct continuation seeds, giving 512 training sources. For each source the frozen base generates a 32-token neutral history. The same history is appended to the neutral and positive prefixes. Their final hidden states are nα΅’ and tα΅’; the student on the neutral prefix and that history produces gα΅’. The target is the raw positive state tα΅’. This is supervision at the end of each sampled history, rather than a loss on every music token.
The generator and critic independently draw batches of eight sources with replacement. Repeated sources share a model forward within a phase but receive independent noise draws. The run seed is 7. The kept Female and Male runs use the expanded vocal cues and the h13 noise hold recorded in their metadata. The failed one-word Male run and the superseded gender candidates are excluded.
Normalize the paired edit
Let eα΅’ = tα΅’ β nα΅’ and H = 2048. Compute sample standard deviations over the 512 paired edits, then choose a scalar gain so their median normalized row RMS is one:
The implementation subtracts the mean positive target from each hidden state
and divides coordinatewise by s. The mean cancels in the paired
error, so the critic sees (gα΅’ β tα΅’) / s. In v1 the coordinate scales came
from absolute target states. Here they come from the desired edit, so small
vocal edits are not measured against the much larger spread of unrelated
hidden states. Median row RMS is one; overall RMS E generally differs from
one. teacher-audit.json records both the normalization and initial noise.
Paired-error game and noise
For each source draw one shared Gaussian vector for its real/fake pair:
The v2 set spans three recorded schedules. Metal has T = 8000 and its
original exponential anneal (the final recorded noise is about 1.537).
Pop and Hip-Hop have T = 1600 with a fixed hold of 1; Pop resumed with
the hold after update 462. The other 13 controls, including both kept gender
runs, use T = 1600 and hold the scheduled noise at a minimum of 1.3E.
For those holds the exponential start exceeds the hold. Consequently the
release does not reach noise 0.03 at update 1600. Per-run traces and
normalization audits are included in evidence/particle-gmix-1600-v2/.
Global-mix critic
The critic learns a linear map from all 2048 error coordinates into eight 48-dimensional tokens and adds learned token positions. One four-head attention block processes those tokens. Each token can mix the entire hidden state; contiguous hidden coordinates are not assumed to be meaningful patches. The block uses RMS normalization, residual self-attention and a residual MLP. Afterward the critic independently RMS-normalizes the mean and elementwise maximum across tokens, concatenates them, and predicts a bounded score:
The actual checkpoint setting is gmix_t8_w48_l1, with four heads and score
bound 8. Later trainer defaults are not a description of these trained files.
Objectives, gradient cap and moving average
Using paired relativistic logistic losses:
Every fourth update applies the lazy gradient cap
It is zero on other updates. The factor four compensates for the lazy frequency. A fresh subset of 64 particles, sampled without replacement, gives sample covariance C with denominator 63. Its variance/covariance penalty is
There is no extra output MSE, lyric preservation or ending loss in teacher training. Adam uses betas (0, 0.999), zero weight decay and constant rates: 0.0006 for the adapter branches, 0.006 for the shared particles, and 0.0009 for the critic. After each update all learned adapter parameters enter an EMA:
The release uses final EMA weights, not a listening-quality selection. Training cosines and critic losses diagnose hidden-state optimization; they do not establish perceptual quality, lyric preservation or natural endings.
Ordinary LoRA distillation and ComfyUI
For each exact v2 teacher, regression fits a rank-8 linear down projection to its routed features while retaining its up projection. Hidden-state refinement then optimizes both matrices, selecting on a reserved training lyric sheet. The two evaluation lyric sheets stay excluded from fitting and selection. See DISTILLATION.md for the regression, refinement and held-out error equations and the resulting measurements.
An ordinary student adds sBAx. ComfyUI fuses Q/K/V by concatenating their down matrices and placing their up matrices on a block diagonal: fused QKV rank 24, O rank 8, with alpha/rank preserved. Conversion is exact before BF16 rounding; teacher-to-student distillation is an approximation. Standard Load LoRA, MODEL 0, CLIP 1, also affects acoustic-prefix processing. The native particle custom node applies its correction only during AR generation. The gallery's Distill examples apply the ordinary LoRA only during autoregressive composition.
Reproduction and previous versions
Release tools create the ordinary distills from the versioned teacher catalog, then convert and validate the exports. The v2 manifest records file hashes; each student records its exact teacher hash. Training and evaluation prompts contain sound descriptions. Weights retain CC BY-NC 4.0 terms.
V1 particle weights Β· V1 ordinary LoRAs Β· V1 README and math at their original revision. The live Space continues to run its pinned v1 particle deployment.
Model tree for ntc-ai/yue2-concept-sliders
Base model
m-a-p/YuE2-3B