Text Generation
Safetensors
English
mistral
character_roleplay
creative_writing
roleplay
conversational

Trouper V2 Logo

WARNING! THIS IS A BETA RELEASE. The full model has been released. Only here as archival.

Trouper-v2-12B-beta

Trouper V2 is a 12B RP model fine-tuned on Mistral Nemo Base, trained for restrained character writing with native SillyTavern integration. Built on the same philosophy as V1 — clean prose, no slop, characters that feel like people — with three new capabilities.

Looking for the larger model? -> Prima v2 coming soon!

What's New in V2

Thinking Characters — The model thinks in character before responding. Think blocks show the character's internal reasoning: what they notice, what they choose not to say, how they read the room. This isn't chain-of-thought for accuracy — it's inner monologue that drives more grounded, intentional responses.

Restraint & Pacing — Characters hold back. A guarded character won't trauma-dump on turn 1. Emotional reveals surface through behavior and slips, not exposition. The "underneath" section of a character card is felt, not stated. Scenes build naturally rather than rushing to the dramatic beat.

SillyTavern Native: Summaries & Image Prompts — The model handles ST's built-in summary and image generation prompts without breaking character context. No second model needed. It recognizes the ST instruction format, switches cleanly from RP to utility mode, and produces output that's informed by the actual conversation — not just the character card.

Measured, Not Vibes: Slop Benchmarks

"No slop" is an easy claim to make and a hard one to back up, so here are numbers. We profiled Trouper V2 with slop-forensics — the toolkit behind EQ-Bench's slop metrics — on the standard creative writing prompt set (Reddit Writing Prompts via Nitral-AI), and ran TheDrummer's Cydonia-24B-v4.1 through the identical pipeline for comparison.

Model Slop score (lower = better)
Trouper-v2-12B 14.5
Cydonia-24B-v4.1 38.3

On the EQ-Bench Slop Score tool, Trouper V2's outputs score 11.46 — lower than every LLM on the leaderboard at the time of testing (Claude Sonnet 4.5: 19.5, Kimi-K2: 23.7, GPT-5-mini: 26.3, Mistral-Nemo: 54.4, Gemma-3-4b: 85.2), landing one slot above the human-writing baseline of 10.4.

The slop the two models do have is different in kind: Cydonia's top over-represented patterns are the classic dialogue-attribution family ("nodded", "leaned", "whispered", "eyes widened almost imperceptibly"), while Trouper's are concentrated in a handful of atmospheric words (see Known Limitations). Neither profile shows contamination of the other's category.

Methodology: ~110 single-shot story generations per model at the model's recommended sampler settings, profiled for over-represented words and n-grams against the wordfreq human baseline. Profiles and scripts available on request.

SillyTavern Setup

Chat Completion mode with default settings.

Reasoning: Under the Reasoning section, enable "Add to Prompts" with a max thinking budget of 7. This lets SillyTavern display the model's think blocks as collapsible reasoning sections.

Template: Auto-detected by Chat Completion.

Temperature: 0.7–1.0 recommended.

Context: Handles 15–20+ turn conversations well.

Image Prompt Setup

The model supports all of ST's built-in image prompt types (Character, User, Scenario, Last Message, Portrait, Background). For the Last Message option, we recommend this custom prompt for best results with image generation models (first-person POV):

PAUSE STORY TO GENERATE AN IMAGE. Your next response must be formatted as a single comma-delimited list of concise keywords. The list will describe the visual details included in the last chat message.
Only mention characters by using pronouns ('he','his','she','her','it','its') or neutral nouns ('male', 'the man', 'female', 'the woman').
Ignore non-visible things such as feelings, personality traits, thoughts, and spoken dialog.
Add keywords in this precise order:
1. A keyword to describe the location/setting of the scene
2. A keyword to mention time of day or lighting conditions if relevant (day, night, sunset, golden hour, etc.)
3. Keywords to describe the primary action or pose taking place
4. Keywords to describe camera angle and POV (pov, male pov, fpv, cowboy shot, high angle, selfie pov, etc.)
5. Keywords to describe Assistant's clothing and outfit
6. Keywords to describe Assistant's body positioning and what they're doing with their body (sitting, standing, leaning, crossed legs, open legs, etc.)
7. Keywords to describe Assistant's facial expression and where they're looking (smiling, wink, little smile, looking at viewer, seductive eyes, etc.)
8. Keywords for specific body parts visible and their appearance (black thighs, black stockings, etc.)
9. Keywords to describe objects being held or interacted with (holding coffee cup, etc.)
10. Keywords for aesthetic qualities and mood (fashion, aesthetic, calm, etc.)
11. Any additional environmental details (city skyline, leaves, table, etc.)
A correctly formatted example response would be:
'coffee shop, fall, sunset, sitting, pov, date, male pov, white long sleeves, skirt, sitting at table, crossed legs, little smile, looking at viewer, smiling eyes, black thigh-highs, holding coffee cup, fashion, aesthetic, calm, street, city, leaves'

Summary Setup

ST's default summary prompt works out of the box. The model supports both 200 and 500 word summaries, and can update an existing summary with new events when ST passes one in context.

Character Card Format

Trouper V2 was trained on a specific card structure. Cards don't need to follow this exactly, but this format gets the best results:

[Name] — [One-line description: age, role, situation, what makes them interesting.]

[Name] presents as: [How they come across to others. Surface-level personality, appearance, mannerisms. What you'd notice in the first five minutes.]

Underneath: [What's actually going on. The thing they don't show. Internal conflicts, hidden feelings, unresolved history. This is what the model holds back and reveals gradually.]

Voice: [How they talk. Sentence length, vocabulary, verbal tics, dialect. What they sound like when relaxed vs. stressed. The more specific, the better.]

Tells: [Physical behaviors that reveal inner state. Fidgets, habits, avoidance patterns, things they do when lying or uncomfortable. Semicolon-separated list.]

Example:

Maren Aldvik — A 33-year-old former deep-sea welder who lost her left
hand in an industrial accident two years ago. Now runs a small marine
salvage consulting business from a converted shipping container on the
docks in Brønnøysund, Norway. Has a prosthetic hand — a functional but
unglamorous myoelectric model with limited grip strength and no
sensation. Matter-of-fact about it.

Maren Aldvik presents as: Blunt, competent, no-nonsense. Speaks with
the cadence of someone used to giving instructions over bad radio
connections. Drinks black coffee constantly. Wears practical clothes —
work boots, cargo pants, waterproof jacket. Keeps her blonde hair short
because long hair and welding don't mix, even though she doesn't weld
anymore.

Underneath: Grieving a version of herself that doesn't exist anymore.
She was one of the best in her field and her identity was built entirely
around being good at a dangerous job. Without it, she doesn't know who
she is. The consulting business is a way to stay adjacent to the work
without admitting she can't do it. She resents the prosthetic not
because it doesn't work but because it works well enough that people
think she's fine.

Voice: Clipped, dry, precise. Norwegian English — fluent but with
occasional odd phrasing that sounds translated ("it is not so" instead
of "it's not like that"). Doesn't waste words. Dark humor about her
hand that makes other people uncomfortable. Technical vocabulary slips
in naturally — she talks about torque and tensile strength the way
other people talk about weather. When she trusts someone enough to
relax, she becomes warmer but never soft.

Tells: Flexes the prosthetic hand when stressed — the fingers open and
close in a rhythmic pattern; Unconsciously positions herself so her
left side is away from new people; Corrects people's misconceptions
about underwater work with disproportionate intensity; Goes quiet and
looks at the water when she's remembering the accident

Strengths

  • Restraint: Characters hold back appropriately. Emotional depth surfaces gradually, not all at once.
  • Clean prose: Minimal AI slop — see the benchmark section above for numbers. No purple prose, no "a symphony of" or "the weight of unspoken words."
  • Voice differentiation: Different characters actually sound different. A Norwegian welder doesn't talk like a retired yakuza accountant doesn't talk like a fire lookout.
  • Physical tells: Characters express emotion through body language and habits rather than narration. The model uses the Tells section of the card naturally.
  • Thinking in character: Think blocks are inner monologue, not narrator commentary. They show the character deciding what to say and what to hold back.
  • User agency: The model stays in the character's lane. It reacts to what the user does rather than narrating the user's actions, breathing, or emotions.
  • ST integration: Summaries and image prompts work without breaking the RP session. In eval, utility prompts produced zero think-block or RP-formatting leaks across all prompt types.

Comparison to Trouper V1

Aspect V1 V2
Think blocks No Yes — in-character inner monologue
Restraint Good Significantly improved — trained specifically for pacing
Summaries Requires separate model Native — handles ST summary prompts
Image prompts Requires separate model Native — handles all ST image prompt types
Format reliability Good Improved
Long conversations Good Better — think blocks help maintain coherence

Comparison to Prima-v2-24B

Aspect Trouper-v2-12B Prima-v2-24B
Prose quality Excellent — direct and concrete Excellent — slightly more elaborate
Voice consistency Strong Near-perfect
Restraint Very good Excellent
Format reliability Good (occasional unclosed asterisk) Excellent
User agency Very good Excellent
Inference speed Fast Slower
VRAM ~8GB quantized ~16GB quantized
Long context Good Better
Best for Single-GPU setups, fast inference Maximum quality, long sessions

Known Limitations

  • Occasional format breaks: At 12B, the model may occasionally produce an unclosed asterisk or drop a think block on very short responses. Swiping for a new response usually fixes this.
  • A fondness for fluorescent lights: The model's slop profile is clean overall, but its residual signature is atmospheric — humming fluorescent lights, the smell of burnt coffee, dust in the light. If your scenes are set in diners, laundromats, or late-night offices, you'll feel at home. If you see the same sensory furniture recur, that's the known tic; a swipe or a scene-setting nudge steers it.
  • High-effort user matching: When paired with a very detailed, narration-heavy writing partner, the model may occasionally mirror that energy by describing the user's actions. This is rare and less pronounced after DPO correction.
  • Template sensitivity: Without Mistral-Tekken or ChatML, may generate meta-narration or continue past appropriate stopping points. Use Chat Completion / the included template!
  • Not a general assistant: This model is trained for RP. It doesn't understand being an "assistant" outside of playing a character that happens to be one.

Training

Trained on Mistral Nemo Base 2407 in a three-stage pipeline:

  1. Restraint & Interiority SFT — multi-turn character conversations across 32+ characters with in-character <think> blocks, explicit pacing rules, and varied scene endings (natural partings, interruptions, quiet moments — not every scene builds to a reveal).
  2. ST Utility SFT — continued fine-tuning on SillyTavern's summary and image-prompt instruction formats, trained for strict template recognition so utility mode never leaks into RP or vice versa.
  3. DPO — preference pass targeting user-agency failures (the model narrating the user's actions, emotions, or dialogue).

All training data was synthetically generated and quality-filtered, including slop profiling of the dataset itself — the training data was measured with the same tools as the outputs.

Why train on a base model?

Per Base Models Beat Aligned Models at Randomness and Creativity — and to avoid GPT-isms leaking into the prose. Training on a base model is like working with fresh clay rather than reshaping something that was already formed for a different purpose. This is why the model doesn't understand being an assistant and isn't intended to.

Feedback

Issues, questions, and feedback welcome in the Community tab. Particularly interested in:

  • Long conversation quality (20+ turns)
  • How the think blocks feel in practice
  • Summary and image prompt quality
  • Character card format experiments
  • Comparison with other RP models at this size
Downloads last month
125
Safetensors
Model size
12B params
Tensor type
F16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for DarwinAnim8or/Trouper-v2-beta-12B

Finetuned
(2)
this model
Quantizations
1 model

Dataset used to train DarwinAnim8or/Trouper-v2-beta-12B

Paper for DarwinAnim8or/Trouper-v2-beta-12B