docs/source/en/model_doc/muse_glimmer_assistant.md
This model was contributed to Hugging Face Transformers on 2026-08-09.
<div style="float: right;"> <div class="flex flex-wrap space-x-1"></div>
MuseGlimmerAssistant is the DFlash drafter for MuseGlimmer. It is not a standalone language model. It has 5 sliding window layers and no embeddings of its own. It borrows the main model's input and output embeddings, and reads the main model's hidden states at target_layer_ids (layers 1, 13, 25, 37, and 49 by default) as context.
Rather than drafting one token at a time, the drafter denoises a whole block of block_size masked tokens in a single forward pass, like a diffusion window. The main model then verifies the block in one step. Meta reports 3.1x faster decoding on an RTX 5090 and 1.5-1.8x on Apple M-series chips.
Pass the drafter to [~GenerationMixin.generate] as assistant_model and set speculation_type="dflash". The drafter must be loaded in the same dtype and on the same device as the main model.
from transformers import AutoProcessor, MuseGlimmerAssistantModel, MuseGlimmerForConditionalGeneration
processor = AutoProcessor.from_pretrained("meta-models/Muse-Glimmer-30B")
model = MuseGlimmerForConditionalGeneration.from_pretrained(
"meta-models/Muse-Glimmer-30B",
device_map="auto",
)
drafter = MuseGlimmerAssistantModel.from_pretrained(
"meta-models/Muse-Glimmer-30B-assistant",
device_map="auto",
)
messages = [
{
"role": "user",
"content": [{"type": "text", "text": "Write a bash one-liner that counts lines of Python in a repo."}],
},
]
inputs = processor.apply_chat_template(
messages,
add_generation_prompt=True,
tokenize=True,
return_dict=True,
return_tensors="pt",
).to(model.device)
input_len = inputs["input_ids"].shape[-1]
outputs = model.generate(
**inputs,
assistant_model=drafter,
speculation_type="dflash",
max_new_tokens=256,
)
response = processor.decode(outputs[0][input_len:], skip_special_tokens=False)
print(response)
generate forces output_hidden_states=True for the target model when speculation_type="dflash".[[autodoc]] MuseGlimmerAssistantConfig
[[autodoc]] MuseGlimmerAssistantPreTrainedModel
[[autodoc]] MuseGlimmerAssistantModel - forward