◈ Latent
Strixy

Qwen3.8-27B Model Card Architecture

This reference documents the architectural specifications of the Qwen3.8-27B model, including its hybrid attention mechanism and native context length capabilities.

Published 2026-10-03T15:26:43.609643+00:00

Source-reported / officially documented. No local benchmark is implied.

Architectural Composition

The model card identifies Qwen3.8-27B as a causal language model with a vision encoder. It specifies a hidden dimension of 5120 and 64 layers. The architecture utilizes a specific layout combining Gated DeltaNet and Gated Attention blocks. This documentation defines the structural components without detailing specific hardware compatibility or installation procedures for local deployment environments.

Context and Attention Limits

The source states the model supports a native context length of 262,144 tokens, extensible to one million. It details 48 linear attention heads for V and 16 for QK in the DeltaNet component. These figures represent documented model parameters. The evidence does not provide performance benchmarks or verify runtime stability on specific local hardware configurations.

Sources & applicability

  • Qwen · model card
    Original date: Not supplied · Retrieved: 2026-10-03T15:24:11.020996+00:00
    Versions: Qwen3.8-27B
    SHA-256 57e4bdb258ee1a7d2635c5174ebd4e56abe392505cdb5f8bbb356b0dc4293641

Immutable revision 6d9d3d737ff3205a4e8bdcbb570b6168485a966faa59d7c6ddf7af65f131f7f9