The license is the real thing. Every prior Meta open release carried a Llama community license, criticized for terms such as its 700 million monthly user threshold. Apache 2.0 has no revenue cap, no user cap, no copyleft, and an explicit patent grant. Artificial Analysis rates the release 44 on its Openness Index, high among open models. The 82 percent is a calibration metric, not an error rate. It measures how often a model guesses rather than abstains when it lacks an answer. Frontier models score badly on it too. The fair complaint is narrower: at this size, Qwen manages 49 percent, so the number reflects a choice, not a constraint. The disclosure standard is above the field. A vendor that hands competitors their most favorable number, prints the rows it loses, and publishes its own scoring deviations is making itself easier to audit. Offline has a threat model of its own. Air-gapped inference removes a whole class of exposure: no prompts leaving the building, no vendor retention question, no dependency on someone else's uptime. For regulated workloads that trade is often worth a few benchmark points. |
| | 08 | if you are evaluating this week |
| Benchmark the quantization you will actually ship. The published table is a high-reasoning run with quantization losses reported separately, so treat those numbers as the unquantized ceiling. Your build is a 4-bit variant with a drafter attached. Measure your own long-horizon task completion at K-Quant-17GB before trusting the 1.0 percent. Treat the harness as the security boundary. The model contributes 28.4 and 26.4. Everything better comes from your tool allowlist, workspace scoping, loopback binding and a confirmation step before any irreversible action, which is Meta's own recommendation. Assume overconfidence and design for it. An 82 percent guess-rather-than-abstain profile means retrieval grounding and a verification step matter more here than with a hosted frontier model. Reserve it for tasks with a checkable output. Match the workload to the strength. Schema-driven tool calling and multi-turn orchestration are where it leads. Desktop control, terminal work and document parsing are where Qwen3.6-27B leads in Meta's own numbers. Read both documents, and watch for the flagship. Apache 2.0 governs the weights, and a separate usage policy ships in the same repository. Training data and the full pipeline are not released, so open weights is the precise term for what you have. If Muse Spark 1.2's weights arrive as stated, re-run your evaluation against it before committing a year of tooling to the 30B. | Meta shipped the capability and the receipts on the same day. The capability fits in 24GB. The receipts say close to three in ten adaptive injections land, private attributes reach the wrong party at more than twice Gemma's rate, and the model would rather answer than admit it does not know. None of that makes it a poor model for its size. It makes the harness, not the checkpoint, the thing you are actually choosing. |
| |
|