This is not a contest of size. It is a contest of sovereignty, integrity, and accuracy: Barq Al-Sagheer (Barq Kid) — a grounded sovereign Arabic model that answers from verified knowledge, not from memory — facing giant models that answer from their memory.
Transparency first: we declare what we measure and what we do not, so the benchmark is defensible before any critic.
Each model's size is shown for context only, and it is the first piece of information — but it is not the metric. The metric is the trust and accuracy the model delivers to the Arabic user.
① Sovereignty (owned/private vs. hosted abroad) · ② Fabrication (inventing an ayah/number/chain of authority) · ③ Accuracy (correctness with verifiable evidence).
Barq Al-Sagheer is an integrated sovereign system, grounded by design — built to answer from verified knowledge rather than memory, so it does not fabricate. (Its internal architecture is proprietary and reserved.) The other models are the original base models (raw, native) — called directly: a question and an answer, with no tools, no web, no wrapper — answering from their parametric memory.
The Arabic language, the Quran, and Islamic knowledge — where sovereignty and zero-fabrication become decisive, not a luxury.
Size is the first fact — not the final verdict.
| Model | Score /64 | Total | ✅ Correct | ⚠️ Partial | ❌ Wrong | 🤷 I don't know | 🚩 Fabrication (danger) | 📎 Grounded | Sovereignty |
|---|---|---|---|---|---|---|---|---|---|
| برق الصغير | ٤٢/٦٤ | ٣٢/٣٢ | ٢١ | ٠ | ١٠ | ١ | ٠ ✅ | 28/32 | ✅ |
| DeepSeek‑V4 | ٥٤/٦٤ | ٣٢/٣٢ | ٢٧ | ٠ | ٤ | ٠ | ١ 🚩 | — | ❌ |
| Qwen2.5‑7B | ٣٢/٦٤ | ٣٢/٣٢ | ١٦ | ٠ | ١٣ | ٠ | ٣ 🚩 | — | ⚠️ |
| GPT‑5.4‑mini | ٦٠/٦٤ | ٣٢/٣٢ | ٣٠ | ٠ | ٢ | ٠ | ٠ ✅ | — | ❌ |
| # | Model | Accuracy 50% | Integrity 50% (each fabrication −20) | Sovereign Score /100 |
|---|---|---|---|---|
| 🥇 | GPT‑5.4‑mini | 43.8/٥٠ | ٠ → 50/٥٠ | 93.8/١٠٠ |
| 🥈 | برق الصغير | 26.6/٥٠ | ٠ → 50/٥٠ | 76.6/١٠٠ |
| 🥉 | DeepSeek‑V4 | 40.6/٥٠ | 1 → 30/٥٠ | 70.6/١٠٠ |
| ٤ | Qwen2.5‑7B | 26.6/٥٠ | 2 → 10/٥٠ | 36.6/١٠٠ |
We hide nothing — this is an honest reading of a first round
«مَن قال لا أعلمُ فقد أفتى» — Being small in size does not diminish honesty
Chosen to cover seven categories of increasing complexity — from simple facts to traps that expose fabrication. Under each question is the documented "Answer key," then each model's answer (to be filled in after the run).
The benchmark does not ask "who is bigger?" but "when the model does not know — what does it do?". This section justifies every methodological choice and sets it before the harshest of critics.
Everything Barq 373M lags in is a capacity constraint scheduled for a fix — not a flaw in the methodology. Three stations, with integrity constant in all of them:
Zero-fabrication is not a phase to be surpassed — it is the foundation upon which every coming size is built. We grow the model so it answers more, not so it fabricates more.