Nanolite Models: Cutting The Reasoning Models Don’t Need
Published: October 4, 2026
By: Sparset Team
Reasoning models think out loud before they answer. They write a chain of intermediate steps first, and every step is a generated token that takes time. On a large hosted model that overhead is easy to miss. On a small model running on a single modest GPU, it is most of the wait.
Much of that thinking is not needed. A model that has already solved a problem will often work it again, reconsider its wording, or check the output format one more time. The answer at the end does not change. The user just waits longer for it.
Nanolite is our experiment in cutting that overhead. We took two small open reasoning models, Qwen3-0.6B and Qwen3.5-4B, and adapted them on worked solutions so they would reach an answer in fewer steps. The architecture and parameter count stayed the same. We wanted to know two things: how much shorter generation gets, and what it costs in answer quality.
On a 40-prompt workplace pilot, Nanolite (0.6B) cut mean generation time from 14.59 seconds to 4.45 seconds, a 3.28x speedup. A fixed-length control showed where that came from: fewer tokens, not faster computation per token. Nanolite V2 (4B) used approximately 92% fewer generated tokens in its final attempts on the GSM8K and BBH questions both models completed.
Quality is the more conditional result. The 0.6B pilot did not establish broad quality preservation. The 4B model matched its base model's overall accuracy on a completion-selected subset, with small gains on two benchmarks and a regression on BBH. An overall score alone would have hidden that tradeoff. These are results for these models, workloads and hardware, not general speedups.
Nanolite: a shorter path to an answer at 0.6B
A model can reach the right answer and still spend tokens repeating the calculation, reconsidering its wording, or checking the requested format. We wanted to see whether adaptation could reduce that overhead without changing the model’s parameter count.

Forty internal workplace prompts, two model orders, 80 natural-task measurements per model. Tesla T4; F16 GGUF loaded through Transformers; thinking enabled; greedy decoding; batch size one; 1,024-token maximum. The fixed-128-token control has 16 observations per model. These are GPU measurements, not laptop or phone results. Underlying measurements.
The archived training recipe uses response distillation through LoRA: supervised learning from worked solutions and final answers, rather than matching a teacher’s logits. Its 5,000-example teaching set draws on NVIDIA’s Nemotron math, science and applied-reasoning corpora, with recorded teacher provenance evenly split between GPT-OSS-120B and DeepSeek-R1-0528. The recipe specifies one epoch, rank 16, and greater loss weight on final-answer tokens than reasoning tokens. This was a reasoning dataset, not a dedicated workplace-assistant dataset.
The same extracted fields, fewer tokens
One prompt asked:
Extract the customer, invoice number, due date, and amount: Acme Ltd invoice INV-2048 for $1,275 is due 18 September 2026.
Both models produced the same visible answer:
- Customer: Acme Ltd
- Invoice Number: INV-2048
- Due Date: 18 September 2026
- Amount: $1,275In the first model-order pass, the base model generated 318 tokens, including 273 reasoning tokens, and took 12.44 seconds. Nanolite generated 88 tokens, including 43 reasoning tokens, and took 3.46 seconds. This pair illustrates the opportunity; it is not a representative estimate by itself.

Across the pilot, mean total generated tokens fell by 69.56%, and mean reasoning tokens fell by 74.07%. Yet throughput in the fixed-length control was almost unchanged: 25.41 versus 25.73 tokens per second. Nanolite simply finished sooner because it usually generated less in both reasoning and output, not because it computed each token three times faster.

The 3.28× figure is the ratio of mean generation times. A paired bootstrap over the 40 prompts gives a 95% interval of 2.66× to 4.16×. That interval describes variation across this workload, not across hardware or future users.
Where the efficiency result stops
The historical automated checks passed 29/40 base answers (72.5%) and 30/40 Nanolite answers (75.0%). Seven prompts passed only for Nanolite; six passed only for the base. Repeating the greedy run produced identical responses, so these are 40 independent quality prompts, not 80.
That one-answer advantage is not a demonstrated quality improvement: its paired bootstrap interval spans −15 to +20 percentage points. The checks also combine simple answer rules and output constraints, not a completed semantic review. For example, Nanolite genuinely missed a two-sentence requirement, while another check incorrectly rejected a correct base-model calculation because of LaTeX formatting. We retain these as historical automated-check rates, not true-accuracy estimates.
A separate Q4_K_M deployment test shows why the runtime matters:

Separate Tesla P100 run with llama.cpp and Q4_K_M, 40 prompts and two model orders. Hardware, runtime and precision differ from the F16 pilot; this is not a controlled estimate of quantization damage.
The lower mean time did not translate into a uniformly better experience: check pass rate fell and tail latency rose. A favorable research configuration must be validated again in the actual deployment package.
Nanolite V2: testing the tradeoff at 4B
The 0.6B experiment suggested a practical opportunity: reduce unnecessary generation while preserving the useful answer. The next question was whether a larger small model could show that behavior on more demanding reasoning tasks. Nanolite V2 explored it with Qwen3.5-4B and a broader teaching set.
The checkpoint’s saved training report identifies an 8,150-example v3.0.0 dataset: 2,000 foundations, 2,045 retention examples, 1,840 frontier corrections, 2,248 standalone demonstrations and 17 review rewrites. Training again used response distillation through LoRA, with one epoch, rank 32, alpha 64, learning rate 5 × 10−5, a 2,048-token sequence limit and unquantized BF16 training.
The demonstrations were not uniformly brief: recorded reasoning had a median of 170 words and a mean of 228.85 words. Prompt tokens were masked from the loss, and final answers received greater weight than reasoning. Shorter generation is an observed outcome; without ablations, we cannot attribute it to one design choice.
First, compare answers that actually finished
The original evaluation planned 500 questions: 125 GSM8K, 285 MMLU-Redux and 90 across nine BBH tasks. GSM8K and BBH used thinking, with 1,536 new tokens initially and 2,560 on an extraction-triggered retry. MMLU-Redux used a shorter, non-thinking profile.

However, the base model frequently exhausted its reasoning budget without a completed answer. Counting those unfinished attempts as ordinary incorrect answers would overstate Nanolite’s reasoning advantage. The validity gate blocked that interpretation. Therefore, we instead inspected the 395 questions completed end-to-end by both models.
This is a useful diagnostic, but a narrower one due to 105 planned questions being excluded. The following scores measure accuracy conditional on both models finishing, not full-suite accuracy or what the base model would achieve with a sufficient budget.
Same retained question IDs for both models: 92 GSM8K, 250 MMLU-Redux and 53 BBH. Corrected final-answer normalization is applied symmetrically. These completion-selected results are not full-suite leaderboard scores. Per-question scores.

Correct answers: 89 → 90 on GSM8K, 187 → 192 on MMLU-Redux, 51 → 45 on BBH, and 327 → 327 overall. Differences are percentage points.
The audit also corrected a parser bug: an answer such as (C) Sam could be rejected against reference (C). The same normalization on both models’ final predictions restored 11 Nanolite BBH answers and changed none of the base answers. We did not search the reasoning for a correct answer, change references, or treat semantic correctness as proof of strict format compliance.
Nanolite won ten questions the base missed and lost ten the base answered correctly. The overall tie is descriptive, not proof of equivalence; its paired-bootstrap difference interval is −2.28 to +2.28 percentage points. The small positive task differences do not establish a general reasoning gain. The BBH regression remains a material result.
Much less generation, not a universal speedup
Total generated tokens include reasoning and answer text in the retained final attempts. Change is calculated from each pair of means. Thinking is enabled for GSM8K and BBH, disabled for MMLU-Redux. Earlier failed attempts are not included. Exact measurements and timing definitions.

The thinking-enabled tasks show approximately 92% fewer generated tokens. The non-thinking task moves in the opposite direction: from a mean of 5.00 to 7.70 tokens. A single efficiency claim for every workload would hide that distinction.
Our older timing figures, 14.4× on GSM8K and 15.3× on BBH, were averages of per-question ratios of allocated time. The harness divided batch time among its questions and retained only the final attempt; earlier retry time was omitted. Ratios of the mean allocated times are 13.33× and 12.59×, respectively. These are substantial reductions in recorded final-attempt time, but not isolated request latency, complete retry-inclusive cost, or laptop speed guarantees.
A useful shortcut and a real failure
For a GSM8K question about a book priced at $19.50 after a 25% discount, both models recovered the correct original price of $26. The base attempt used 2,303 tokens; Nanolite used 105. Nanolite’s generated solution was direct:
The discount is 25% of the original price, so the sale price is 75% of the original. Let the original price be x dollars. Then 0.75x = 19.50. Solving for x gives x = 19.50 / 0.75 = 26.00. Therefore the original price was $26.00.
The base output repeatedly revisited the solved calculation and answer format. In this selected example, the extra generation did not improve the answer.
A BBH ranking question exposes the other side. Three golfers, Ana, Amy and Eli, finished in order. Eli was second and below Amy, so Amy must have finished first. The base answered correctly in a 1,780-token attempt. Nanolite reversed the relation, put Amy third and answered Ana, using 126 tokens. That is a reasoning error, not a grading artifact. These illustrative pairs are preserved in the full example record; they do not establish that brevity caused the failure.
What the two experiments taught us
Both models show that adaptation can sharply change generation behavior without shrinking the architecture. At 0.6B, an unchanged token-generation rate combined with much shorter output produced a substantial latency reduction. At 4B, shorter completed solutions sometimes reached the same correct answer with far fewer tokens—but an overall tie concealed a loss on relational reasoning.
For Forge, the useful target is fewer unnecessary steps while retaining the steps the task needs. That requires measuring quality and efficiency together, on the actual runtime. Neither a shorter answer nor a lower training loss is sufficient evidence.
Our next tests should use larger independently reviewed task sets, full completion budgets, retry-inclusive timings, and both thinking and non-thinking base-model baselines. We would protect relational and state-tracking examples during training, audit train–test overlap, and test ordinary prompts without benchmark-specific answer labels. These are proposed follow-ups, not additional results reported here.
We also learned to separate model behavior from scorer behavior. Fixing an unambiguous parsing error is appropriate; dropping instruction-following requirements to improve a score is not. The evaluation has to preserve the work the model is meant to do.
Training and release provenance
The historical 0.6B repository was not publicly accessible during this review. Its archived notebook establishes the intended recipe, but a completed training report tying it to immutable evaluated weight bytes was not recovered.
The current Nanolite V2 experimental repository contains an adapter, merged checkpoint and F16 GGUF. Its saved report identifies dataset v3.0.0. A later v3.1.0 dataset removed a repeated Final answer: prefix, but we do not attribute the evaluated results to that later revision.
The historical benchmark did not pin an immutable model revision, so the current repository is not guaranteed to reproduce the exact evaluated weights. The saved GGUF report leaves runtime smoke verification pending. We make no deployment-readiness or blanket commercial-license claim.
Methods and evidence contains the source ledger, configurations, scoring corrections, uncertainty calculations and provenance limits. The accompanying evidence files retain exact values and source hashes. The measured results in this article are from the August 2026 experiments; this is a revised account under the Sparset name.
Build your own model around your workflow
Nanolite is one example of the research behind Sparset Forge, where we adapt models for specific tasks and evaluate the tradeoffs in quality, speed, and deployment requirements.
If you’re exploring a custom model for your workflow or a model that runs on your own infrastructure, reach out to Sparset or book a meeting to discuss what you need: https://www.sparset.ai/book