One question, from every angle: how much useful capability can a fixed intelligence get per unit of computation, and what moves that number? Each paper holds the model fixed, measures capability with an external verifier, and changes one thing: search, the evaluator, the judge, taste, feedback, labels, the internal ceiling, experience, symmetry, representation.
git clone https://github.com/LouayAlsakka/efficient-thinking && cd efficient-thinking bash experience/reproduce_matched_compute.sh # base at the head's compute vs the head, n = 300, paired
It refuses to print a table if any gate fails and exits non-zero if the published interval is not reproduced. An independent run on a second machine by someone who had not touched the work read +8.0 [+3.0, +13.3] against the published +9.3 [+3.7, +15.0].
| # | Title | State | Read |
|---|---|---|---|
| I | Measuring What Capability Costs — search substitutes for size on a frozen chess evaluator | measured | html · pdf · md |
| II | Where Search Pays and Where It Can't — the evaluator's ceiling, in games and in language | measured | html · pdf · md |
| III | Efficient Judging — when an LLM judge is worth its compute | measured, draft | pdf · md |
| IV | Search Where Taste Is the Evaluator | machine arms measured; rater arms running | proposal |
| V | The Exchange Rate of Feedback | concept | concept |
| VI | The Label Ceiling — why self-play plateaus | measured, unwritten | concept · ledger |
| VII | The Elicitation Gap — what a fixed system can get from inside | concept; probe running | concept |
| VIII | Experience Priors — the constraint and the carrier | closed, reproduced | html · pdf · md |
| VIII-b | Accumulation — does experience move the frontier again | measured; paper in progress | record |
| IX | Symmetry — invariance as thought compression | concept | concept |
| X | Machine-Native Representation | concept | concept · plan |
Every paper opens with a state line — MEASURED, REGISTERED, CONCEPT or IDEA — so that a concept is never read as a finding. Withdrawn claims stay in the record, dated. Readings are written before runs. Result artefacts name their instruments by hash. The series plan and the one-question rule are in series-completion-plan.md.
There is no venue, no referee and no gate. Review is public and runs through the repository's issues. The four places the author thinks Paper VIII is weakest, so you know where to aim: an alternative explanation for the matched-compute frontier shift; leakage in how the experience was constructed; a flaw in the paired statistics; a reason the second search structure does not establish transfer of the mechanism. A fatal flaw is worth more than praise, and the author will run any experiment you name.
Louay Alsakka. An electrical engineer from the semiconductor industry; this is a first research project, done outside any institution with AI assistance, and published directly. The chess engine of Paper I is where the two interests meet.
Repository: github.com/LouayAlsakka/efficient-thinking