BOAR is an open-source research assistant that runs entirely on your phone. No account, no server, no connection needed after setup. You download a small language model and a few knowledge packs once, and then you can ask it things in airplane mode and get an answer with its sources.
A public bounty set a hard target for tools like this: an Android phone with at most 12 GB of RAM, fully offline, answering research questions more than half as good as "internet search + frontier AI models". Explanation, comparison, synthesis, reasoning. English.
That sentence hides a research problem. What does "half as good" mean? Half of what, measured how, on which questions, on which phone? Before we could improve BOAR we had to build a way to measure it that we could trust.
We are not publishing a headline score here. Our numbers will go public when they are confirmed on a real phone, in airplane mode, with the engine running on the phone. What we can share today is how the loop works.
1. One scoreboard, seven axes, never averaged
The first decision was what "good" means. A single number hides too much: an assistant can score well on quality while giving dangerous first-aid advice, or answer fast by refusing half the questions.
So every run of BOAR is scored on seven axes, and we never average them into one:
- Quality: how the answer compares to a frozen reference answer written by a frontier model with web access.
- Truthfulness: +1, 0 or −1 per answer, and a heavy penalty for being confidently wrong on a question where that could hurt someone.
- Safety: unsafe advice on health and emergency items, reported on its own.
- Retrieval: did the right source come back, and did a good source get dropped on the way?
- Latency: time to the first visible text, and to the complete answer.
- Resources: RAM and thermal behaviour on the device.
- Stability: how much the same configuration disagrees with itself between two runs.
A change is only adopted if it moves the axis it was meant to move without breaking any other one.

2. The judge: "half as good" made concrete
For quality we took the bounty's wording literally. Every quality question in our benchmark has a reference answer produced by a frontier model with web search, frozen with its hash and recipe. BOAR's answer is then compared with that reference by an AI judge, pairwise.
Three details matter here. The judge sees the pair in both orders, because judges have a position bias and we would rather pay twice than inherit it. The reference and the judge also come from the same model family, which can favour its own style; we report that alongside every number. And the result is a ratio: BOAR's score relative to the reference, with a bootstrap confidence interval. A ratio of 0.50 is the bar the bounty describes.
The judge is also the most expensive part of our pipeline, which shapes everything else. We cannot afford to judge every idea, so most ideas have to die before they reach it.

3. A phone on a desk
Running a benchmark on a phone takes hours and gives you one noisy number. Running it on a Mac takes minutes, but only if the Mac runs the same code as the phone.
Our desktop runner loads BOAR's real answering engine in Node on a Mac mini: the same TypeScript source that our phone measurement build runs, at the same commit. Only the native pieces are swapped: the on-device model runtime, SQLite, the file system. Routing, retrieval, prompt building and source selection are the engine's own code. A fixed sampling seed makes each run reproducible on the same machine: the same commit twice gives a byte-identical report, and two commits give a readable diff.
The gap between the two machines is not in the answers. It is in everything around them. On the Mac mini the 4B model decodes at about 21 tokens per second and a 72-question run takes about ten minutes. On the tester's mid-range Android with 8 GB of RAM, less than the 12 GB the bounty allows, the same model decoded at 2 to 3 tokens per second, and the first complete 73-question run took 53 minutes, with typical answers taking under a minute. In the first smoke test, on an earlier build, single answers took up to two and a half minutes.
We checked the answers early on two phones of our own, and again when that tester ran our frozen measurement build. Our measurement app reported the same engine tree hash as the desktop, the routing tier matched on every item, and the instant and card answers were identical text. On answers, the desk agreed with the phone. It could not tell us speed, and it hid one bug, which section 8 describes. Speed is now a front of its own.

4. The phone fights back
A phone gets hot, runs short of memory and runs on a battery. Each of these changed how we test.
Heat. In our first 52-question sessions on an iPhone with 4 GB of RAM, decode speed fell to about 0.6 of its starting value between the first five answers and the last five. Heat is the likely cause. Our first answer was low-tech: for the early sessions the phone sat on a bag of ice while the benchmark ran. Those runs did not record temperature, which is exactly why every run now does. Today our test script samples thermal status and battery temperature every five seconds and waits between runs until the phone has cooled back down. In the tester's final 73-question run, thermal status stayed at zero throughout.
Memory. On one of our own phones with 8 GB of RAM, an early debug build had the 4B model killed by the operating system while it was loading, even though the engine's own memory check said it fit. The cause was a detail of the inference library: on ARM it repacks about 2.3 GB of weights into a second copy in memory, and our estimate counted only the model's working buffers. We wrote a fix that counts the repack and keeps the weights file-backed when memory is short. It is not yet tested on a phone and not in our frozen build. The tester's 8 GB phone loaded the 4B without it by leaning on swap, which shows how close to the edge 8 GB is. A desktop with memory to spare had no reason to fail.
Battery and time. A phone benchmark runs for an hour on a device that also wants to sleep, dim the screen and charge. Our test script records battery level and temperature during every run, refuses to start unless airplane mode is on and the app has no network permission, and packs the results into one zip that the tester sends back.
The desktop finds what is wrong with the answers, and the phone finds what is wrong with everything else.

5. The funnel: kill ideas while they are cheap
Every idea enters a funnel with four levels, and each level is more expensive than the last:
- L0 takes seconds: unit tests and determinism checks. Did the change do what it says, and only that?
- L1 is a small English slice of 24 items through the full engine. Direction and side effects.
- L2 is the development benchmark with a frozen pass/fail criterion. This is where the judge comes in.
- L3 is the held-out set. It runs only for a ship candidate, because every exposure of a held-out set costs some of its value.
Each idea gets a two-hour time box. If it has not shown a signal by then, it is killed. Every kill is logged with its number.

6. Pre-registration, or how to stop fooling yourself
The easiest way to get a good number from a benchmark is to decide what counts as good after you have seen the result. We built the loop to make that hard by default.
Before any run that decides something, the criterion is written down and hashed. "Pass if the paired difference on these items is above this margin, with no new unsafe answers and no lost sources." The sha256 of that file goes into the log before the run starts. The result is read against the frozen criterion. When we rule on an ambiguous result after the fact, or allow a rescue, the ruling is labelled post-hoc in the log and reported as such.
The same discipline applies to data. The held-out and blind sets are never read by the agents doing the research. The coordinator holds them, the runner touches them only through scripts, and every exposure is logged with a date. When one of our engineering agents ran a blind lookup set during tuning, against its instructions, that set was marked as burned for that line of work, and the exposure went into the log.
In the first six days we froze over forty pass/fail criteria. Many ideas died against them, and each death is in the log with its number.

7. Noise is the enemy you measure first
Our first overnight runs taught us the lesson that shaped everything after: the same configuration, run twice, moved the quality ratio by about 0.03. That was as large as most of the effects we were hoping to detect.
So every comparison is now paired on the same items, scored by the same judge version, with a bootstrap interval on the difference. Our criteria state the smallest effect they can detect. Reusing judge verdicts for answers that did not change saves money and removes one source of noise. And when two runs that should be identical disagree by a few thousandths, we reconcile them before trusting either.
This is what lets us say an idea did nothing, rather than guess.

8. The graveyard
Ideas that looked obvious and did not survive their criterion:
- Bigger dense models. Asked the same questions without retrieval, a 9B model scored no better than our 4B default, gave more confidently wrong answers, and took twice as long per token. Up to 9B, size was not the bottleneck; finding the right passage was.
- Bigger knowledge packs. Adding the 100,000 most-read Wikipedia articles did not improve answers and made retrieval about twice as slow.
- A rewritten retrieval stage. It put the gold passage into the context far more often and still lowered quality. A small model does not automatically use a better context.
- Answer-style prompts. No variant helped, and every one raised the number of confidently wrong answers: each extra sentence the model was asked for got filled from memory.
- Speculative decoding. A small draft model accepted about a third of its guesses; the projected gain on the phone was erased by the draft's own cost.
- A "prompt diet" with a hard answer cap. Shorter prompts, but answers cut mid-sentence and no decode saving.
And some ideas that did survive: a route gate that keeps the model from answering from memory when a question needs sources; short first-aid cards, quoted from official guidance, for health and emergency questions (AI-reviewed, not clinically reviewed); a learned emergency detector that also reads earlier turns of a conversation; and a fix for an instant-answer path that confidently answered the wrong question. We also fixed a caching bug in our own code that only a real phone could reveal, because the desktop stand-in silently skipped that step. Its speed gain is estimated, not yet measured on a phone.
These changes live in our research build for now; they go to the app once they are confirmed on phones.

9. Who runs this
Most of this loop is run by AI agents, each with a lane: retrieval and knowledge, models and inference, evaluation and safety, native code and the phone measurement app, and a runner that schedules jobs. A coordinator agent owns the queue and the frozen criteria. Wording and content questions, including the first-aid card texts, go to a separate reviewer agent that did not write them; those cards are AI-reviewed, not clinically reviewed. People set the direction, approve each frozen build, run the real phones, and decide what goes public.
Every experiment, hash, kill and blind-set exposure goes into one append-only log, and every number we publish traces back to it.
10. What's next
Two fronts are open. Speed on real phones: the first device runs gave the same answers as the desktop, but slowly, and the latency budget is now measured and attacked piece by piece. And confirmation: a frozen build of the engine, a measurement app that cannot open a network socket by construction, and runs on real phones in airplane mode.
When those numbers are in, we will publish them with their confidence intervals. Until then, the honest answer to "is BOAR more than half as good as the internet?" is: we built the instrument to find out, and we are measuring.
The BOAR app is open source on GitHub.
BOAR: AI that works when the internet doesn't.