Field Guide FG-02 / Pre-arrival profile / Rev 2026.10.06

TIINY AI
POCKET LAB

A computer around power-bank size that promises serious offline inference. The idea is unusually prosumer-friendly. The evidence is not ours yet. So this page separates what Tiiny says, what follows from the architecture, and what the parcel still has to prove.

Pre-arrival, not review~300 g claim80 GB LPDDR5X claimdNPU ecosystem is the bet
TIINY POCKET LAB  ✚  300 GRAMS OF AMBITION  ✚  FITS ≠ USABLE  ✚  THE CONVERTER IS THE DOOR  ✚  TEST WHEN IT ARRIVES  ✚  TIINY POCKET LAB  ✚  300 GRAMS OF AMBITION  ✚  FITS ≠ USABLE  ✚  THE CONVERTER IS THE DOOR  ✚  TEST WHEN IT ARRIVES  ✚ 
SEC.01

The hardware, translated.

Every attractive sentence gets an evidence label

Spec / claimEvidenceProsumer translation
14.2 × 8 × 2.53 cm · ~300 gVendor→If the shipping unit matches the claim, this is genuinely jacket-scale compute — the jacket test, finally. The full kit still includes power and cables.
ARMv9.2 · custom SoCVendor→Efficiency first. Also: no CUDA comfort blanket. The software ecosystem matters as much as the silicon.
dNPU · ~190 TOPSVendor→Potentially excellent watts-per-token — but only for models the vendor toolchain can convert. You are buying the converter pipeline too.
80 GB LPDDR5X · 1 TB NVMeVendor→Huge capacity for the size. The important unanswered question is bandwidth and how the CPU/NPU memory split behaves in real workflows.
~30 W AI workload / ~65 W system ceilingVendor→Exactly the kind of power envelope an always-on private companion wants. Wall-meter verification comes after delivery.
“Runs up to 120B”Vendor→Fits is only question one. Question two: which 120B, what quantisation, what context, and how many tokens per second?
SEC.02

Why a small box at all?

The doctrine is compelling even before the benchmark

The always-on one

Chat, note search, voice and quick questions are all-day, low-intensity jobs. A small efficient node can own the interactive day while a bigger GPU machine sleeps.

Offline by design

The pitch is private, local inference with encrypted storage and no token meter. That is exactly the right promise for this category — and exactly what we should verify packet by packet.

Swarm, not monolith

One device present all day, another waking for heavy work. Different boxes doing the jobs that fit their shape instead of one giant machine doing everything badly.

Presence may be the feature.The machine that is simply there, quiet and awake, may get used more than the magnificent one that needs ceremony.
SEC.03

Claim map, not benchmark yet.

Pre-arrival classification; bars are not measurements

≤14B LLMs
Inferred
Should be easy
~30B class
Vendor
Claimed interactive
70B class
Unverified
Need exact runs
120B MoE
Vendor
Headline claim
Image / diffusion
Unverified
Format support first
Training / finetune
Inferred
Not the job
The most important benchmark is currently missing: ours.That is fine. “Not tested yet” is a valid result.
SEC.04

Questions the parcel has to answer.

Fine print before the pledge button

01

The converter is the door

If a great new open model arrives and the dNPU converter does not, excellent silicon can still be a generation behind. We need to measure converter cadence as part of product maturity.

02

80 GB is not one sentence

Capacity, bandwidth and CPU/NPU partitioning are different facts. The brochure gives the cup; we still need to inspect the straw.

03

120B: which one, exactly?

Model, quant, context, throughput, first-token latency, power and sustained behaviour. One headline number cannot answer all seven.

04

What survives without the vendor?

APIs, local files and weights should remain useful if the company, desktop app or conversion service changes. Long-term autonomy is a feature, not an apocalypse fantasy.

SEC.05

Eight-test plan

No pass badges before our units arrive

TestPre-arrival callEvidenceWhat we will actually do
First resultWaitVendorTime unboxing → first useful answer on our own PDF. Count every install, login and hidden step.
ReadingWaitVendorRun exact 30B/70B/120B models; capture first-token latency, sustained tok/s and context.
LibraryWaitUnverifiedMeasure dB(A), surface temperature and sustained speed after twenty minutes.
BackpackLikelyInferredWeigh the complete useful kit, not just the 300 g chassis; count charger and cables.
OfflineWaitVendorDisconnect WAN and inspect outbound packets, activation checks and local workflow completeness.
Meeting roomWaitInferredRun a private document workflow from another screen with the internet unavailable.
ElectricityWaitVendorWall-meter idle, interactive and sustained loads; cost the 24/7 case.
Next doorWaitUnverifiedCheck API portability, model-converter cadence, Olares/network interoperability and vendor independence.
300 grams of private AI is a strong sentence. It is not yet our result.

That is why this page is a pre-arrival Field Guide rather than a review. When the units arrive, every attractive claim above either graduates to OBSERVED / MEASURED, gets qualified, or fails. All three outcomes are useful.

← ManifestoThe user-side definition this guide is scored against.Olares One →The big lab sibling: lived with, heavier, more mature, a very different job.