Tether QVAC open-sources VisionPsy-Nano — company claims top sub-0.5B on-device scores

Another sub-half-billion vision-language model just hit Apache 2.0 with a phone-first pitch. Tether Data’s QVAC initiative open-sourced VisionPsy-Nano, a ~460M-parameter VLM aimed at on-device and edge runs, and used a PR Newswire release to claim the top normalized score among evaluated sub-0.5B on-device VLMs.
Read that carefully: the “best in class” framing is company-asserted, not an independent leaderboard verdict. The weights and two variants are real. The ranking is Tether/QVAC’s own scorecard until someone else reproduces it.
What VisionPsy-Nano claims on sub-0.5B benches
Per the PR Newswire announcement, Tether says VisionPsy-Nano-460M posted an overall normalized score of 62.3, ahead of Liquid AI’s LFM2.5-VL-450M (59.6), Hugging Face’s SmolVLM2-500M (52.5), and Tether’s own base nanoVLM-460M-8k (54.9). The company says the 460M variant beat those peers on 16 of 17 benchmarks in its suite.
Category claims, also company-asserted:
- Visual perception: +4.6% relative margin over the next-best ~0.5B model in Tether’s comparison
- Reasoning & knowledge: +7.4% relative margin over competing ~0.5B models
- Instruction following: on MM-IFEval (42.3) and POPE (87.9), Tether says VisionPsy-Nano-460M beats larger models including FastVLM-0.5B (759M), Qwen3.5-0.8B (873M), and InternVL3.5-1B (1061M)
Document understanding / OCR (flowcharts, financial reports, infographics) is listed as a core strength. Evaluation configs under VLMEvalKit are included for reproducibility — the right move if Tether wants the numbers treated as more than press copy.
Until independent runs land, treat 62.3 and the “No. 1” language as Tether-asserted via PR Newswire, not as settled field fact.
Quality vs Flash: which variant to ship
Two checkpoints ship day one:
- VisionPsy-Nano-460M — quality-focused reference at ~460M
- VisionPsy-Nano-460M-Flash — latency-tuned; Tether claims ~99% of full quality (61.4 normalized) with far fewer visual tokens
Flash latency claims (again, company-asserted): first-token generation ~19–23× faster than SmolVLM2-500M and the base nanoVLM on Pixel 9, Galaxy S23, and Galaxy S25 Ultra, and up to ~36× faster on iPhone 15, with lower latency than LFM2.5-VL-450M and Qwen3.5-0.8B across those four devices. Demos are linked for Galaxy S25 Ultra, Pixel 9, and iPhone 15.
Practical split: ship 460M when OCR and visual reasoning accuracy dominate the SLA; ship Flash when time-to-first-token on a phone is the product. Do not assume the quality variant is “always better” for UX — on-device VLMs lose users to lag before they lose them to a two-point bench gap.
Apache 2.0 terms that matter for apps
Both variants are open-weight under Apache 2.0. That license allows commercial use, modification, and redistribution with attribution and license notice — the usual app-friendly defaults, without a copyleft infection into your larger binary. Tether’s release text frames the models as intended to support researchers and educational use; Apache 2.0 itself does not restrict commercial apps the way some research-only or non-commercial licenses do. Still read the LICENSE file in the checkpoint repo before you embed weights in a store-bound build.
- Full-precision via Hugging Face Transformers
- Quantized GGUF for phones through
llama.cpp(single-line invocation claimed)
3.
For mobile product teams, GGUF + llama.cpp is the path that matches the Flash pitch.
Where on-device VLMs still fail
A 460M VLM on a phone will not replace a frontier multimodal API for dense documents, adversarial OCR, or long multi-image reasoning. Even Tether’s own narrative is “local-first as a viable pathway,” not parity with data-center models. Expect brittle spatial counting, weak fine print on complex PDFs, and quality cliffs once you leave the marketing demos.



