Accessibility Adjustments

Use these optional tools to adjust reading and display preferences. These tools cannot resolve every accessibility barrier. Please contact the website owner if you need assistance.

  • Text adjustments
  • Content scaling 100%
  • Font size 100%
  • Line height 100%
  • Letter spacing 100%
  • Colour adjustments
  • Orientation adjustments

Tether QVAC open-sources VisionPsy-Nano — company claims top sub-0.5B on-device scores

Listen to this article

Another sub-half-billion vision-language model just hit Apache 2.0 with a phone-first pitch. Tether Data’s QVAC initiative open-sourced VisionPsy-Nano, a ~460M-parameter VLM aimed at on-device and edge runs, and used a PR Newswire release to claim the top normalized score among evaluated sub-0.5B on-device VLMs.

Read that carefully: the “best in class” framing is company-asserted, not an independent leaderboard verdict. The weights and two variants are real. The ranking is Tether/QVAC’s own scorecard until someone else reproduces it.

What VisionPsy-Nano claims on sub-0.5B benches

Per the PR Newswire announcement, Tether says VisionPsy-Nano-460M posted an overall normalized score of 62.3, ahead of Liquid AI’s LFM2.5-VL-450M (59.6), Hugging Face’s SmolVLM2-500M (52.5), and Tether’s own base nanoVLM-460M-8k (54.9). The company says the 460M variant beat those peers on 16 of 17 benchmarks in its suite.

Category claims, also company-asserted:

  • Visual perception: +4.6% relative margin over the next-best ~0.5B model in Tether’s comparison
  • Reasoning & knowledge: +7.4% relative margin over competing ~0.5B models
  • Instruction following: on MM-IFEval (42.3) and POPE (87.9), Tether says VisionPsy-Nano-460M beats larger models including FastVLM-0.5B (759M), Qwen3.5-0.8B (873M), and InternVL3.5-1B (1061M)

Document understanding / OCR (flowcharts, financial reports, infographics) is listed as a core strength. Evaluation configs under VLMEvalKit are included for reproducibility — the right move if Tether wants the numbers treated as more than press copy.

Until independent runs land, treat 62.3 and the “No. 1” language as Tether-asserted via PR Newswire, not as settled field fact.

Quality vs Flash: which variant to ship

Two checkpoints ship day one:

  • VisionPsy-Nano-460M — quality-focused reference at ~460M
  • VisionPsy-Nano-460M-Flash — latency-tuned; Tether claims ~99% of full quality (61.4 normalized) with far fewer visual tokens

Flash latency claims (again, company-asserted): first-token generation ~19–23× faster than SmolVLM2-500M and the base nanoVLM on Pixel 9, Galaxy S23, and Galaxy S25 Ultra, and up to ~36× faster on iPhone 15, with lower latency than LFM2.5-VL-450M and Qwen3.5-0.8B across those four devices. Demos are linked for Galaxy S25 Ultra, Pixel 9, and iPhone 15.

Practical split: ship 460M when OCR and visual reasoning accuracy dominate the SLA; ship Flash when time-to-first-token on a phone is the product. Do not assume the quality variant is “always better” for UX — on-device VLMs lose users to lag before they lose them to a two-point bench gap.

Apache 2.0 terms that matter for apps

Both variants are open-weight under Apache 2.0. That license allows commercial use, modification, and redistribution with attribution and license notice — the usual app-friendly defaults, without a copyleft infection into your larger binary. Tether’s release text frames the models as intended to support researchers and educational use; Apache 2.0 itself does not restrict commercial apps the way some research-only or non-commercial licenses do. Still read the LICENSE file in the checkpoint repo before you embed weights in a store-bound build.

  1. Full-precision via Hugging Face Transformers
  2. Quantized GGUF for phones through llama.cpp (single-line invocation claimed)
    3.

For mobile product teams, GGUF + llama.cpp is the path that matches the Flash pitch.

Where on-device VLMs still fail

A 460M VLM on a phone will not replace a frontier multimodal API for dense documents, adversarial OCR, or long multi-image reasoning. Even Tether’s own narrative is “local-first as a viable pathway,” not parity with data-center models. Expect brittle spatial counting, weak fine print on complex PDFs, and quality cliffs once you leave the marketing demos.

Maya Chen
Maya Chen

Maya Chen is focused on covering AI models, research, and the evidence behind new capabilities. Maya follows model launches, benchmarks, open weights, and scientific uses of AI with one question in mind. What changed, and how would we know? The voice is curious and exacting, with a soft spot for elegant technical ideas and little patience for a leaderboard without context.

Leave a Reply

Your email address will not be published. Required fields are marked *

Gravatar profile