Skip to content
AI Atlas
PaperActive

Can Edge-Deployable Vision-Language Models Identify Species?

arxiv.org/abs/2609.11916

quality89

Updated 56 min ago · first seen 12 Sept 2026

paper_01M29X34R3W5QQ2BMB23BNE1V4

Published
12 Sept 2026
T1 · 56 min ago
arXiv
2609.11916
T1 · 56 min ago
Category
cs.AI
T1 · 56 min ago

Abstract

Camera traps often run in the field on edge hardware with limited or no connectivity, making small, locally-deployable vision-language models (VLMs) -- not frontier-scale ones -- the practically relevant class to evaluate for species identification. We test whether models in this deployment-relevant 2--8B range carry genuine taxonomic knowledge, evaluating four such VLMs (Qwen3-VL 2B/4B/8B, Gemma3 4B) against the domain-specific specialist BioCLIP (300M parameters) on a 96-species task, comparing clean iNaturalist photographs against camera-trap imagery from 6 LILA.science collections, on two independently-sampled evaluation sets. All models identify species far above chance, but every model -- general-purpose or specialist -- degrades sharply on field imagery (domain gaps of 9.6--26.6 percentage points, consistent across taxonomic levels and both evaluation sets), indicating the degradation reflects general image legibility rather than fine-grained discrimination failure. BioCLIP substantially outperforms every VLM tested (by 33.2--59.2 percentage points across an expanded 200-image sample for every model) despite its far smaller size, suggesting the gap reflects specialized training data rather than model scale; yet BioCLIP's own domain gap (18.0 points) is statistically indistinguishable from the best VLM's (22.3 points), suggesting the clean-to-field degradation itself is a property of the image-quality shift rather than a general-purpose-model weakness. Under open-set prompting, 5.9--9.6% of responses are syntactically valid but taxonomically nonexistent species names; the relative fabrication-rate ranking across models replicates exactly across both evaluation sets, a more robust finding than any single point estimate.

Authors 5

Mayukha Siripuram, William Zhou, Xiao Yan, Yi Ding, Ziqi Liu

Specification

Official page

Source:arXiv (Atom API + RSS)T1observed 56 min agohigh

Arxiv announce type
new

Source:arXiv (Atom API + RSS)T1observed 56 min agohigh

arXiv id
2609.11916

Source:arXiv (Atom API + RSS)T1observed 56 min agohigh

Categories
cs.AI

Source:arXiv (Atom API + RSS)T1observed 56 min agohigh

PDF

Source:arXiv (Atom API + RSS)T1observed 56 min agohigh

Primary category
cs.AI

Source:arXiv (Atom API + RSS)T1observed 56 min agohigh

Published
12 Sept 2026

Source:arXiv (Atom API + RSS)T1observed 56 min agohigh

Each value shows its source, tier and observation time. Conflicting claims are kept side by side and flagged — never averaged. How AI Atlas records facts →

Provenance

Attributed facts

9

Source tiers

T19

Freshest observation

56 min ago

Conflicts

None