Open-vocab document classification · benchmark
Zero-shot document-type classification — labels supplied at inference, not baked into a
head. Ranked by doc-type mean macro-F1 across DocLayNet, Forms and Tobacco on the
held-out
document-classification-benchmark.
Nutrient's in-house small models vs. cloud frontier VLMs, all scored by the same
score_classify.py (unweighted macro-F1, argmax over candidate labels).
Reading it: the in-house models match/lead the cloud on the visual tracks (DocLayNet, Forms) at zero API cost, and trail on Tobacco — a read-the-header task where a VLM's text reading wins. A dash (—) = track not run for that model.
Nutrient · commercial Nutrient · open-weight cloud VLM Highlighted rows are Nutrient models. Numbers are provisional pending a clean scoring pass.
Models benchmarked: nutrientdocs/document-classification-v2 · nutrientdocs/document-classification-v1 · google/siglip2-so400m-patch16-512 · google/siglip2-base-patch16-512 · google/siglip-base-patch16-224 · openai/clip-vit-large-patch14-336 · openai/clip-vit-base-patch32 · laion/CLIP-ViT-H-14 · apple/DFN5B-CLIP-ViT-H-14-378 · timm/eva02_large_patch14_clip_336. Cloud VLMs (no public HF page): GPT-4o / 4o-mini / 4.1 / 4.1-mini · Gemini 2.5 Pro / Flash / Flash-Lite · Claude Opus / Sonnet / Haiku.