Qwen3.8-Max Visual Grounding Tested on 27,083 Real Logos
Blog post from Voxel51
An evaluation of Qwen3.8-Max using FiftyOne and the QMUL-OpenLogo dataset of 27,083 images and 352 logo classes found that its visual-grounding performance is highly variable across bounding boxes, keypoints, polygon outlines, rotated boxes, and classification. Although the model achieved strong box IoU scores of up to 0.94 on some images and could produce plausible rotated quadrilaterals for tilted objects, it frequently changed coordinate conventions among normalized values, pixel coordinates, and a 0–1000 grid despite explicit prompts, causing major localization errors. Results also varied substantially with thinking mode, image-detail settings, prompt wording, and sometimes identical inputs, including one repeated Red Bull test that changed from three correct matches to none. Polygon tracing was qualitatively capable but costly in reasoning tokens and lacked mask ground truth, while detection, keypoint, and polygon calls sometimes disagreed on the number of logo instances, partly because OpenLogo annotations omit real secondary occurrences. Untargeted logo classification identified several genuine background and sponsor brands but also misread a watermark as a brand with high confidence. In contrast, ordinary visual question answering remained stable across detail settings, suggesting that the observed instability is concentrated in tasks requiring precise spatial coordinates.
| Trend | Post Mentions | Total Month Mentions | Posts | Companies | MoM |
|---|---|---|---|---|---|
| Observability | 8 | 625 | 152 | 84 | -84% |
| AI Guardrails | 2 | 96 | 30 | 18 | -81% |
Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.