Beyond the Aggregate Score: Per-Country Domain Shift in the GWHD Wheat Head Detection Model Zoo
Blog post from Hugging Face
A follow-up evaluation of nine YOLO and RF-DETR wheat-head detection models fine-tuned on the Global Wheat Head Dataset examines performance separately across images from Australia, China, Japan, Mexico, Sudan, and the United States, revealing domain differences obscured by aggregate mAP scores. Using the same evaluation pipelines as the original model release and a newly assembled country and growth-stage metadata manifest, the analysis found that China was every model’s strongest-performing subset, while the countries where models performed worst differed substantially by architecture: most YOLO variants struggled most on the heavily represented US subset, whereas all RF-DETR variants performed worst in Australia. YOLOv26m and YOLOv11x showed the most consistent country-level performance, while RF-DETR Nano had the widest variation, indicating that deployment region could affect its results more than model selection within the benchmark. Because the US accounts for nearly 44% of test images, aggregate rankings may disproportionately reflect performance in a region where several models are weakest. The analysis cautions that results for small Sudan and Japan subsets are less stable, does not include genotype-level comparisons due to unavailable metadata, and adds per-country tables and raw evaluation files to each model card, with growth-stage analysis planned next.
No tracked trend matches for this post yet.
Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.