Accelerating researchers and developers building multilingual AI with a new open dataset
Blog post from GitHub
GitHub has released the GitHub Multilingual Repositories Dataset, a metadata collection aimed at helping researchers and developers discover public repositories with non-English content, emphasizing the importance of multilingual developer collaboration as AI becomes integral to software development. This dataset includes language classifications for READMEs, issues, and pull requests, relying on classifiers like fastText, gcld3, and lingua-py to provide confidence scores, and encompasses over 80 million classification rows across more than 40 million repositories. While Portuguese is the most common non-English language in READMEs, Korean is prevalent in issue texts, highlighting language distribution variances. Designed as a transparent discovery tool, it allows users to determine precision and recall tradeoffs for their research, offering insights into non-English developer communities and supporting more inclusive AI tools. Released under CC0-1.0, the dataset aligns with Microsoft’s European Digital Commitments to enhance multilingual data accessibility and encourages contributions from researchers and developers to extend and critique the data for broader applications.
| Trend | Post Mentions | Total Month Mentions | Posts | Companies | MoM |
|---|---|---|---|---|---|
| AI Coding Assistant | 1 | 2,161 | 541 | 167 | +20% |
Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.