Home / Companies / Hugging Face / Blog / Post Details
Content Deep Dive

So, let’s make our own dataset

Blog post from Hugging Face

Post Details
Company
Date Published
Author
tegridydev
Word Count
5,936
Company Posts That Month
82
Language
-
Hacker News Points
-
Post removed?
No
Summary

A practical tutorial explains how to build a small local dataset of public Hugging Face dataset-repository metadata using a standalone Python 3.10+ script and the Hub API, without extra packages or authentication. The script requests up to 1,000 repositories ranked by downloads over the previous 30 days, preserves the unmodified API response and collection details in a RAWDATA folder, then produces processed JSONL records containing repository IDs, URLs, downloads, likes, tags, extracted languages, task categories, formats, citations, revisions, and timestamps. It also creates CSV count reports, owner aggregates, a summary, terminal insights, and a README documenting fields, methods, and limitations. The guide emphasizes validation, retry and rate-limit handling, safe file writing, preservation of missing values rather than converting them to zero, and the distinction between API-reported tags and verified repository contents. It cautions that the collection is a top-downloads sample rather than a representation of the entire Hub, that download counts are rolling 30-day totals rather than unique users or quality measures, and that missing metadata does not prove information is absent elsewhere. Suggested extensions include building dataset finders or dashboards, comparing snapshots by repository ID, and enriching selected records with independently verified DOI or citation information.

Trends Found in this Post
Trend Post Mentions Total Month Mentions Posts Companies MoM
Serverless 4 156 54 28 -80%
LLM 1 747 162 79 -85%
Use This Data

Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.