Lessons from building a 9,010-video sign language dataset
Work in progress. This is a draft outline — the full post is being written.
Papers make dataset building sound like a procedure. It’s logistics: scheduling humans, fighting storage, and discovering that your labelling convention from week one fails on a sign you meet in week six. We collected 9,010 videos of continuous Bangla Sign Language across 530 word classes; the model that consumes them took a fraction of the effort the dataset did. This post is the honest accounting.
Outline of the full post:
- Picking 530 words — frequency lists vs. communicative coverage, and who should make that call (hint: signers, not engineers).
- Continuous vs. isolated recording — why we chose the harder option, and what it cost in annotation time.
- The annotation pipeline: tooling, label alignment, and the QC pass that caught the most errors.
- Storage, backup, and the boring infrastructure that saved the project twice.
- What I’d redo: things I’d standardise on day one, and the metadata I wish we’d captured.