Work in progress. This is a draft outline — the full post is being written.

Papers make dataset building sound like a procedure. It’s logistics: scheduling humans, fighting storage, and discovering that your labelling convention from week one fails on a sign you meet in week six. We collected 9,010 videos of continuous Bangla Sign Language across 530 word classes; the model that consumes them took a fraction of the effort the dataset did. This post is the honest accounting.

Outline of the full post:

  1. Picking 530 words — frequency lists vs. communicative coverage, and who should make that call (hint: signers, not engineers).
  2. Continuous vs. isolated recording — why we chose the harder option, and what it cost in annotation time.
  3. The annotation pipeline: tooling, label alignment, and the QC pass that caught the most errors.
  4. Storage, backup, and the boring infrastructure that saved the project twice.
  5. What I’d redo: things I’d standardise on day one, and the metadata I wish we’d captured.