Milestone map
Milestone map
3 milestones
Collect your own labelled data
1–2 weeks
Record your own dataset from a real sensor — sounds, gestures or images — with at least 3 labels. On-device models live or die by their data, so you collect it yourself.
Proof required
Link the dataset (or a summary with sample counts per label) and a short note on how and where it was recorded.
What gets checked
- At least 3 labels and 50+ samples each
- Recorded by you on the target sensor
- Counts per label reported honestly
Common mistakes
- Downloading a ready dataset
- One label with far more samples than others
- All samples recorded in one setting
Powstik Guide
A small dataset you recorded yourself.
Steps
- Pick a task with 3+ labels.
- Record on the sensor the model will run on.
- Vary where and how you record.
- Count samples per label.
- Hold back a test set before training.
Template
Dataset Label | samples | where recorded Test set held back:
What gets sent back
- A downloaded dataset.
- Fewer than 3 labels.
- No held-back test set.
Resources
Foundationstart here
Depthgo deeper
What a verifier looks for
- Checked by the verifiers you invite to your panel (up to 3). Pick people who can judge the work, not friends.
- Check sample counts per label
- Ask how the data varies (people, places, noise)
You'll sign in first, then come straight back here.
Train and test the model
1–2 weeks
Train a small model and test it on the held-back set. Report accuracy per label with a confusion matrix, plus model size and memory, so it can fit on the device.
Proof required
Submit the confusion matrix on your held-back test set, model size in KB, and the training notebook or project link.
What gets checked
- Tested only on held-back data
- Per-label errors shown, not just one accuracy number
- Size and memory fit the target device
Common mistakes
- Testing on training data
- Reporting one headline accuracy
- A model too big for the device
Powstik Guide
A trained model with honest test results.
Steps
- Train a small model.
- Run it on the held-back set only.
- Build the confusion matrix.
- Check size against device memory.
- Note the weakest label.
Template
Results Accuracy (test): Weakest label: Model size (KB): Device RAM/flash:
What gets sent back
- Results on training data.
- No confusion matrix.
- No size check.
Resources
Foundationstart here
Depthgo deeper
What a verifier looks for
- Checked by the verifiers you invite to your panel (up to 3). Pick people who can judge the work, not friends.
- Confirm the test set was not trained on
- Ask which label confuses the model most and why
You'll sign in first, then come straight back here.
Run it live on the device
1–2 weeks
Deploy the model to the microcontroller or single-board computer and run it on live input. Measure how long one prediction takes. Explain it to a reviewer who tests it live.
Proof required
Link a video of live predictions on the device (reviewer supplies at least one input), latency per prediction, and the review log.
What gets checked
- Runs on the device, not a laptop
- Latency measured
- Works on an input the reviewer chose
Common mistakes
- Inference actually runs on a PC
- Only pre-recorded inputs
- No latency number
Powstik Guide
A model predicting live on the device.
Steps
- Export for the device.
- Flash and run on live input.
- Time 20 predictions and average them.
- Let the reviewer try their own input.
- Log what failed and why.
Template
Review session log Date: Duration: Reviewer name: Reviewer role and experience: Challenge 1 (on-device model): What they asked: My answer: Challenge 2: What they asked: My answer: Change I made after the session:
What gets sent back
- Laptop inference.
- Pre-recorded demo only.
- No reviewer input.
Resources
Foundationstart here
Depthgo deeper
What a verifier looks for
- Checked by the verifiers you invite to your panel (up to 3). Pick people who can judge the work, not friends.
- Give it one input of your own
- Ask what they would change to cut latency
You'll sign in first, then come straight back here.
Part of