Hello,
The biggest thing holding back ball and player tracking is training data. I’d like to build a shared dataset from real match footage so tracking gets better for everyone.
You’d optionally contribute footage from your matches, and in return tracking improves for everyone, and down the line, better models on your own games, since the training has fit to your footage.
Many of you have already sent me your games, and some of you authorized me to use them for training. Thank you.
The question I want your opinion on: should this be a public dataset? That way anyone can train on it themselves, labeling is much easier to share (ball + players, eventually field markings), it scales to other sports without me being the bottleneck, and it would make Reco a real reference for the sports community.
This also goes beyond tracking: better data and models are exactly what unlock analytics and highlights, and I don’t think the scale needed is that large. Here’s a video of what I mean: https://www.youtube.com/watch?v=aBVGKoNZQUw
Regarding existing datasets, in my experience, they don’t fit Reco’s constraints: a fixed rig, with the AI running on the raw unstitched footage.
I am concerned about footage with children in it, whether on or around the pitch. There’s also a control trade-off: a public dataset is hard to take back, since anyone can copy it. Kept private, I can delete your footage whenever you ask.
So, public or private, and what would make you comfortable either way?