SportsFirst

Plain-English query tool over combined training load and wellness data

Data & reportingPlatform module3 month first phaseAnalyse data

Problem

When a head coach asks how many sessions a group missed to load management this month, or wants total high-speed running across the last training block, the answer requires three separate exports. Someone logs into the GPS platform for a CSV, into the wellness survey tool for another, and checks the medical status spreadsheet the medical team keeps by hand. The files use different athlete name spellings and different date formats, so joining them takes an afternoon and the totals rarely match what was reported the week before. This usually lands on a Friday, ahead of a Monday selection meeting, and whoever does it is guessing at some of the joins by the third page. The organisation already holds the data. Nobody has an afternoon spare to reconcile it every time a question arrives.

Product idea

A data model that pulls in periodic exports from the GPS platform, the wellness survey tool, testing hardware and the medical status log, and matches athletes across them on a single roster ID rather than name text. On top sits a plain-English query box: "which academy players logged a load spike in the last two weeks without a matching wellness dip", "total sessions missed to injury this term by position". Answers come back as a table or chart with the underlying rows shown, so a sceptical reader can check the join rather than trust the summary. Any query can be saved and re-run, or scheduled as a weekly digest to a named inbox. It does not raise alerts and does not predict injury risk — it answers the question that was asked, on data already collected.

Who it is for

Head of Performance and sports scientists building reports for selection meetings and season reviews, performance analysts fielding ad hoc questions, and the head coach who asks them on a Friday afternoon.

Possible first version

A unified data model built from weekly manual CSV uploads of GPS session summaries, wellness survey scores and a medical status export, with athletes matched on a shared roster ID maintained by hand. A query box handles a fixed set of question templates — load by period, missed sessions by reason, wellness trend by group — plus free-text questions mapped onto the same model, returned as a table or simple chart with source rows visible. Saved queries and a weekly email digest. Out of scope for version one: live API integration with any vendor platform, predictive modelling, and real-time alerting.

Build classification
Platform module
Rough effort
3 month first phase
Roles involved
Head of Performance, Sports scientist, Performance analyst, Head coach
Relevant to
Professional club, Collegiate athletics, Academy & youth, Federation / governing body
Systems in play
GPS and wearable tracking platforms, Wellness survey tools, Medical and injury records, Force plate and testing hardware
Product framing
Analyse data

Questions we get asked

How much historical data do we need to load before this is useful?

Enough to answer the question someone actually asks, which in practice means one full training block, roughly four to six weeks, uploaded from each source. Less than that and trend questions such as load progression have nothing to compare against. The first upload is also where roster ID mismatches between systems surface — spelling differences, retired athletes, mid-season transfers — so budget time to clean that mapping once rather than discovering it query by query.

Does this replace the GPS platform or wellness survey tool?

No. Those stay the system of record for capturing the data in the first place — this only reads what they already hold, on a periodic export, and joins it with the other two sources. If a number looks wrong the fix happens in the source system, not here. Think of it as the layer that answers the question your coach asked, not a replacement for any tool already paid for.

Our athlete names never match cleanly between systems anyway — won't that break this?

That mismatch is the main reason this kind of question currently takes an afternoon, so it is treated as the core problem rather than an edge case. Version one uses a manually maintained roster ID map rather than trying to auto-match names, which is more work up front and far more trustworthy than a fuzzy match that silently drops or merges the wrong athlete.

Why doesn't it alert us automatically when something looks wrong?

Because that is a different product with a different failure mode. An alerting tool has to be right in real time or staff stop trusting it. This is a retrospective analysis tool: it is opened when someone has a question, not watched continuously, and it is judged on whether the answer to that question is correct and explainable, not on how fast it notices something.

Is this your workflow?

Tell us one sports workflow that still runs on paper, spreadsheets, WhatsApp or an outdated system. We will map it and show you what a simpler product looks like.

Tell us about it

More in Athlete performance & sports science