Tag: Featured

From Splits to Heat Strain: A Ten-Year Triathlon Study Built End-to-End with AI

Regular readers will know I have spent the last couple of months playing with AI agents to scrape and visualise publicly available sport data — first football muscle injuries, then a live dashboard for the Giro d’Italia. Those were fun, low-stakes experiments to learn what the tools could do. This post is the point where the experiment turned into something more serious: a full research study, now posted as a preprint, built almost entirely with the same family of AI tools.

I planned, conducted and executed this study — and developed the accompanying digital twin — using Claude Sonnet 5 and Claude Opus, together with Claude Cowork, Claude Code and Claude Design. Cowork orchestrated the data gathering and organisation across ten years of race results and weather records, Code built the statistical models and the digital twin engine, and Design shaped how the outputs are presented. I stayed firmly in the loop throughout as the domain expert: setting the research questions, checking the physiology and the methodology at every step, validating the data, and deciding what the numbers actually meant. I also modified some of the coding as things progressed to develop various analytics steps.

Why triathlon, why heat

Triathlon is an Olympic sport, and elite Olympic-distance racing is shaped by the interplay of swim-bike-run pacing, transition efficiency, the quality of the field, and — increasingly — the heat athletes race in. Major championships are more and more often held in hot conditions, which matters enormously for how athletes and support staff plan training, pacing, cooling and acclimatization. Despite there being a decade of publicly available race results and weather records out there, nobody had linked the two together at scale. That gap was the starting point.

The study had four aims: characterise the performance signature of a podium finish; reconstruct the thermal environment of championship venues over ten years; identify which athletes seem resilient to heat; and, building on all of that, prototype a digital twin that could support race planning.

Ten years of racing, in numbers

A few things stood out. Bike and run splits each contribute roughly equal unique variance to total race time, but the run leg is the real discriminator between the podium and the rest — it was the single most important feature in the podium-prediction models, at 45.5% importance for men and 48.9% for women. Somewhat counter-intuitively, the slope linking heat to performance did not reach statistical significance for either sex (men p=0.065, women p=0.104), which is a useful reminder not to over-interpret heat effects from headline temperature alone. DNF rates, on the other hand, told a clearer story, ranging from 12.7–16.3% for men and 11.7–18.8% for women across World Triathlon’s Green and Blue flag heat-risk categories.

Building the digital twin

The last aim — and the part I am most excited about — was turning ten years of descriptive analysis into something forward-looking. The digital twin prototype couples a performance model with a thermo-physiological heat-strain model, so it can be used to explore pacing, cooling and acclimatization decisions ahead of a race rather than just explaining results after the fact. On a temporal hold-out (training on earlier years, testing on later ones — the fairest test for something meant to inform future decisions), it achieved a mean absolute error of 0.490 z-units for men and 0.468 for women (r=0.383 and 0.295 respectively). Those are honest, prototype-grade numbers, not a finished predictive tool, and I say so explicitly in the paper.

The usual health warning

This is a preprint. It has not been peer reviewed, and I want to be upfront about that rather than let the AI-workflow angle overshadow it. The underlying data are scraped from publicly available sources, so they carry all the usual caveats about completeness and accuracy that come with that. What I can say is that the statistical approach, the modelling choices and the interpretation were all reviewed and directed by me at every stage — the AI tools accelerated the mechanics of gathering, structuring, analysing and building, but the scientific judgement was, and had to be, human. Critical thinking stays firmly the job of the person in the loop, whatever is doing the typing.

Why this matters to me

I hope this prototype is the beginning of something bigger rather than a one-off curiosity. There is an enormous amount of publicly available data across sport that nobody has the time to properly interrogate and organise — results archives, weather records, GPS feeds, injury registries. What changed for me this year is that a single person, using these AI tools well, can now plan, run and analyse a study of this scale in weeks rather than requiring a full research team and many months. I want to keep using that capability to ask more questions like this one across sports science and sports medicine, and to be transparent about the process as I develop more tools and research questions.

If you want to dig into the full methods, results and figures, the preprint is here: From Splits to Heat Strain: A Ten-Year Analysis of Performance Determinants and a Digital Twin Prototype for Olympic-Distance Triathlon Developed Using a Generative AI Workflow (DOI: 10.21203/rs.3.rs-10216961/v1). As always, comments and critique are very welcome — that is rather the point of putting it out as a preprint while it goes through formal peer review.

New paper: Physical Predictors of Skeleton Performance

This week I had another paper published. This paper was part of the PhD studentship of Dr. Steffi Colyer in partnership with Bath University, GB Skeleton, UK Sport and my previous role at the BOA.

In this work we looked at the testing battery for strength and power assessment of bob skeleton athletes and identified predictors of skeleton performance. The analysis approach revealed that 3 tests scores can obtain a valid and stable prediction of bob skeleton start performance. More work from Dr Colyer’s excellent PhD will be published soon, so follow her work as I am sure more applied approaches in other sports will be followed in the next years. I enjoyed working with a great group of colleagues, athletes and coaches for this project and the publication reminded me of how fortunate I was in my time in the UK.

This project is a good example of how some applied sports science projects can advance understanding of specific performance issues as well as provide meaningful advice for the coaches and practitioners involved in this particular sport.
The abstracts is below:
Int J Sports Physiol Perform. 2016 May 1. [Epub ahead of print]

Physical Predictors of Elite Skeleton Start Performance.

Abstract

PURPOSE:

An extensive battery of physical tests is typically employed to evaluate athletic status and/or development often resulting in a multitude of output variables. We aimed to identify independent physical predictors of elite skeleton start performance overcoming the general problem of practitioners employing multiple tests with little knowledge of their predictive utility.

METHODS:

Multiple two-day testing sessions were undertaken by 13 high-level skeleton athletes across a 24-week training season and consisted of flexibility, dry-land push-track, sprint, countermovement jump and leg press tests. To reduce the large number of output variables to independent factors, principal component analysis was conducted. The variable most strongly correlated to each component was entered into a stepwise multiple regression analysis and K-fold validation assessed model stability.

RESULTS:

Principal component analysis revealed three components underlying the physical variables, which represented sprint ability, lower limb power and strength-power characteristics. Three variables, which represented these components (unresisted 15-m sprint time, 0-kg jump height and leg press force at peak power, respectively), significantly contributed (P < 0.01) to the prediction (R2 = 0.86, 1.52% standard error of estimate) of start performance (15-m sled velocity). Finally, the K-fold validation revealed the model to be stable (predicted vs. actual R2 = 0.77; 1.97% standard error of estimate).

CONCLUSIONS:

Only three physical test scores were needed to obtain a valid and stable prediction of skeleton start ability. This method of isolating independent physical variables underlying performance could improve the validity and efficiency of athlete monitoring potentially benefitting sports scientists, coaches and athletes alike.
PMID:
27140284
[PubMed – as supplied by publisher]

#TrainingLoad16 Videos Now Online

All the videos of the #TrainingLoad16 conference organised in Doha by Aspire Academy’s department of Sports Science are now online and free to access for everyone here.
The conference was a great success and showed how much work has been done in this field as well as how much we need to do in order to provide more and better tools to improve our decision making when planning training activities in different sports.

Below is the video of my talk.

https://player.vimeo.com/video/160709539 Dr. Marco Cardinale (QAT) – Monitoring Athlete Training Loads – The Hows and Whys from Aspire Academy on Vimeo.