Your Wearable Does Not Measure Your Mind

From sensor signals to individual performance tools in elite tennis

Elite tennis has become exceptionally good at producing data. A modern performance team can follow training load, movement, sleep, recovery, heart rate, acceleration, stroke characteristics and match events with a degree of resolution that would have been difficult to imagine two decades ago. Point-by-point data from professional matches add another layer. We can reconstruct the score, the type of error, the stage of the match and what happened on the next point. This has created a natural expectation that mental performance should become equally measurable. Wearables are increasingly discussed in that context, and psychologically loaded terms such as stress, readiness, nervousness, flow or being „in the zone“ are beginning to appear alongside physiological and behavioral data.

The attraction is obvious, but the language can hide an important methodological distinction. A sensor records a physical signal. Depending on the device, that signal may reflect heart activity, temperature, acceleration, rotation or another measurable property of the body or its movement. A psychological construct is something different. If a system labels a pattern of sensor data as stress, readiness or a mental state, an additional inferential step has already taken place. A model, an operational definition and some form of validation connect the measured signal with the psychological label. That can be scientifically useful. It can also be highly predictive. What it cannot do is turn the label back into the thing that was directly measured.

Wearables remain valuable precisely when we are clear about the level at which they operate, particularly when the goal is to support one elite athlete rather than describe athletes in general. The distinction becomes more consequential when classification is treated as explanation. A model may estimate a psychological state accurately and still tell us very little about how that state developed in this athlete, why it mattered in this situation, what the athlete did in response and why a similar situation produced a different outcome on another day. For high-performance sport, that gap is not a philosophical detail. It determines what kind of performance tool we are actually building.

What Tennis Data Can Already Tell Us

Tennis is an unusually strong environment for studying performance under pressure because so much of the competitive sequence is observable. Harris, Vine, Eysenck and Wilson (2021) analyzed 3,552 Grand Slam matches played between 2016 and 2019, covering 658,068 individual points. Their point-by-point analysis examined whether situational pressure and prior errors predicted subsequent errors. Pressure was operationalized from the match situation, including factors such as break points and the stage of the match. Higher situational pressure increased the probability of a performance error, and an error on the preceding point increased the probability of another error.

The study is valuable precisely because it is so detailed, but it also helps clarify the boundary between an observed sequence and a psychological process. The researchers could identify the score, the error and the subsequent point. Their pressure index represented the objective importance of the situation as defined by the model; it was not a continuous reading of each player’s subjective experience. That distinction leaves open a question that matters enormously in applied work: what happened psychologically inside this particular player between the first error and the next point?

Suppose a player double-faults at 4-4, 30-40 in the third set and loses the next two points quickly. The match data can tell us exactly what happened. They cannot, on their own, tell us whether the player interpreted the double fault as a technical mistake, a threat to the match, evidence of choking, a sign of fatigue or nothing more than an error that required immediate tactical adjustment. They do not tell us whether the player became angry, anxious or unusually passive, whether attention shifted to the consequences of losing, whether a between-point routine was abandoned, or whether an interaction with the box changed the next decision. These are not decorative details around the „real“ data. If the question is mental performance, they are part of the process we are trying to explain.

A Revealing Wearable Study: Trying to Detect “The Zone”

Research by Havlucu and colleagues is particularly useful because it approaches the problem directly from tennis. In 2017, the group reported an online survey of 1,567 tennis players through the Turkish Tennis Federation, followed by interviews with 20 professional and international players. One of the most striking findings was the players‘ interest in feedback about mental states from future wearables (Havlucu et al., 2017). The demand therefore did not originate only in technology companies. Players themselves were asking for information that went beyond physical metrics.

The researchers later explored whether wearable technology could provide such information. Their 2022 study, „Toward Detecting the Zone of Elite Tennis Players Through Wearable Technology,“ involved four professional players and two elite coaches. The players wore Apple Watches that provided inertial measurement unit data. Coaches observed the players and labeled periods according to whether they considered the player to be „in the zone“ or „out of the zone.“ A recurrent neural network was then trained to predict those coach labels from the IMU data.

Havlucu et al. explicitly presented the work as a preliminary exploration rather than proof that a commercial wearable could read a player’s mind. The individual-player models achieved an average test accuracy of 84.30 percent, while an aggregated model across players achieved 83.24 percent. Those results suggest that movement information can contain patterns that are useful for predicting coach-labeled psychological states within the studied setting. The more revealing result for our purposes came from the attempt to generalize the model to a player it had not seen before. Average accuracy was 51.02 percent, essentially at chance level. The authors concluded that the models did not generalize to new players and that individual training was needed. They explicitly discussed the highly personalized nature of the zone and psychological states (Havlucu et al., 2022).

There is another methodological point in the study that deserves attention. The target was not an independently measured internal state. The model predicted what expert coaches had labeled as „the zone.“ That does not invalidate the work; expert observation is a legitimate source of information and the authors are transparent about their procedure. It does mean, however, that the result should be described accurately. The Apple Watch measured movement-related signals. The machine-learning system learned to classify a psychological label supplied by coaches. The study therefore demonstrates the possibility of mapping sensor patterns onto an operationalized psychological category within a particular context. It does not establish that the sensor itself measured the player’s subjective psychological state.

For elite sport, the failure to generalize across players may be more interesting than the headline accuracy. It points toward a problem that is often hidden by the word personalization. If a model has to be trained on a particular athlete before it becomes useful, then the relationship between observable signals and psychological meaning is at least partly person-dependent. Yet even a perfectly personalized classifier would still answer a relatively narrow question: given the data available now, which state is this athlete most likely to be in?

For applied performance work, that estimate may be useful, yet it still leaves another set of questions unanswered: why did this athlete arrive in that state, what did the athlete do next, and what can be learned from previous occasions when a similar situation developed differently?

Personalization Is Not the Same as an Individual Performance Tool

This distinction is central to the work we have been developing at Sixpack Mind AI. The term „personalized“ is now used so broadly that it can describe very different systems. A population model can be trained on many athletes and then adjusted for one player. A dashboard can display different recommendations based on a user’s score. A classifier can be fine-tuned on an individual’s sensor data. All of those approaches may deserve to be called personalized. Our interest is different because the single athlete is not the final customization layer. The single athlete is the unit of analysis from the beginning.

We develop individual performance tools for elite sport. The relevant question is not primarily whether an athlete resembles a population profile, nor whether a general model can be made slightly more accurate by adding some of that athlete’s data. We want to know what can be learned from the athlete’s own longitudinal history: which situations repeatedly precede a change, which subjective appraisals matter, which emotional responses occur, what the athlete tries to regulate, what follows behaviorally and when the expected pattern fails to appear.

That last point is important. An individualized system should not simply collect examples that confirm an athlete’s preferred explanation. It should actively preserve contradictory cases. If a player believes that double faults destroy confidence, the useful evidence is not only the matches in which a double fault preceded a collapse. It is also every comparable match in which the player double-faulted and recovered immediately. Those counterexamples force a more precise question. Perhaps the relevant factor is not the double fault itself but the meaning assigned to it at a particular score, an accumulation of previous errors, fatigue, communication, a specific opponent, a change in self-talk or the failure of a regulation routine.

This is where person-specific process knowledge differs from a state label. A label compresses information. Process knowledge tries to preserve the sequence and its conditions. In methodological terms, the difference also reflects a broader problem in behavioral science. Fisher, Medaglia and Jeronimus (2018) showed that relationships observed between people do not necessarily reproduce the relationships that operate within an individual over time. Their argument is not that group research is useless; it is that group-to-individual generalizability should be demonstrated rather than assumed. For high-performance practice, this matters because the person who has to serve at 5-5 in the third set is not an average of 200 athletes.

Population-level sport psychology remains essential. It identifies robust mechanisms, establishes intervention evidence and tells us which variables deserve attention. But an athlete-specific performance system has an additional task: determine whether those relationships actually describe this athlete, under which conditions they appear and where they break down. That requires repeated observations within the same person. It also requires us to treat the athlete’s own history as evidence, not merely as anecdote.

What Tennis Psychology Already Shows

Tennis psychology already provides methods that move in the direction of situation-specific and within-person analysis. Fritsch and colleagues (2022) video-recorded 20 competitive tennis matches and, shortly afterwards, confronted players with concrete situations from their own match. The players reported their self-talk and rated the intensity of emotions and outward emotional reactions. Rather than asking only how tennis players generally feel under pressure, the design connected subjective reports to identifiable moments in real competition.

Their later study examined the relationship between emotional experiences and emotional expressions in 20 tennis players during competition (Fritsch et al., 2024). Again, players reviewed pre-selected points from their own matches. The analyses focused on within-person associations. Emotional experience and observable expression were related, but the authors also emphasized that many experienced emotions were not externally displayed. They therefore warned against relying exclusively on observation and argued for multiple methods, including self-report, observation and physiological measures, to capture different components of emotion. They also suggested in-game assessment of mental states at several points during competition as a direction for future research.

This work is directly relevant to the wearable discussion because it demonstrates why observable behavior cannot simply stand in for subjective experience. At the same time, it shows that self-report should not be treated as an inferior substitute for sensor data. If the scientific question concerns appraisal, emotion, self-talk or regulation, the athlete’s report is part of the phenomenon under study. The challenge is not to eliminate subjectivity; it is to collect subjective information systematically, close enough to the relevant event, and repeatedly enough to distinguish an isolated interpretation from a recurring individual process.

A systematic review by Nijenhuis and colleagues (2024) offers a limited but instructive view of the broader measurement landscape in racket sports. Among 32 multidimensional or longitudinal studies in talent identification and development, 28 assessed physiological characteristics, 21 anthropometric characteristics and 16 technical characteristics, whereas only four included psychological characteristics. The scope of that review is talent research, not professional tennis monitoring, so it would be wrong to generalize the numbers to elite practice as a whole. Still, the distribution illustrates an asymmetry in one relevant research domain: longitudinal measurement has been developed far more extensively for physical and technical characteristics than for psychological ones.

From “What Happened?” to “Does This Explanation Survive the Athlete’s Own History?”

In applied sport psychology, post-match explanations are unavoidable and often useful. A player may say, „After the double fault I was mentally gone.“ A coach may agree immediately because the sequence looked obvious. The statement may be correct. The problem is that a plausible retrospective explanation is not yet the same as a tested person-specific explanation.

An N=1 approach changes what happens next. Instead of treating the statement as the end of the analysis, it becomes a hypothesis that can be compared with the athlete’s own previous cases. How often has a double fault occurred in a comparable score situation? What happened on the following points? Was the reported emotional response similar? Was the athlete physically fatigued? Which thoughts were reported? Was the between-point routine used? Who was involved in the interaction around the event? Did the athlete attempt a regulation strategy? Did it work? Most importantly, where are the exceptions?

Over time, those comparisons can produce a different type of knowledge. The athlete may discover that the double fault itself is a poor predictor of what follows. The more relevant pattern may be that performance deteriorates when a double fault occurs after a specific sequence of missed opportunities and is followed by a particular interpretation, such as „here we go again.“ Another player may show no such pattern at all. A third may become more effective after an error because anger temporarily sharpens commitment. The purpose is not to find a universal psychological law for double faults. It is to identify which process, if any, is stable enough within this athlete to become useful for future performance decisions.

This is what we mean by an N=1 performance study. The athlete is followed longitudinally, and mental, emotional, physical and contextual information is documented repeatedly. Situations are not reduced to isolated questionnaire scores. They are connected to what preceded them, how they were appraised, what the athlete reported feeling or thinking, what regulation was attempted, what behavior followed and how the episode developed. Successful and unsuccessful situations are both necessary because without successful counterexamples the system would simply become a sophisticated archive of failures.

The aim is practical. We want individual performance tools that become more relevant as the athlete’s own evidence base grows. That is a fundamentally different design objective from creating a mental score that can be displayed for millions of users. Mass-market systems need generalization. Our work is optimized for the opposite problem: extracting useful, testable patterns from the history of one elite athlete. The question is not whether the same pattern exists in everyone else. The question is whether it recurs reliably enough in this athlete to influence what the athlete and performance team do next.

This is also why we are cautious with the word „explanation.“ Longitudinal N=1 data do not magically establish causality. Repeated associations can still be confounded, self-reports can be biased, and hypotheses can be wrong. The advantage is more modest and, in our view, more useful: explanations can be confronted with a growing set of within-person observations, including cases that contradict them. Instead of relying on a single memorable defeat, the athlete can ask whether the story told after that defeat is consistent with the rest of the athlete’s own history.

Why Generative AI Changes the Feasibility of This Idea

The methodological idea is not new simply because AI is involved. Intensive longitudinal assessment, N=1 analysis, ecological momentary assessment and idiographic research all have established scientific traditions. The practical difficulty has always been scale. A serious person-specific performance history can quickly contain hundreds or thousands of observations in different formats: spoken reflections, ratings, training and competition contexts, physical status, social interactions, interventions, setbacks, successes and changing goals. Integrating that material manually over months or years is expensive and difficult.

Generative AI may change the economics and technical feasibility of that task. In our concept, AI is not a psychological theory and it is not an independent source of truth about the athlete. Its potential role is to make large amounts of heterogeneous person-specific information retrievable and comparable over time. It can help structure reports, surface previous situations that resemble the current one, organize hypotheses and counterexamples, and make a longitudinal record usable in everyday performance work. Whether a particular pattern is meaningful still depends on the quality of the data, the design of the analysis and, where appropriate, professional interpretation.

We call the underlying methodological framework Operationalized Self-Research with AI. A scientific vision paper describing the approach – „Operationalized Self-Research with AI: A Sport Psychology Vision for Person-Specific Process Knowledge in Elite Sport“ – is currently in the arXiv submission process. It is deliberately a vision paper, not an efficacy study. The same distinction applies to Sixpack Mind AI. We are developing a service around this architecture; we are not claiming that a finished, empirically validated AI system has already demonstrated performance effects.

At the same time, the direction is not speculative in the sense of being newly invented for this article. Our work has developed over nine years around one persistent problem: elite athletes are surrounded by increasingly sophisticated physical and technical data, while the mental side of performance still lacks an equivalent person-specific longitudinal evidence base. Over those years, our working hypothesis has become increasingly specific. If the goal is to improve one athlete, then the performance tool should be built around that athlete’s own process data from the outset, rather than beginning with the assumption that the most useful knowledge is the average pattern across many people.

What Sixpack Mind AI Is Trying to Build

Sixpack Mind AI develops personal AI performance assistants for elite sport. We are not building a mass-market mental wellness application and we are not trying to create another universal readiness score. Our objective is to build individual performance tools whose knowledge base grows from the longitudinal history of a single athlete. The system under development is intended to document self-reported mental, emotional, physical and contextual information over time and to make individual trajectories, recurring relationships, exceptions and potential tipping points available for exploratory analysis.

The reference point is therefore not primarily a norm value or another athlete. It is the athlete’s own history. This does not isolate the athlete from established sport psychology, coaching knowledge or medical expertise. Sixpack Mind AI is explicitly not intended to replace coaches, sport psychologists or medical professionals. The purpose is to create an additional information layer that can support better questions and better-informed decisions within the performance team.

Our initial market focus is the return from injury in elite sport, because this period makes the need for longitudinal individual data particularly visible. Physical rehabilitation, expectations about return, fear of re-injury, pressure from the sporting environment, setbacks, identity and daily fluctuations in confidence can interact over months. A single assessment at the beginning or a conversation after a setback may be useful, but neither can represent the process as it develops. A longitudinal N=1 record is designed to make that process more visible without pretending that every fluctuation has a simple cause.

For tennis, the broader implication is straightforward. Wearables can become more sophisticated, and psychological inference from sensor data may become much better than it is today. That would be valuable. But it would still answer a different question from the one that motivates our work. A system that estimates whether a player is stressed, ready or in the zone is classifying a state. An individual performance tool should also help examine how performance-relevant states arise, how the athlete responds to them, which responses are associated with successful outcomes, and whether the explanation that seems convincing today survives comparison with the athlete’s own history.

The boundary is therefore quite specific: wearable data can be highly useful, while an inferred psychological label still should not be mistaken for a person-specific account of the athlete’s performance process. Elite sport has spent years building longitudinal data systems around the body. Our proposition is that mental performance deserves its own longitudinal, person-specific data logic – not as a copy of physiological monitoring, but as an evidence base built around the experiences and performance history of one athlete.

References

Fisher, A. J., Medaglia, J. D., & Jeronimus, B. F. (2018). Lack of group-to-individual generalizability is a threat to human subjects research. Proceedings of the National Academy of Sciences, 115(27), E6106-E6115. https://doi.org/10.1073/pnas.1711978115

Fritsch, J., Jekauc, D., Elsborg, P., Latinjak, A. T., Reichert, M., & Hatzigeorgiadis, A. (2022). Self-talk and emotions in tennis players during competitive matches. Journal of Applied Sport Psychology, 34, 518-538. https://doi.org/10.1080/10413200.2020.1821406

Fritsch, J., Fiedler, J., Hatzigeorgiadis, A., & Jekauc, D. (2024). Examining the relation between emotional experiences and emotional expressions in competitive tennis matches. Frontiers in Psychology, 14, 1287316. https://doi.org/10.3389/fpsyg.2023.1287316

Harris, D. J., Vine, S. J., Eysenck, M. W., & Wilson, M. R. (2021). Psychological pressure and compounded errors during elite-level tennis. Psychology of Sport and Exercise, 56, 101987. https://doi.org/10.1016/j.psychsport.2021.101987

Havlucu, H., Bostan, I., Coskun, A., & Ozcan, O. (2017). Understanding the Lonesome Tennis Players: Insights for Future Wearables. Proceedings of the 2017 CHI Conference Extended Abstracts on Human Factors in Computing Systems, 1678-1685. https://doi.org/10.1145/3027063.3053102

Havlucu, H., Akgun, B., Eskenazi, T., Coskun, A., & Ozcan, O. (2022). Toward Detecting the Zone of Elite Tennis Players Through Wearable Technology. Frontiers in Sports and Active Living, 4, 939641. https://doi.org/10.3389/fspor.2022.939641

Nijenhuis, S. B., Koopmann, T., Mulder, J., Elferink-Gemser, M. T., & Faber, I. R. (2024). Multidimensional and Longitudinal Approaches in Talent Identification and Development in Racket Sports: A Systematic Review. Sports Medicine – Open, 10, 4. https://doi.org/10.1186/s40798-023-00669-2

Further Information

If you are interested in how we operationalize this N=1 approach in elite sport, or would like further information about the methodology and the development of Sixpack Mind AI, please contact us. We are particularly interested in discussions with elite athletes, coaches, performance teams and research partners who are working on the same problem from different perspectives.

Frank W. E. Stockmann
Head of Mental Data and Performance
Sixpack Mind AI

„`

Operationalized Self-Research with AI:
A Conceptual Position Paper on Person-Specific Empirical Research in Elite Sport

Schreiben Sie einen Kommentar

Ihre E-Mail-Adresse wird nicht veröffentlicht. Erforderliche Felder sind mit * markiert