CARE v1.0: A large multimodal video dataset of speech and non‑verbal behaviour across 12 medical conditions
Researchers present CARE v1.0, a curated public dataset of short interview videos designed to help study how speech and non‑verbal behaviour change with illness. The corpus contains about 143.5 hours of video made from 4,281 clips. These clips come from 612 people and cover 12 medical conditions plus a control group.
The material was drawn from the Health Experience Insights (HEXI) archive of filmed patient and caregiver interviews. The authors selected profiles that had usable video and written narratives. After filtering for corrupted files, privacy restrictions, and other issues, the corpus includes 622 profiles representing 612 unique individuals. Patient conditions include asthma, chronic pain, cleft lip and palate, COVID‑19, depression, epilepsy, fibromyalgia, lung cancer, motor neuron disease, Parkinson’s disease, psychosis, and stroke. Controls are often caregivers, bereaved relatives, or health workers.
Each video clip comes with many pre‑computed descriptors that capture both speech and visible behaviour. These include measures of speech production, facial activity, gaze patterns, and upper‑body movement. The dataset also contains structured metadata extracted from participants’ written narratives and demographics. That metadata lists things such as medication references, life impacts, and expressed emotions. The authors extracted those fields automatically using an open‑source large language model (gpt‑oss‑120b) run with a deterministic setting and then converted the outputs into CSV files.
This resource matters because earlier public datasets were often small, focused on a single disease, or audio‑only. CARE is broader and multimodal, so it can support research on automatic disease and symptom detection from both speech and non‑verbal cues. It can also be used to study how people speak and move when talking about emotionally charged experiences, and to explore disease trajectories and coping over time.