VISTA-PD captures 14 acoustic biomarkers from a 3-minute voice protocol and aids primary care clinicians in differentiating early Parkinson's disease from Essential Tremor.
Protocol
A clinician-administered voice protocol captures the acoustic signatures associated with hypokinetic dysphonia and vocal tremor — two hallmarks that separate PD from ET years before motor symptoms become unambiguous.
Patient holds the /a/ vowel for 5 seconds. Eight Praat-derived features are extracted: jitter, shimmer, HNR, CPPS, F0 mean/std, and 4–8 Hz vocal tremor.
Patient reads a standardised passage (EN/ES/ZH/HI/PA). Six features are extracted: F0 prosodic std, pause count, pause-to-speech ratio, and pause duration statistics.
Audio is processed on HIPAA-covered Cloud infrastructure. Features are extracted; audio is cryptographically wiped before any data leaves the server.
Phase 1 builds the labelled dataset. Phase 2 returns a PD vs ET probability with top-3 SHAP biomarker drivers for the clinician to review.
Signal Science
Each feature maps to a known pathophysiological mechanism. The classifier weights them jointly — no single feature is diagnostic.
| Feature | Task | Clinical significance |
|---|---|---|
| Jitter % | Phonation | Micro-period perturbation from basal ganglia hypokinesia |
| Shimmer % | Phonation | Amplitude flutter from reduced vocal fold closure force |
| HNR (dB) | Phonation | Harmonic-to-noise ratio — reduced in breathy PD voice |
| CPPS (dB) | Phonation | Cepstral peak prominence — key hypokinetic dysphonia marker |
| F0 mean / std | Phonation | Monotone pitch (reduced std) characteristic of PD |
| Vocal tremor 4–8 Hz | Phonation | Resting tremor frequency band in the voice signal |
| F0 prosodic std | Reading | Reduced prosodic variability in connected speech |
| Pause / speech ratio | Reading | Motor planning gaps — festination and speech freezing |
| Pause count / duration | Reading | Bradykinesia signature in speech timing |
Privacy Architecture
PHI never leaves HIPAA-covered infrastructure. Audio never leaves the analysis server. Nothing identifiable is stored anywhere.
The web UI (vista-pd.com) stores nothing. No cookies, no local storage, no database. Study IDs only — no names, DOB, or MRN.
Audio streams to GCP Cloud Run, is written to an ephemeral /tmp path, features are extracted, then the file is overwritten with cryptographic random bytes before deletion — before any response is returned.
The dataset bucket stores a JSON object of 14 numeric features per participant — never audio. A study ID links the row to the clinical outcome label provided by the enrolling clinician.
The VISTA-PD mobile app records uncompressed 44.1 kHz WAV, transmits once over TLS, then discards the recording from device memory. Mic permission is granted only during recording.