Feature Selection from Lyme Disease Patient Survey Using Machine Learning

Vendrow, J., Haddock, J., Needell, D., & Johnson, L. (2020). Feature Selection from Lyme Disease Patient Survey Using Machine Learning. Algorithms, 13(12), 334. https://doi.org/10.3390/a13120334

Neural-network architecture used in the machine-learning analysis. Grey nodes represent dropout used for regularization.

Researchers at the University of California Los Angeles partnered with the MyLymeData patient registry to better understand why some people with Lyme disease improve with treatment while others continue to struggle. Applying advanced machine learning tools, the study analyzed data from thousands of patients enrolled in the registry. Patients reported whether their condition improved, worsened, or stayed the same after antibiotic treatment using the Global Rating of Change (GROC) scale. Various machine learning methods were applied to assess the ability to predict the effect of individual features such as treatment approach, clinician expertise, and duration of treatment on participant GROC responder status. These insights may be valuable to medical professionals in determining the factors that are most predictive of treatment response.

Background. Important questions remain about why some Lyme disease patients improve after treatment while others remain ill. This study applied machine-learning methods to MyLymeData survey information to identify the patient and treatment features most closely associated with self-reported treatment response.

Methods. Researchers analyzed 2,162 participants with persistent Lyme disease and 215 survey features. Linear regression, support vector machines, neural networks, entropy-based decision trees, and k-nearest-neighbor models were used to distinguish High Responders from Nonresponders and rank the most informative features.

Results. The strongest signals were concentrated in features involving antibiotic effectiveness, antibiotic-treatment duration, symptom severity, fatigue, and the type of clinician providing care. A reduced group of 30 key features performed similarly to—or better than—the complete 215-feature dataset.

Conclusion. Machine learning can help identify survey features that are most useful for predicting treatment response and may also improve future registry design by reducing redundant questions.

Study Purpose

The study investigates whether machine-learning techniques can identify the patient characteristics, symptoms, treatment approaches, and care factors that are most useful for predicting Global Rating of Change responder status.

Dataset and Methods

The analysis used MyLymeData Phase 1 survey responses from participants with persistent symptoms after antibiotic treatment. Researchers compared High Responders and Nonresponders using multiple regression and classification models and repeatedly tested the predictive value of individual survey features.

Key Findings

Features connected with antibiotic effectiveness and duration, symptom severity—especially fatigue—and treatment by a clinician focused on tick-borne diseases ranked among the most informative. The selected top features retained most of the predictive information in the full dataset.

Why This Matters

The findings demonstrate how patient-generated registry data and machine learning can help researchers focus on the variables most likely to explain treatment response. They may also help clinicians identify questions that deserve closer investigation in future prospective studies.

The complete publisher page contains the full reference list, model descriptions, tables, figures, appendix, and links to the study’s supporting materials.

Study Details

Share This Study