Interpretable Machine Learning Can Help Tell Benign and Malignant Ovarian Tumors Apart
Researchers tested five computer programs to help tell whether ovarian tumors are benign or malignant using patient data. A model called a support vector machine scored the highest accuracy, and the team also tested a system to explain the computer's choices.
In short: Researchers tested five computer programs to help tell whether ovarian tumors are benign or malignant using patient data. A model called a support vector machine scored the highest accuracy, and the team also tested a system to explain the computer's choices.
Finding out if an ovarian tumor is dangerous or safe as early as possible is vital, and computer programs are now being tested to help doctors look at medical data.
What happened, in plain words
Researchers at the Third Affiliated Hospital of Soochow University studied data from 349 patients with 171 malignant and 178 benign ovarian tumors. They tested five different machine learning classifiers—which are computer programs that learn from data to make predictions—using a method called Bayesian hyperparameter optimization to tune them. The support vector machine classifier achieved the highest accuracy of 93.7% and sensitivity of 95.3%. They also optimized a method called LIME to help explain the computer's decisions, which achieved 91.9% accuracy and pointed to specific blood markers like CA125 and CA19-9 as major predictors of malignancy.
Key points
- Studying patient records: The study used a group of 349 patients from a single center in China, looking at both malignant and benign tumors.
- Testing multiple computer models: Five classifiers were evaluated, including support vector machines, k-nearest neighbors, discriminant analysis, decision trees, and logistic regression.
- Top performing model: The support vector machine model reached an accuracy of 93.7% and a sensitivity of 95.3% using leave-one-out cross-validation.
- Explaining the predictions: The team optimized an explanation tool to show why the computer made its choices, identifying CA125 and CA19-9 as key positive predictors.
Terms explained
- Machine learning — A type of computer technology where software analyzes data to learn patterns and make predictions without being explicitly programmed for every step. Example: A streaming service looking at what you watch to suggest a new show you might like.
- Support vector machine — A specific type of computer classification model that draws a dividing line to sort data points into distinct categories. Example: Sorting an email inbox into two folders, spam and not spam, by drawing a clear boundary based on message words.
- Sensitivity — A statistical measure of how well a test correctly identifies people who actually have a specific condition. Example: A security scanner correctly noticing every single prohibited item in a bag.
Why it matters
Making computer models easier for doctors to understand is an important step toward building helpful clinical decision-support tools. However, stronger predictive results have already been reported on this same dataset by other means.
What we still don't know
External validation remains necessary, and the authors state that prospective, multicenter, and clinician-in-the-loop validation is required before these models can be used to support clinical decisions.
Based on reporting from Scientific Reports. This is an independent explainer, written in our own words with AI assistance; Scientific Reports has not reviewed or endorsed it. Read the original for the full details.