Clinical AI is fast becoming a part of daily clinical practice.
AI is being employed by physicians in the search of medical evidence. Hospitals are doing trials with models for diagnosis and risk prediction. The use of generative AI in documentation and clinical decision support is being introduced.
Technologies are advancing rapidly.
Yet there's one question healthcare has no right to turn a blind eye to later.
Does Clinical AI apply to all patients in need?
Imagine that it's a simple question. It is not.
For most patients, an AI system might give a fantastic response, but for a smaller number of patients, it would fail in a clinically relevant manner.
Health equity must be integrated into the design of Clinical AI system development.
It should be taken into account from the outset.
AI is trained with the information that humans have created.
That means medical text books. Research papers. Clinical guidelines. Electronic health records. Images. Physician documentation.
They are not completely balanced sources.
An example of dermatology is a field that is helpful.
Investigating widely-used medical textbooks, images of brown and black skin tones constituted approximately 10.5 percent of the skin images analyzed. That is important because there can be a difference in the way that many skin disorders look on each skin tone.
If the educational content continues to display lighter skin as the assumed norm, then an AI being trained on the same type of content will pick up on the same pattern.
This issue is illustrated in a recent article on Doximity about a psoriasis query. The clinical AI tool was initially developed for lighter skin to indicate the presence of the disease. The answer was only completed when dark skin was stated in particular.
The response might have sounded like it was a good response.
However, confidence isn't a guarantee of completeness.
This is one of the major obstacles to Clinical AI.
The issue is not the system's issue.
What the system excludes.
The typical reaction to AI bias is "we need more diverse sets of data.The standard reaction to AI bias is that we need more diverse sets of data.
We do.
However, that's just the start.
Suppose that there is a data set containing patients drawn from several racial groups. That sounds representative.
Suppose some of those patients in the past did not receive easy access to specialists.Assume that some of those patients in the past could not find easy access to specialists. Less diagnostic testing was done on some. Others had trouble getting medications or follow up care.
The data set could be heterogeneous.
Even though the overall experience of the data set can be equal, it may not be the same for each person.
Clinical AI is able to learn those patterns.
Hence, developers need to think outside of the demographic box.
They must inquire about the process used to generate data.
What did the patient do prior to the time that they were entered into the dataset?
Who received testing?
Who received treatment?
Who was referred?
Who was lost to follow up?
These questions are important since AI can transform past healthcare trends into future predictions.
If the past was unequal then blindly learning from the past can reproduce that inequality.
Fairness is often discussed as a mathematical problem.
Researchers may compare sensitivity between patient groups. They may measure false positive rates or examine whether predictions are equally accurate across populations.
Those measurements are important.
But health equity asks a broader question.
What happens to the patient after the prediction?
A model could produce similar accuracy across racial groups and still contribute to unequal outcomes.
Consider two hospitals using the same Clinical AI system.
One hospital has specialists available around the clock. It has strong case management and reliable outpatient follow up.
The other hospital has limited staffing and fewer specialists.
The algorithm may identify the same level of risk in both places.
The patients may not receive the same care.
That is why algorithmic fairness should not automatically be confused with equitable healthcare.
Recent research on Clinical AI fairness has highlighted this gap. Much of the field still focuses on differences in model performance while paying less attention to what happens when those models enter real clinical systems.
The distinction matters.
A fair prediction is useful.
A fair outcome is better.
Frequently, the reply to inquiries regarding Clinical AI is straightforward.
Ensure there is a human in the chain.
That is important.
The use of AI for recommendations should not be taken as gospel by physicians.
It is impossible for any one clinician to detect all AI issues.
Physicians are currently analyzing lab tests. Imaging reports. Patient messages. Alerts. Medication changes. Documentation requests.
Now imagine being asked to spot subtle bias in each of their AI generated answers.
There are some that will be apparent.
Others will not.
It is possible to write a correct answer in which none of the statements are incorrect. It may just not talk about a critical diagnosis or treatment plan.
This kind of mistake may not be easily identified.
One of the bigger clinical AI safety assessments was conducted by Doximity and included over 1,100 cases developed by physicians in 10 clinical specialties. One significant finding was that serious safety issues frequently related to recommendations that were not made, but not to obviously inaccurate recommendations.
That's changing the part of the human watchman.
The use of the tool cannot be done by the doctor alone.
There also needs to be oversight on the model, at the health system level and throughout its life.
A Clinical AI system shouldn't be tested once and be deemed safe for all eternity.
Healthcare changes.
Patient populations change.
Treatment guidelines change.
Models change.
Clinical practice can change even with the use of a tool.
This implies that equity monitoring must be ongoing following deployment.
Health systems should consider if there are disparities in performance by patient population.
They should track the clinical recommendations that are accepted.
They must consider false positive and false negative.
They should review additional testing and/or treatment after an AI recommendation.
Most importantly they should develop a method whereby unusual patterns can be reported by clinicians.
This overall strategy is supported by recent implementation research. DQ and model bias are recognized as key issues for reviews. They also point to workflow design and continuous monitoring as key components of ensuring safe implementation of AI.
Equity is not a feature, therefore.
It is a process.
A Clinical AI model created at a large academic hospital might not perform as well at a community hospital.
Patient population might be different.
Resources can vary.
Clinical processes can vary.
The record-keeping of information can also evolve.
Hence the need for local validation.
Before assuming that a model is effective for a patient population, hospitals should first determine if a model is effective for their patients.
This is even more significant in safety net and underserved environments.
The bottom line is that health systems don't need to assume that a predictive model is fair because it is accurate, and NYC Health + Hospitals has proven that they can explicitly consider sociodemographic performance differences.
It ought to be normal to have such kind of local evaluation.
Clinical AI should be shown to be effective in the real world of clinical practice.
There's another facet of the equity discussion that doesn't get as much attention.
Clinicians require spaces for discussing observations.
In some patients physicians might find that the AI tool is missing part of the differential diagnosis, regularly.
A nurse may perceive that recommendations are out of alignment with the practical application on the bedside.
A pharmacist can recognize that a med advice may not be optimal for a specific population.
One clinician might believe that the issue is confined.
A healthcare community can show another person they are seeing the same thing.
Platforms like MedSocially can come in handy in that capacity.
Healthcare communities can bring together physicians and nurses, researchers and other health care professionals. They are able to reflect on the role of Clinical AI in an unaided clinical assessment.
This is not an alternative to the formal process of governance.
It's another information source.
Conversations about the real world can reveal issues that the performance dashboard may not.
In order to make Clinical AI more equitable, there needs to be a way for the people using it to have a meaningful way to share what they're experiencing.
Health equity is not a "compliance at the end of the day" issue in Clinical AI.
It should influence the whole process.
Who are the subjects of the data?
Who is missing?
Is there variation in the model performance for patient groups?
Does it miss out any key information?
What happens after its recommendation?
Is it possible for all hospitals to follow that recommendation?
Are problems reported by clinicians?
Does someone continue to monitor after the tool is implemented?
These questions pose a challenge to the development of AI.
This doesn't always sound negative.
Healthcare is hard.
Not only should a Clinical AI system be able to give an answer quickly, but the answer should be impressive as well.
That answer should be evaluated based on whether or not it is safe and helpful for the patient in front of us.
In particular, the one patient who has been neglected in the past.
The point is not to create functional AI, it is to create a usable, useful, and responsive AI.
It's important to work towards creating Clinical AI for all.