We Have Heard This Before. The Answer is to Learn.

Ten years ago, a question from the audience unsettled me enough to change the direction of my career.

It was 2016. I was still in training, speaking at a national conference about the nuances of workflow optimization in imaging IT. After my talk, a data scientist from IBM walked to the microphone and asked, in effect: why bother with human radiologists when algorithms are clearly better at the work?

She also mentioned a renowned scientist who had recently gone on record to say the same. The smartest people in the world were asking this question.

I did not have a satisfying answer. I also did not know much about machine learning or the work of data scientists. But I had spent years of my life preparing for a profession that someone was suggesting might be unnecessary. The question bothered me.

So I decided to learn everything I could about it.

A familiar question, ten years later

I thought about that moment when I read Robert F. Kennedy Jr.’s recent remarks that AI could give patients a second opinion “much better informed than any doctor in the country.”

My reaction as a radiologist was less shock than recognition. We have been hearing versions of this argument for a decade.

As the argument reaches more of medicine, colleagues in other specialties are encountering a question that has shaped radiology’s professional conversation—and much of my career.

Six physician organizations, including the American Medical Association, responded with a joint statement emphasizing clinical context, professional judgment, and responsibility for care. They warned that portraying AI as inherently better informed could undermine patients’ trust.

They also explicitly supported physicians helping lead responsible AI development and use.

Those concerns are legitimate. “Better informed” is not the same as more accurate, and being more accurate on a defined task is not the same as delivering better care.

Which system, doing what, for which patients, with what evidence? Those are necessary questions, not evasions.

But for individual physicians, the response cannot end with explaining why the claim is too broad.

If the possibility bothers you, let that discomfort become a reason to learn.

What happened after that conference

Over the next several years, I learned everything I could about the technology behind that question.

Between 2016 and 2018, I worked through An Introduction to Statistical Learning (it still holds up). I learned that cleaning the data and defining the predictive task were more important than training the model itself. I became very familiar with DICOM.

I also placed on the private leaderboard of a Microsoft-sponsored machine-learning competition, using a boosted random forest model and a Gates Foundation dataset to predict women’s health risk.

Along the way, a co-founder and I created an NLP solution to automatically audit clinical notes for discrepancies, hoping to catch medical errors before they happened. We became finalists in UPenn’s business competition and launched a startup (it didn’t do well—another story for another time).

In 2018, I published a peer-reviewed study I led with Hanna Zafar, Maya Galperin-Aizenberg, and Tessa Cook, using natural language processing and machine learning. It was one of the earlier radiology studies to systematically evaluate machine learning and NLP for large-scale mining of unstructured radiology-report data.

Later that year, I joined the Cleveland Clinic.

Today, I get to spend much of my working day thinking about AI: how it changes healthcare workflows and whether it can reduce the burdens that contribute to clinician burnout. I get to talk about what it means to be responsible for AI in the same discussion as being responsible for patients. And best of all, I get to do this work with folks who really, really care.

The question that frightened me did not become irrelevant.

It became part of the fuel that shaped what I now call a career.

Earn the disagreement

I am not suggesting that every physician needs to become a machine-learning researcher. My path was one path. But there is a difference between disagreeing with a claim because we understand its limitations and disagreeing because we dislike its implications.

And when rigorous evidence shows that an AI system performs better than we do on a task, we should be the first to recognize it. A commitment to evidence-based medicine cuts both ways.

Learning also requires more than trying a chatbot once. It means examining evidence, working with people who understand the methods, and observing how a tool behaves in an appropriately governed setting.

And it means building. It is the widest gap you cannot bridge by reading papers. “Vibe code” an app for your personal life. Make an AI agent to answer questions about something you know by heart. Do it with the utmost care and follow hospital policies—but you must try. After all, that’s what our patients will be doing for their own health.

Familiarity is a beginning; expertise requires knowing when a tool’s apparent competence does not apply.

Ten years ago, I could have treated that audience question as something to rebut and forget. Instead, it pushed me to learn a technology that now shapes much of my work.

If you are a physician worried about what AI means for your patients or your profession, take the concern seriously. Learn enough to disagree for good reasons, and enough to recognize when the technology has something to teach you.

Our relevance will not come from insisting that AI cannot do our work. It will come from understanding it well enough to transform what we do.

Systems Thinking: What Doctors Do to Prepare for the AI Era

When people talk about preparing doctors for the AI era, the conversation often goes straight to tools: which model, which vendor, which use case. Those questions matter. But there is an earlier question that may be more useful: what kind of problem is this system actually presenting?

Medicine is full of problems that look similar from a distance and behave very differently up close. Treating each one as a technical problem is an easy way to create a technically impressive solution that does not improve much.

Continue reading →

RSNA 2026 Knee MRI Challenge: Early Momentum

The RSNA 2026 Knee MRI Challenge is gaining remarkable momentum, and I am deeply grateful to everyone contributing. It’s been both a fulfilling and a humbling experience. As of August 29, more than 16,700 people have joined, with nearly 2,900 competitors forming over 2,600 teams and submitting more than 24,700 models.

Almost 500 public notebooks and 85 forum topics reflect a community that is not just competing, but actively teaching and learning. Participation and iterative activity continue to rise, while both leading and median public leaderboard performance have strengthened impressively.

What excites me most is the conversation. Data scientists, trainees, and licensed radiologists from around the world have posted on the discussion boards about the challenge design, the dataset, and the realities of clinical data. Discussions about why there is no beautifully crafted “ground truth,” and what to do with nuanced, free-text radiology reports, go directly to the educational mission of this challenge. These are not side questions; they are central to building useful, clinically meaningful AI.

I am extremely impressed by the competitors’ performance so far, and thankful for the rigor, curiosity, and generosity participants are bringing to the challenge.

Learning AI From Radiology’s Messy Middle

I am frequently asked by high school and college students what AI in radiology means, and by residents how to enter the technical side of imaging AI models. More importantly, how to build one from the real life of radiology?

This year’s RSNA Knee Abnormality Detection Challenge is one answer.

It is not a pristine tutorial dataset. Competitors will work from carefully pooled knee MRI images and their reports in 10 languages, sourced from 19 sites across 5 continents: the hedges, imperfect descriptions, and occasional errors that make clinical data both difficult and familiar.

The task is also not accuracy at all costs. Models must be efficient, and they must address a representative set of 12 knee abnormalities, not one model for one diagnosis.

That tension is the point. Building useful AI means confronting what radiology actually produces, then deciding what is worth measuring and improving.

As of August 17, there have been 12,853 joined users, 1,890 competitors across 1,761 active teams, and 11,456 submissions; the public leaderboard’s top, median, and last-listed scores are 0.951, 0.8745, and 0.468, respectively.

I am grateful to the RSNA staff, the Challenge Task Force, and especially Dr. Errol Colak for their leadership. I am equally thankful to my co-lead, Dr. Naveen Subhas. This Challenge could not have been pulled together by just one team: we owe this work to over 100 radiologists, biostatisticians, and data scientists across the world who volunteered their time reviewing reports, annotating images, offering advice, and contributing data.

We hope this challenge gives learners a practical place to begin and a clearer sense of the questions that remain.

AI Can Raise the floor, But People Determine The Ceiling.

AI readiness is often treated as a technical skill, as though one could study for it, take a test, and be done. It is a pitfall particularly in education – what does it even mean to teach residents to be ready for the AI age?

Researchers at UT Austin asked 523 early-career professionals to do client-like work with an AI agent, then compared their results with an AI-only baseline. About half were “AI Amplifiers”: they did better than AI alone. A quarter were “Delegators”: their work was about as good as the AI’s. The rest, “Apprentices,” did worse, even though they were not necessarily less capable.

Yes – about a quarter of the early-career professional appeared to produce worse outputs when using AI compared to than just having it done by AI alone. (The reality is probably that the researchers were better prompt engineers than the apprentices).

Strong users framed the problem, used real domain knowledge, checked the output, and refined it over several rounds. AI was used as a collaborator that needed direction and oversight.

This feels familiar in resident education. An algorithm can produce a useful first pass. Safe clinical value still depends on whether the clinician asks the right question, spots what does not fit, and knows what to do next.

The lesson is not to teach longer prompts. It is to teach judgment in the workflow: frame, check, refine, decide.

AI can raise the floor. People still determine the ceiling.

Source: McCombs School of Business, University of Texas at Austin.

Navigating AI Decisions: Going Beyond “Buy vs Build”

Conceptual illustration of an AI buy-versus-build crossroads with governance and monitoring decisions

“Should we buy or build?” is usually the first question organizations ask about radiology AI. It is useful, but it is no longer the whole question.

Today, you can choose from cleared clinical algorithms, enterprise AI platforms, generative-AI tools, and internally developed workflows. The real challenge begins after a contract is signed or a prototype works. AI becomes another clinical system: one that can help, distract, confuse, or quietly change behavior.

Anyone who has lived through an enterprise IT rollout, PACS replacement, or hospital merger (let alone lead an aspect of it) knows the pattern. Installation is only one part of deployment. The hard work is making a tool fit the people, data, incentives, exceptions, and failure modes of the place where it will be used.

Continue reading →

AI Reimbursement Is Becoming a Workflow Problem

CMS has proposed a new way to recognize and potentially pay for certain software-based clinical services (always start with the fact sheet – far more readable than the full text). This may matter in radiology, but it does not mean every AI tool will suddenly be reimbursed by Medicare.

The proposal calls these tools Software as a Medical Service. Examples include software that extracts useful information from medical images, such as heart blood-flow analysis, fracture-risk scores, or brain MRI comparisons. CMS would give some of these services their own payment category while it learns more about how they are used.

Medicare is beginning to acknowledge that some software can provide clinical information beyond a simple workflow aid, and it augments rather than replace physician work (Radiologists have been saying this for a while, and other -ologies expected to follow as product lines expand). Still, the proposal is temporary and limited. Many AI tools would not qualify for separate payment, and payment will depend on appropriate ordering, documentation, billing, and medical necessity.

Clear language matters – CMS differentiates between assistive, augmentative, and autonomous. Software that helps a radiologist work faster (assistive) is different from software that provides new, clinically useful measurements (augmentative). Both differ from a system that makes a diagnosis without a clinician (autonomous). Administrative tools such as scheduling or drafting messages can be useful, but they are not diagnostic services.

The bigger challenge is workflow. Buying or building a tool is only the beginning. All of it involve some form of cost. Hospitals must integrate it into their systems, review privacy and security, train staff, check its performance, document its use, bill correctly, manage denials, and show that it improves care. On the flip side, a billing code alone does not guarantee revenue.

For radiology leaders, the practical next step is to take inventory: Which software tools are already in use? What do they add to patient care? Are eligible services being documented and billed correctly? What do they cost, and what clinical decisions do they improve?

CMS’s proposal is a meaningful step, but not a windfall. The opportunity is to identify the software services that provide distinct value, and build reliable workflows around them.

“The Moon is Orange” – When the Air Gets Visible

On the way to summer camp yesterday morning, my kids pointed to the sky and said “look the moon is orange!”

The strange thing about bad air is that it is usually abstract until it is not. It turns out, that was the sun at 8 am on a ‘sunny day.‘ There was so much ash in the air that you could directly stare at the sun with the naked eye, even mistaking it as the moon.

On Thursday morning, Cleveland’s sky made air quality very concrete. The sun looked filtered, the horizon was soft, and everything had that slightly apocalyptic orange-gray cast that we have now learned to associate with Canadian wildfire smoke.

A dim orange sun over Cleveland streets through heavy Canadian wildfire smoke.
I took this photo from my car at 8am. Smoke from Canadian wildfires dimmed the Cleveland sky on July 16.

And so I wondered if we could turn this situation into a little project of our own.

Continue reading →

FDA AI Guidance and the Hard Part of Transparency

The least glamorous part of AI in radiology may turn out to be the most important: telling people what changed.

That sounds simple. It is not. A diagnostic AI tool may be trained on one dataset, validated on another, deployed inside a PACS or reporting workflow, monitored after release, and then updated when the model, threshold, input, interface, or intended environment changes. Somewhere in that chain, a radiologist is expected to decide whether to trust a box, a score, a contour, a triage flag, or a sentence.

Continue reading →

AI Readiness as an Educational Obligation

I wrote a guest editorial in Academic Radiology about AI readiness as a practical educational obligation, rather than an optional informatics side quest.

The piece responds to a study of medical undergraduates and radiology trainees in China, but the issue is broader: enthusiasm and exposure do not equal competence. Radiologists need enough AI literacy to recognize workflow fit, automation bias, governance gaps, and the limits of machine suggestions. Most trainees will not become model developers.

They still need to become safe, skeptical operators who can use AI with confidence and accountability in real clinical environments, worldwide, across diverse resource settings today.