← back to posts
2025-08-21
GPT-5Healthcare

From GPT-4 to GPT-5 in Medical Language Understanding

If you’ve been following the progress of large language models in healthcare, you know things move fast. There’s always a new benchmark, a new claim, a new set of numbers.

This month OpenAI released their latest model: GPT-5. Unlike the GPT-4 era models, GPT-5 exposes a single interface to its users without distinction between fast models and reasoning models. The model itself decides when to think more carefully.

I recently worked on evaluating the performance of GPT-5 in healthcare. More specifically, I performed an apples-to-apples comparison between GPT-4 era models and GPT-5 in the Stanford MedHelm benchmark. The results are enlightening.

Here’s the link to the study.

Studies like this are important for society. They let us see the objective progress in the AI field beyond a polished public image.