Journal club: STEP 8 and the fairness of the comparator dose — setting out what I have, and where I think it stops being reliable.
Comparing SURMOUNT-2 (Lancet, 2023) with STEP 8 (JAMA, 2022) and finding the comparison harder than it looks.
Different populations, different durations, different endpoints defined slightly differently, and in one case a different estimand. People compare the headline percentages anyway, including me until recently.
Is there a defensible way to put these side by side, or is the honest answer that there is not and we should stop?