Revisiting: Why the absence of controlled human data on BPC-157 matters more than people admit Writing it up because I had to work it out twice and would rather nobody else did.
Comparing SURPASS-4 (Lancet, 2021) with STEP 1 (N Engl J Med, 2021) and finding the comparison harder than it looks.
Different populations, different durations, different endpoints defined slightly differently, and in one case a different estimand. People compare the headline percentages anyway, including me until recently.
Is there a defensible way to put these side by side, or is the honest answer that there is not and we should stop?