There is a version of this argument that is about misconduct, and it is not the one I am making. Nothing in the BPC-157 literature looks fabricated to me. The surgeries appear to have been done, the tendons appear to have been pulled until they failed, and the numbers appear to be the numbers. The problem is structural and much more boring than fraud, which is precisely why it has survived thirty years without anyone having to defend it: an enormous share of what is known about this peptide was produced by people who already worked together, using the same model, in the same place.

That is a description, not an accusation. It is also, if you care whether any of this is true, the most important fact about the field.

What the rat tendon work actually reports

The canonical experiment is Achilles tendon transection in rats. The tendon is cut, the animal is allowed to heal over a defined period, and the repaired tissue is then assessed — usually by biomechanical testing, where the tendon is pulled to failure and the load it withstands is recorded, and usually alongside histology, where a section is stained and scored for collagen organisation, cellularity, and vascularity. Sometimes functional endpoints such as gait analysis are added. Variations of this design appear in a rat medial collateral ligament transection model and in rat muscle crush and transection models, and a separate cluster of work uses gastric and colonic lesion models in rats.[1]

Read individually, these papers are unremarkable in form. They report treated animals healing faster, stronger, or with better-organised tissue than controls. The effect sizes are often large. The direction of effect is almost uniformly favourable, across tissue types, across injury mechanisms, and across two decades — which is the first thing that should give a reader pause, because that degree of consistency is rare in any biological literature and is more commonly a signature of how a body of work was assembled than of how nature behaves.[2]

It is also worth being clear about what a rat tendon is not. Rat Achilles tendon is smaller, differently vascularised, and loaded differently than the human equivalent, and a quadruped's gait distributes force through it in a way bipedal walking does not. Rodents heal cutaneous and connective tissue injuries faster than humans do, with a different balance of contraction and remodelling. A transected tendon in an eight-week-old rat housed in a cage is not a partially degenerated tendon in a fifty-year-old, and the difference is not a matter of scale. Every one of those gaps sits between the experiment and any sentence about a person, and each one has to be crossed by assumption rather than by data.

Group sizes are small, typically at the low end of what is standard for rodent surgical work. Blinding is inconsistently reported: many of these papers do not state whether the person scoring the histology knew which animal had been treated, and histological scoring is exactly the kind of judgement that unblinded assessment inflates. Randomisation of animals to arms is frequently unstated. None of this is unique to BPC-157 — it describes a large fraction of preclinical surgical research — but the usual corrective for it is that somebody else, somewhere else, runs the experiment and reports what they get. That corrective is what is missing here.

Abstract scatter plot in which each point carries a tall vertical error bar, several spanning much of the plot area, illustrating wide uncertainty around individual estimates.
With small groups, a statistically significant result is not a precise one. The estimate that clears the threshold is systematically larger than the true effect — an artefact known as the winner’s curse.

Why one laboratory is a structural problem, not a moral one

A laboratory is not a neutral instrument. It is a set of hands that have done a surgery several hundred times, a particular strain of rat from a particular supplier, a housing arrangement, a scoring rubric someone wrote in the 1990s, a way of deciding when an experiment has gone wrong and should be discarded, and a shared sense of what the answer is probably going to be. Every one of those is a variable. None of them is recorded in the methods section.

When the same group runs the experiment forty times, those variables are held constant by definition. That is excellent for internal consistency and useless for establishing that the finding exists outside the room. Repetition within a laboratory is not replication; it is the same measurement repeated with the same systematic error. The whole epistemic function of independent replication is to vary the things nobody thought to write down.

There is a second, subtler effect. A group that has published a favourable result has an interpretive commitment to it. Not a dishonest one — an ordinary human one. Ambiguous histology gets read in the direction of the prior. An experiment that fails gets attributed to a technical problem and repeated, while an experiment that succeeds gets written up. Over thirty years, that asymmetry alone is sufficient to produce a literature that looks overwhelming and means very little. No individual decision in that chain is misconduct. The aggregate is still a distorted record.[3]

Repetition inside one laboratory is not replication. It is the same measurement, repeated with the same systematic error.

A citation record is not a verification record

The most common defence of this literature is its size. There are a great many papers. They agree with one another. Surely that counts for something.

It counts for less than it looks like, because of how citation works. When a review article summarises a finding, it inherits that finding's evidentiary status without inheriting its caveats — the sample size, the unstated blinding, the single-institution provenance. A second review then cites the first. By the third or fourth generation the claim is being supported by a chain of documents none of which involved a rat. The count rises; the evidence does not. Our contributor Rina Iyer has traced this pattern in other literatures and can usually identify the precise sentence where a hedge was dropped, and it is almost always inadvertent — a summariser compressing “was reported to accelerate healing in a rat model” into “accelerates healing.”

There is also a filtering effect on what enters the record at all. A group that runs an experiment and observes nothing is unlikely to publish it, because null results are difficult to place and career-negative to produce. If several unaffiliated laboratories had quietly attempted this experiment and found nothing, the published literature would look exactly as it currently looks. That is the uncomfortable part: the visible record is compatible both with a real effect and with an absent one, and it cannot distinguish between them. This is precisely what the publication you are reading is named after.

What independent replication would actually require

It is worth being concrete about this, because “more research is needed” is the phrase that lets everyone stop thinking. A replication that would move the BPC-157 tendon question forward has a specific shape.

It would be run by a group with no prior publications on the compound and no collaborative relationship with the originating lineage. It would pre-register the protocol, the primary endpoint, and the analysis plan before the first animal was operated on, so that the outcome could not be selected after the fact. It would blind both the surgeon and the assessor to allocation, and it would say so explicitly in the methods. It would power the study to detect an effect substantially smaller than the ones already reported, on the assumption that published effect sizes in a small-n literature are inflated. It would use a primary endpoint chosen in advance — load at failure, say — rather than reporting whichever of six measures reached significance. And it would be published whether the result was positive, negative, or null, which in practice means it needs funding that does not depend on the answer.

Almost none of that is exotic. It is standard practice in clinical trials and increasingly standard in preclinical work through initiatives that have pushed for pre-registration and reporting guidelines in animal research.[4] The reason it has not happened for this compound is mundane: replication is expensive, unglamorous, difficult to publish, and career-neutral at best. The incentive structure of academic science funds novelty. Nobody gets a grant to find out whether a thirty-year-old rat experiment holds up.

Abstract composition of vertical bars of varying width and tone resembling the spines of bound journal volumes shelved together.
Citation counts grow whether or not anyone repeats the experiment. A frequently cited finding and a frequently tested one are different things, and the literature does not distinguish them.

What follows from none of this

Here is where a piece like this normally resolves. I do not have a resolution, and I would rather say so than manufacture one.

The honest position is that the rat tendon findings for BPC-157 are real observations that have not been independently verified, produced under conditions that tend to inflate effects, in a species whose tendon healing differs from ours in vascularity, loading, and scale. They constitute a hypothesis. They do not constitute evidence of an effect in people, and there are no published peer-reviewed randomised human efficacy results to fill that gap — not weak ones, not preliminary ones. There are none. Any sentence about what this compound does in a person is currently an extrapolation across a species boundary from an unreplicated literature, and it should be read as one.

What frustrates me is that this is a genuinely testable question. Someone could settle a large part of it with one well-designed, pre-registered, blinded, adequately powered rat study from an unaffiliated group. Thirty years in, that study still has not been published. Until it is, the accurate description of this field is not “promising” and not “debunked.” It is unfinished — and it has been unfinished for long enough that the failure to finish it is itself the finding.

References and further reading

Described generically rather than by citation. We list what to search for and where the work sits, not identifiers we have not verified line by line.

  1. The primary rodent literature on this compound — rat Achilles tendon transection, rat medial collateral ligament transection, rat muscle injury, and rat and mouse gastrointestinal lesion models — published across the 1990s, 2000s, and 2010s, largely by a research lineage based in Zagreb and its collaborators. Indexed in the standard biomedical databases under the compound name.
  2. The general methodological literature on effect-size inflation in small preclinical studies, including work on the winner’s curse and on the relationship between statistical power and the positive predictive value of a published finding. Widely reviewed in the meta-research literature of the past two decades.
  3. Empirical studies of reporting quality in animal research, examining how often randomisation, blinding, and sample-size calculation are described in published in vivo experiments, and the association between their absence and larger reported effects.
  4. Reporting and design guidelines for in vivo research, and the associated literature on pre-registration of preclinical studies and on multi-laboratory replication projects in animal work.