EVIDENCE MATRIX / FOUR FILES
A Scorecard for What the Studies Can Bear
The same four questions applied across metabolic, mitochondrial, immune, and copper-peptide research.
The short version
These four peptides do not compete for one biological job. They belong together because their evidence lives at different stages. Retatrutide has randomized human trials with direct metabolic endpoints. MOTS-c has strong mechanistic and animal work but no human efficacy trial. Thymosin Alpha-1 has decades of clinical study and a large modern sepsis trial that did not confirm the hoped-for survival effect. GHK-Cu has plausible repair biology, a handful of topical human signals, and a delivery problem that constrains interpretation.
The matrix asks four questions: Was there a control group? Was the sample large enough for the claim? Did the endpoint matter to people or only to a laboratory model? Has the finding been reproduced? No single column declares truth. Together they show what a result can carry. The broad conclusion is simple: evidence maturity follows study design and repetition, not the novelty of a mechanism or the size of a headline.
The data-quality matrix
| Compound | Controls | Sample size | Endpoint quality | Reproducibility | Present reading |
|---|---|---|---|---|---|
| Retatrutide | Randomized placebo and active-controlled human trials [3][4][5][6] | Moderate Phase 1 and Phase 2 cohorts | Direct weight, HbA1c, and MRI liver-fat measures | Metabolic signal repeats; long-term outcomes pending | Strongest clinical efficacy evidence here; still investigational |
| MOTS-c | Controlled cell and mouse experiments [8][11][12] | Model-dependent; human cohort small [9] | Binding, signaling, animal function, human association | Mechanism coherent; no human intervention replication | Translational hypothesis, not established human benefit |
| Thymosin Alpha-1 | Large blinded placebo-controlled sepsis trial plus earlier studies [13][17] | Large modern sepsis cohort | Mortality is direct and clinically meaningful | Largest trial did not reproduce earlier hoped-for benefit | High-quality null for sepsis deserves greatest weight |
| GHK-Cu | Small topical or combination-product studies; laboratory models [18][20][21] | Small human samples | Cosmetic, hair-count, gene, and penetration endpoints | Limited independent standardized replication | Plausible topical research; systemic claims unsupported |
Controls: the counterfactual column
A control group asks what would have happened without the compound. Retatrutide's central studies use randomized comparators, allowing differences in weight, glycemic measures, and liver fat to be attributed with far more confidence than a simple before-and-after account [3][4][5]. Thymosin Alpha-1's TESTS trial adds double blinding and placebo control, making its negative result especially persuasive [13].
MOTS-c also has controlled experiments, but they answer model-level questions: whether a molecule binds CK2, changes signaling in cells, or affects performance in mice [8][11][12]. Good control does not erase the species boundary. GHK-Cu's hair study used placebo, yet the tested product combined ingredients, so the control supports the combination more than the peptide in isolation [20].
The practical rule is to identify the missing alternative explanation. No comparator leaves time and expectation unresolved. An active mixture leaves ingredient attribution unresolved. A mouse control leaves human translation unresolved. “Controlled” is the start of appraisal, not its end.
Samples and endpoints: size meets meaning
Sample size must be judged beside the claim. The retatrutide obesity and diabetes trials enrolled hundreds of adults, enough to make their primary changes difficult to dismiss as a handful of unusual responders [4][5]. The liver-fat substudy was smaller and more selected, so its dramatic MRI result needs narrower language [3]. Thymosin Alpha-1's TESTS trial enrolled more than a thousand participants and measured mortality, giving the null result both scale and clinical weight [13].
MOTS-c's human cohort measured circulating peptide in a small dialysis population. It may inform risk prediction but cannot establish treatment benefit [9]. GHK-Cu's randomized hair study measured counts, a concrete endpoint, yet included only a small group and tested a combination [20]. Its gene-expression and penetration findings are still farther upstream [19][22].
A technically precise surrogate can be valuable. It can validate target engagement or explain a pathway. But it should not be translated into “lives longer,” “ages more slowly,” or “recovers better” unless those outcomes were actually measured.
Reproducibility: where stories harden or soften
Retatrutide's metabolic effects recur across separate controlled studies and distinct endpoints, which raises confidence in the core signal [3][4][5][6]. The missing replication lies farther out: broader Phase 3 populations, long-term outcomes, and safety after longer exposure.
MOTS-c shows conceptual repetition across AMPK signaling, nuclear stress response, CK2 activity, and mouse function [8][10][11][12]. Yet much of that literature remains preclinical, and some findings depend on a limited research lineage. The decisive repetition—a controlled human intervention—does not yet exist.
Thymosin Alpha-1 demonstrates replication working as correction. An earlier sepsis trial leaned positive but uncertain [17]; the larger, more rigorous TESTS trial was decisively null [13]. GHK-Cu has recurring repair themes across reviews and laboratory work, but formulations, endpoints, and study scale vary, making a single clinical estimate difficult [18][19][21].
Reproducibility is not mere duplication. It asks whether the result survives changes in team, setting, population, and method. A finding that travels becomes more general. One that changes teaches where its boundary lies.
What this comparison does not do
The scorecard does not rank personal suitability, recommend a compound, or convert animal study conditions into human practice. It does not treat regulatory approval as a perfect proxy for scientific truth, though approval status matters greatly for product quality, oversight, and the scope of established use.
It also resists a common false symmetry. “More human data” does not mean “more positive data.” Thymosin Alpha-1 has a large human trial precisely because the question was tested seriously, and the answer for sepsis mortality was negative [13]. “More mechanism” does not mean “more benefit.” MOTS-c and GHK-Cu have fascinating molecular narratives, but their strongest claims remain limited by translation or delivery [8][18].
The matrix is best read as a map of next questions. Retatrutide needs mature outcomes and longer confirmation. MOTS-c needs controlled human intervention research. Thymosin Alpha-1 needs indication-specific claims that respect the sepsis null. GHK-Cu needs standardized formulations, larger independent topical trials, and route-specific safety evidence. Curiosity survives all four conclusions; certainty is simply kept in its proper column.