← The library
Technical corpus, written for specialists. For the plain-language version of these ideas, start with the framework.
Applications

The Philosophy Experiments Under the Framework

Working draft This public copy has not yet had its final author pass; wording may still move.

A Consistency Stress-Test Across Fourteen Moral Thought-Experiments

Developed in conversation, May 31, 2026 — a structured test of the framework against the ethically-relevant experiments compiled from philosophyexperiments.com


What this document is

The applied documents so far have each taken a single domain — personal morality, politics, animals, the relation between moral psychology and moral grounding — and shown what the framework produces there. This piece does something different in kind. It runs the framework against a battery of thought-experiments specifically engineered to detect inconsistency. The experiments compiled from philosophyexperiments.com are not neutral prompts; most are diagnostic instruments, built to catch a respondent holding two principles that cannot both be true, or applying a standard to one case and quietly suspending it for a structurally identical one.

That makes the battery an unusually sharp test, because the framework's central Tier 2 procedural principle — treat symmetrically situated apertures symmetrically — is itself a consistency demand. The experiments reward exactly the discipline the framework claims to institutionalize. So the test has real stakes in both directions: if the framework cannot keep its own footing across these cases, the symmetry principle is hollow; and if it does keep its footing, the interesting question becomes why it does where unaided intuition does not, and whether the questions designed to force contradiction succeed against it or merely re-expose problems the framework has already named as open.

The short findings, stated up front. The framework answers the battery with a single engine and answers it consistently — the same apparatus that handles the incest case handles the trolley, the lifeboat, the antique, and the Euthyphro. The "trick" questions do not drive it into contradiction, because the contradictions they detect are contradictions in surface-feature reasoning (killing versus letting die, near versus far, touching versus not touching, my-side versus your-side), and the framework's moral variable is never the surface feature. What the battery does do is press the framework, repeatedly and in first-order life-and-death cases, onto the three open problems it already carries — aggregation, the autonomy/paternalism line, and the is/ought residue — and surface one structural seam the foundational documents have not yet made explicit: a latent agent-relativity in the framework's own grounding. Those are the contributions of this exercise, and the rest of the document earns them.


How the experiments are built, and why it matters

A diagnostic moral experiment works by isolating a variable. It presents Case A, records the verdict, then presents Case B which differs from A only in some feature the designer suspects is doing illegitimate work, and asks for the verdict again. If the respondent flips between A and B, the designer has located a feature the respondent treats as morally decisive but cannot defend as relevant. The Loop trolley exists to catch people whose Switch/Footbridge verdicts rest on "did you physically touch the victim." The Vintage Sedan and the Envelope exist to catch people whose duty-to-aid rests on "is the victim near me." The Disposition test's fourteen scenarios exist to catch people whose epistemic standard ("concede when your empirical argument is refuted") is applied to their opponents' arguments but not their own.

The framework's exposure to this style of attack depends entirely on what it treats as the morally relevant feature. If the relevant feature is a surface property, the framework is as vulnerable as any respondent. If the relevant feature is structural — whether and how a stake-geometry is collapsed, whether an aperture is instrumentalized, whether two cases are structurally indistinguishable — then the surface variable the experiment isolates was never load-bearing, and flipping it produces no inconsistency to catch. The framework's verdicts across the battery are best read as a sustained demonstration that its relevant features are structural, which is exactly what immunizes it against consistency-probes built around surface features. This is not the framework getting lucky. The symmetry principle is the consistency demand stated as doctrine; the experiments are that same demand stated as a quiz.


The trolley family: instrumentalization versus forced allocation

The trolley cases — Should You Kill the Fat Man?, the Lifeboat / Loopback items, and Murder in the Playhouse — are the battery's centerpiece and its sharpest consistency trap. The classic finding is that most people permit the Switch (divert the trolley onto a side track, one dies instead of five) and forbid the Footbridge (push a large man off a bridge to stop the trolley, one dies instead of five), while being unable to name a relevant difference. The Loop case is then introduced to close the escape routes: it has the switch-flipping surface of the permissible case but the body-as-trolley-stopper structure of the forbidden one, so any respondent whose distinction rested on "flip versus push" is caught.

The framework is not caught, because its distinction never rested on flip-versus-push, on killing-versus-letting-die, or on the act/omission line. Its relevant variable is instrumentalization. Stake-Geometry grounds the wrongness of treating someone as a mere means structurally: doing so "flattens their stake-geometry to a single dimension (their use to you)... a forced geometric simplification." That gives the framework an independent, non-arbitrary reading of the three cases.

In the Switch, the one person on the side track is not used. They are not the mechanism by which the five are saved; they are killed as a foreseen side-effect of redirecting a threat that exists independently of them. Were the side track empty, the five would still be saved. No geometry is instrumentalized; a genuinely unavoidable harm is being allocated under a forced choice between symmetric stakes. In the Footbridge, the large man is the mechanism — his body is the trolley-stopper, and were he not there (or not large enough) the five could not be saved this way. His geometry is conscripted and flattened to its use-value, which is the paradigm dignity-violation the framework already names. In the Loop, the surface returns to switch-flipping, but the one's body is again causally necessary to stop the trolley before it returns to the five. Structurally the Loop is the Footbridge: the victim is the instrument. The framework therefore verdicts Loop with Footbridge — impermissible — and does so consistently, because the variable it tracked through all three cases (is a geometry being used as a means?) never changed when the surface did.

This delivers the doctrine-of-double-effect verdicts without importing double effect as a brute rule. Where the tradition asserts a difference between intended and merely foreseen harm and is then pressed to justify it, the framework supplies the justification: the morally loaded thing is not a metaphysics of intention but the flattening of a stake-geometry to a single instrumental dimension. The intention/foresight distinction the Murder in the Playhouse items probe falls out of the instrumentalization analysis rather than being posited alongside it.

Two results of this are worth stating plainly, because together they fix the framework's character.

First, the framework is not consequentialist, and the trolley series is where that becomes unmissable. A reader who took the framework's talk of "minimizing collapse" as a summation rule would expect it to push the fat man — five collapses prevented at the cost of one. It does not, because density is not a quantity to maximize and the anti-aggregation commitment is explicit: "an ethics that demands the destruction of one geometry to serve another's is doing the negative thing it claims to oppose" (the vet case in Relations Between Apertures). The Loop is precisely where the framework's own "minimize collapse" pull and its "don't instrumentalize" prohibition collide, and the framework resolves the collision in favor of the prohibition. That is a real deontological bullet — it accepts that letting five die can be better than using one as a means — and the framework should own it rather than soften it.

Second, and in the opposite direction, the framework does lean toward diverting in the Switch, and here it leans on the open aggregation problem. Saying the Switch is permissible because it produces "less total collapse" in a forced choice is a mild aggregative step, and the framework officially holds that conflict-weighting is type-matched, not aggregate — it weighs the particular symmetric stakes, not summed densities. It can reframe the Switch as allocating an unavoidable harm among symmetric parties rather than summing goods, which is honest as far as it goes, but the intuitive force of "five rather than one" is doing aggregative work the official machinery disavows. This is the battery pressing directly on Known Weakness #2. The framework is not contradicted — it gave a coherent verdict — but the Switch sits in the exact gap between "we do not aggregate apertures" and "five symmetric deaths are still worse than one," and the framework's stated apparatus does not cleanly deliver the second without borrowing from the thing it refuses.

The pairing of these two results is the framework's actual signature: more restrictive than consequentialism at the instrumentalization boundary (Loop), and — as the next cluster shows — more demanding than common sense about beneficence. Both flow from the same anti-aggregation, geometry-respecting core, which is why the framework is recognizably neither utilitarian nor strictly Kantian.


Duties to aid across distance: the Singer and Unger cases

Peter Unger's Vintage Sedan and Peter Singer's Bugatti / drowning-child variations are consistency traps of a different shape. They pair a near case (a bleeding hitchhiker you could drive to hospital, at the cost of bloodied upholstery; a child drowning in front of you) with a far case (the Envelope you could mail to save distant children) and ask why the verdicts differ. The near case reliably reads as "you behaved badly if you refused"; the far case reliably reads as permissible to decline. The trap is that the only salient variable is distance, which is hard to defend as morally decisive.

The framework agrees that distance is not decisive, and it agrees for a principled reason: the badness of a stake-collapse is constitutive of the collapsing geometry, not indexed to the geometry's spatial relation to any observer. A distant child's death is the same total collapse as a near one. Reach is a dimension of the agent's stakes — how far the agent's own mattering extends — but it does not make the distant victim's collapse less bad; it makes the agent less likely to register it. So the framework, like Singer and Unger, holds that the common near/far asymmetry tracks a heuristic rather than a moral fact, and it can say which heuristic: recognition-fidelity degrades with mediation and distance (the Relations Between Apertures modifier), so we vividly recognize the apertures in front of us and dimly recognize the ones on a screen. The verdict-flip is the recognition system doing what it evolved to do, not a defensible moral distinction. This is the same diagnostic move the incest analysis makes — name the heuristic, separate it from the structural fact — applied to proximity rather than to disgust.

But the framework does not follow Singer all the way to his demandingness conclusion, and the reason is the same anti-aggregation commitment that restrains it at the Loop. Because density is not a quantity to maximize, the framework does not require giving until the marginal sacrifice equals the marginal benefit. Its positive duty is prevent collapse where you can without collapsing your own geometry — bounded by the protection of the agent's own stakes, which are also a geometry, and "an ethics that demands its destruction to serve another's is doing the negative thing it claims to oppose." On the Sedan itself the verdict is clear and matches intuition: refusing to drive the bleeding man to save upholstery sets a trivial stake against a total collapse, and proportional weighting (the same lever that, in Animal Ethics, makes a trivial human interest yield to a severe animal one) condemns the refusal outright. You behaved badly.

The honest complication is the Envelope. Here the framework's anti-demandingness escape hatch does not obviously apply, and this is worth flagging as a place consistent application bites. The vet-case bound was built for sacrifices that would destroy the agent's capacity — close the practice, collapse the giver. A modest, painless donation that would save a distant child threatens no such thing. So the framework, applied consistently, appears to endorse a duty of beneficence for the affluent that is more demanding than common sense — not Singer's unlimited demand, but well past the comfortable "charity is optional." The framework can locate why the Envelope feels lighter than the Sedan (mediation, uncertainty about efficacy, the diffusion of the proxy duty across many possible givers, the severed recognition of the distant aperture), and it can credit those as real reductions in individual weight. What it cannot do, consistently, is reduce that weight to zero. The result is a substantive commitment the project has not previously owned: a strong, bounded, but genuinely demanding positive duty to prevent distant collapse, with the bound set by the agent's own geometry rather than by distance. Some readers will take that as a bullet. It is at least a consistent one, and it is the mirror of the Loop result — the framework is demanding where consequentialism is permissive (beneficence) and restrictive where consequentialism is permissive (instrumentalization), both for the same structural reason.


The harmless-taboo cluster: the incest analysis, re-run and confirmed

Would You Eat Your Cat?, the taboo items in Morality Play, and several Consent-adjacent provocations are all built on the same engine as Haidt's incest case: present a taboo act with every avenue of harm stipulated away (the pet died naturally; the act is private, consensual, and harms no one), and watch the respondent insist on wrongness while failing to locate it. The framework already has a fully worked treatment of this engine in the Moral Sentimentalism / Incest document, and the most important thing this battery shows is that the treatment generalizes without modification. The framework gives the same answer to cat-eating, to the harmless private acts in Morality Play, and to consensual incest, and gives it for the same reason: a disgust or purity reaction is a Tier-1 gut heuristic (operative, gut-homed, firing ahead of deliberation), and the framework's question is whether the act gratuitously collapses a felt stake-geometry. Where the scenario has been engineered so that no geometry collapses, Tier 2 falls silent, and the framework bites the bullet — no structural wrongness — while explaining the persistent feeling as a well-calibrated detector firing into a vacuum, and defending the blunt Tier-1 rule as worth keeping for its track record across the un-engineered cases.

That the verdict is identical across three superficially different taboos is the consistency result. The cat case adds one wrinkle the framework handles cleanly: a pet that has already died is no longer an aperture, so there is no geometry to collapse in the eating itself; whatever residual wrongness a respondent feels attaches to the living owner's relational geometry (the grief-bearing dimension constituted by the now-absent aperture, per the past apertures note), not to the cat. The framework thus distinguishes "wrong because it harms a stake-bearer" from "felt as wrong because it violates a relation the survivor still carries," and only the first is a Tier 2 matter. This is the harm/taboo separation the experiments are built to elicit, delivered structurally.

The known cost carries over too, and should be restated rather than hidden: many readers will regard "no structural wrongness in the stipulated case" as the framework failing the case rather than diagnosing it. That exposure is identical to the one the incest document already concedes, and the battery does not worsen it — it confirms that the framework pays the same price consistently wherever the harmless-taboo engine is run, which is at least the honest behavior of a theory rather than the ad hoc behavior of an intuition.


Consent and bodily autonomy: authority over one's own geometry

The Consent Experiment and Whose Body Is It Anyway? test whether a respondent's commitment to autonomy survives across cases — whether someone who permits dangerous consensual sport also permits voluntary organ sale, dangerous private choices, and the rest, or whether they smuggle in a purity residue under autonomy's banner. The framework enters this cluster with a strong default and a narrow exception, and the pairing is what lets it stay consistent.

The default is near-maximal respect for an aperture's authority over its own stake-geometry. Coercion is a paradigm collapse precisely because it "narrows the dimensions along which a person can act," and the framework's refinement of the Platinum Rule warns explicitly against overriding a person's stated stakes — the paternalism problem is named as a danger, not a license. So the framework's baseline across the whose-body items is permissive: the aperture is the authority on its own stakes, and a bystander's discomfort is not a structural reason to override. The symmetry principle then does the consistency enforcement directly: if a respondent permits boxing but forbids voluntary organ sale, the framework demands a structural difference between the cases, and will not accept a difference that turns out to be residual disgust. A real structural difference is available in some such pairs — organ markets may involve exploitation and compromised consent under economic duress, which is a genuine difference about whether the consent is free — but the framework forces that difference to be named and defended rather than assumed.

The narrow exception is the framework's one resource for not deferring to a stated preference: where the preference is itself a product of stake-collapse — addiction, coercion, ideological capture — the framework reads it as "what their well-functioning stake-geometry would call for," not "whatever they currently report." This is what keeps the framework from collapsing into pure preference-satisfaction. But it is also precisely where the cluster presses on a known soft spot. The framework holds two commitments that pull against each other: respect the aperture's authority over its own stakes and do not honor preferences issued by a collapsed geometry. The boundary between them is, in the framework's own words, "dangerous," and "historically the second has been dressed in the language of the first." The whose-body experiments do not contradict the framework, but they force it to operate exactly on this seam, and the framework's apparatus under-determines where the line falls. This is the autonomy/paternalism indeterminacy, re-pressed in concrete cases, and it is a genuine open problem the battery makes vivid rather than a contradiction it manufactures.


Reciprocity, desert, and fair dealing

Tit for Tat and the Philosophical Health Check / Antiques items test the consistency of judgments about deserved treatment, retaliation, fair exchange, and honesty under role-reversal — whether a respondent's fairness principle survives when they switch from buyer to seller, or from the party owed retaliation to the party against whom it is directed.

The framework passes the reciprocity and desert items largely by construction, because treat symmetrically situated apertures symmetrically is the demand these experiments enforce, and it is Tier 2 doctrine. The framework also brings a prior critique of pure reciprocity from Applied Ethics: the Reciprocity Rule ("treat others as they treat you") makes one's stake-geometry "parasitic on others' behavior," hostage to the worst actor in the environment — a coherence-collapse, structurally the same defect it diagnoses in envy. So the framework does not merely pass the tit-for-tat consistency check; it explains why strict reciprocity is a poor principle to be consistent about. Desert and fairness it locates as Tier 1 network goods (Justice as the canonical network-bearer good), binding within the network and enforced by the symmetry principle, without inflating them to Tier 2 — which is the right result, since "what is deserved" varies with network agreements in a way "don't gratuitously collapse a geometry" does not.

The retaliation and punishment items land the framework on its hardest open political problem, named in Political Philosophy: punishment "deliberately imposes constriction on an aperture, which the framework treats as prima facie bad," and any justification must show it preventing greater collapse rather than merely answering harm with harm. So the framework treats tit-for-tat retaliation as structurally suspect by default — the desire to see a wrongdoer suffer is not, on its own, a collapse the retaliation prevents — and demands a forward-looking justification. Consistent, and consistent with the foundational documents, but resting on an acknowledged unfinished account of punishment.

The Antiques fair-dealing items draw a clean distinction the framework is well-equipped for: between deceiving and failing to disclose. Active misrepresentation of an item's worth violates the Tier 2 principle "don't deceive other apertures about the contents of their own." Buying an undervalued item without lying is not deception, and the framework does not manufacture a Tier 2 wrong where none of its principles is violated. What remains is a relational and Tier-1 question about exploitation — using an information or power asymmetry to extract from another a stake they would have kept had they participated in the decision on equal footing. That is structurally adjacent to the framework's analysis of theft ("taking from an aperture something they had stakes in without their participation"), but not identical, because the seller did participate — they chose to sell. So the framework grades the antiques cases rather than flattening them: lying about worth is a Tier 2 violation; silent exploitation of asymmetry is a Tier-1 fairness matter whose weight depends on the relation (arm's-length transaction versus fiduciary trust); and a fair-and-informed bargain is no wrong at all. The role-reversal trap fails against this because the grading is symmetric — it gives the same verdict whether the respondent is buyer or seller.


The metaethics tests: Euthyphro and the Disposition battery

The two borderline experiments are the most explicitly philosophical, and the framework's performance on them is asymmetric: on the Euthyphro it is at its strongest in the entire battery, and on the Disposition test it both passes and contributes a refinement.

The Euthyphro dilemma asks whether God can will that torturing innocents for pleasure is good and just. It is built to force a respondent who has affirmed divine sovereignty, omnipotence, and omnibenevolence into one of two horns: either torture would be good if God so willed (morality is arbitrary divine fiat) or it would not (there is a standard independent of God, and God is not the source of morality). The framework dissolves the dilemma rather than choosing a horn, and it does so straight out of existing doctrine. Divine-command morality is a Tier 3 claim — a value holding independent of any network, true in a universe with no consciousness — and Tier 3 is empty by construction. The badness of torturing an innocent is constitutive of the victim's collapsing geometry at Tier 0; it is what the felt loss of mattering is, registered from inside, and no act of will, divine or otherwise, can make a constitutive badness not-bad, any more than a will can make a circle have no ratio. So the framework answers the dilemma's final question with a firm no — God cannot make torture good — and crucially this no does not impale it on either horn. It is not the Platonic horn (an abstract standard floating free of God in Tier 3), because the standard is not free-floating; it is anchored in the existence of any conscious sufferer. It is not the arbitrariness horn, because the good is not whatever is willed. The framework offers a third position the dilemma did not anticipate: value is grounded in the structure of felt stakes — neither divine fiat nor mind-independent Form — and an omnipotence properly understood does not include the power to make a felt collapse not-bad-for-the-sufferer, because that is not a coherent thing to do rather than a hard one. This is the framework's cleanest victory in the battery, and it is worth noticing that the experiment the project flagged as "borderline / out of scope for moral testing" turns out to be the one the framework's metaethics is most decisively built for.

The Disposition test is not a first-order moral instrument; it measures whether a respondent applies the epistemic standard "concede the specific claim when your empirical argument is conclusively refuted" symmetrically across political sides, rather than conceding when an opponent's argument falls and inventing an escape when their own does. The framework's relevant resources are the symmetry principle (applied now to arguments and standards rather than to apertures) and the responsiveness dimension, whose failure mode — organizing one's geometry around a doctrine and refusing to update on evidence — the framework already names as ideological capture, the structural cousin of pride. So the framework passes the disposition test by construction: it is committed to applying the concede-when-refuted standard regardless of which side benefits, and it has a diagnosis ready for the failure (responsiveness-collapse / ideological capture) and a name for why it is a self-regarding harm.

But the framework also refines the test, and this is a small genuine contribution. Each scenario offers Option A (concede the refuted empirical claim) and Option B ("the real commitment was never that empirical claim, so it survives"). The test treats Option B as evasion. The framework's tier structure and its descriptive/normative distinction show that Option B is sometimes legitimate: a moral commitment can be a genuine Tier 0/1/2 value that never reduced to the contested empirical prediction, in which case the refutation of the prediction leaves the value standing. The discriminating question is whether the empirical claim was originally advanced as the argument or was always a proxy for an undisclosed value. Where it was advanced as the argument, the responsiveness principle plus the don't-deceive principle (extended to self-deception) require concession; clinging to the refuted claim is ideological capture. Where it was always a proxy, Option B is honest only if the respondent now defends the value openly rather than continuing to relitigate the dead empirical claim. The framework thus turns the test's binary into a principled three-way sort — concede; or retreat to the value and own it as a value; or, illegitimately, keep the refuted empirical claim alive for motivated reasons — which is a more discriminating instrument than the test itself supplies, and which the framework's symmetry commitment then requires be applied identically to every side.


Can the framework answer the battery consistently?

Yes, and the consistency is the principal finding. One engine — felt stakes as the locus of value, collapse and elaboration as the moral variables, the tier sort, the symmetry and don't-instrumentalize and proxy and don't-deceive principles, the felt/functional line, and the relational modifiers — handles all fourteen experiments, and the verdicts cohere across them. The harmless-taboo verdict is identical for incest, cat-eating, and the Morality Play items. The trolley verdict, the lifeboat verdict, the Playhouse intention/foresight verdict, and the Carneades verdict are all generated by the single instrumentalization analysis. The Sedan and the Envelope are decided by one proportional-weighting lever. The Euthyphro is decided by the same empty-Tier-3 commitment that decides the framework's response to the grounding skeptic's cosmic challenge. This is what it looks like for a theory to have a structural core rather than a list of intuitions: the cases that fracture surface-feature reasoning do not fracture it, because the surface feature each experiment isolates was never the framework's reason.

One caution belongs next to that result, in the framework's own spirit. Passing a consistency battery is not the same as being true. A theory can be perfectly consistent and still rest on the is/ought residue the framework openly concedes it cannot close. The experiments test coherence, and coherence is necessary, not sufficient. The framework's strong showing here is evidence that its principles hang together and apply evenly; it is not, and the framework should not be read as claiming it is, independent confirmation that felt-stake-collapse is what wrongness consists in. That deeper claim is carried by the metaethical argument and admitted to be persuasive rather than airtight.


Are the trick questions driving an unavoidable conflict?

This was the live question, and the answer is no — but the no is specific and worth stating precisely, because it is not a complacent one.

No experiment in the battery forces an internal contradiction: two framework principles delivering opposite verdicts on a single case with no resolution. What the hardest experiments do instead is one of two things. Either they press the framework onto a problem it has already named as open — and there the framework gives a partial, honest answer and points at the gap — or they force it to bite a bullet it has already committed to biting, consistently. Three pressure points and one near-miss are worth naming.

The aggregation problem (Known Weakness #2) is pressed hardest by the Switch and by the Singer Envelope. The Switch wants a mild "five rather than one" aggregation the framework officially disavows; the Envelope wants a stopping-rule for beneficence the framework cannot precisely supply. The framework reframes both away from summation (allocate an unavoidable harm; bound by the agent's own geometry), but the reframings are thin exactly where the numbers and the modest cost are doing intuitive work. This is the battery's sharpest contact with an open problem, and it is real — though it is re-exposure, not a new defeat.

The autonomy/paternalism line is pressed by the Consent and Whose-Body items, which force the framework to operate on the seam between respecting an aperture's authority and refusing to honor a collapsed geometry's stated preferences. The framework's apparatus under-determines the line, and it admits the line is dangerous. Again: a known soft spot made vivid, not a contradiction.

The is/ought residue is pressed by the deepest reading of the Disposition test and by anyone who pushes the Euthyphro to Reading C ("so what if collapse is bad inside every geometry — why does that bind from nowhere?"). The framework concedes it cannot compel the determined skeptic, but locates the concession precisely: the cosmic standard the skeptic demands is empty for every ethics, and the burden flips onto anyone claiming to have exhibited one. The Euthyphro is where this concession is least costly, because the question there is whether an external will can override constitutive badness, and the answer to that is a clean no.

The near-miss — and the one genuinely new structural finding — is the Carneades plank, treated in the next section.

So the trick questions do not drive the framework into unavoidable conflict. They were built to drive surface-feature reasoners into conflict, and against a structural-feature theory they instead function as a map of where that theory's honest open problems lie. That the same three problems recur across a battery of independently designed experiments is itself mild corroboration that the framework has correctly identified its own weak points rather than having hidden others.


What is new, and what is challenged

Four things came out of the exercise that the foundational documents did not already contain in worked form.

A means-based, non-DDE grounding of the trolley verdicts, made explicit. The framework distinguishes Switch from Footbridge from Loop by instrumentalization (forced geometric simplification / treating-as-mere-means), not by the act/omission or flip/push surface, and so verdicts the Loop with the Footbridge. This was latent in Stake-Geometry's dignity section but never run against the trolley series. Its payoff is clarifying: it shows the framework is not consequentialist despite its collapse-minimizing language, and it pins down where the framework is deontological-leaning (the instrumentalization boundary) versus where it is allocation-minimizing (symmetric forced choices). The Loop is the precise point where those two pulls collide inside the framework, and the framework's resolution in favor of the prohibition is a deontological bullet it should state openly.

A latent agent-relativity in the grounding, surfaced by Carneades. This is the most interesting new finding. The Grounding Skeptic establishes an asymmetry: the badness of your own collapse is constitutive and given (Reading A), while your reason to weight another's collapse is persuasive-but-not-airtight (Reading B). The plank case forces this asymmetry to do practical work. Pushing the other sailor off to survive instrumentalizes a symmetric innocent geometry — which the framework's symmetry and don't-instrumentalize principles forbid, so the act is not justified. Yet the framework also grounds your own survival-stake more firmly than your reason to weight the other's, which makes the act excusable in a way it can articulate rather than merely assert. The framework thereby reconstructs the law's justification/excuse distinction structurally — but only by leaning on an agent-relativity that the foundational documents have not made explicit. There is a real tension here that the battery brings to the surface: the consistency argument at Step 4 pushes toward agent-neutrality ("no relevant difference between your geometry and another's"), while the A/B grounding asymmetry pushes toward agent-relativity (your own stakes are grounded more firmly than your reasons to weight others'). The framework can hold both — agent-neutral about what is valuable, agent-relative about the firmness of grounding and the reasons that flow from it — and that combination is exactly what it needs to make sense of self-preservation (Carneades) and of the bounded demandingness that stops it short of Singer. But it is currently implicit, and the cases show it needs an explicit treatment. This is the clearest place the exercise generates new work for the framework rather than merely applying it.

A demandingness commitment the project had not owned. Consistent application to the Singer Envelope yields a strong, bounded, but genuinely more-than-common-sense duty of beneficence for the affluent, because the anti-self-immolation bound does not cover modest painless giving. The framework's character across the battery is therefore more demanding than consequentialism about beneficence and more restrictive than consequentialism about instrumentalization — a single profile flowing from the anti-aggregation core. This is a substantive position worth stating in the framework's own voice.

A refinement of the Disposition test's binary. The framework's tier structure plus its descriptive/normative distinction turn the test's A/B forced choice into a principled three-way sort (concede; retreat to an openly-owned value; or illegitimately preserve a refuted empirical claim), applied symmetrically by the symmetry principle. Modest, but a genuine sharpening of the instrument.

On the challenge side, nothing in the battery defeats the framework, but three pressures are real and have been named: the aggregation gap (sharpened by being shown to bite in first-order life-and-death cases, not only in abstract population ethics); the autonomy/paternalism indeterminacy; and the demandingness bullet, which some readers will not accept. None is new in kind. The exercise's honest summary is that a battery built to manufacture contradictions found the framework consistent and instead returned a precise inventory of its known fault lines — plus one structural seam, the latent agent-relativity, that the framework now has reason to develop explicitly.


Known weaknesses of this analysis

The instrumentalization line is sharp in the stylized cases and blurry in real ones. Switch, Footbridge, and Loop are engineered so that "is this geometry the mechanism of rescue?" has a clean answer. Real cases — collateral harm in war, risky rescue, economic policy that foreseeably kills — do not come pre-sorted into means and side-effects, and the framework's verdict inherits all the difficulty the means/side-effect distinction has always carried. The trolley performance should not be over-read as a general solution.

The forced-choice/instrumentalization reconciliation is asserted more than derived. The framework leans toward diverting in the Switch (allocate unavoidable harm) and against pushing in the Footbridge (don't instrumentalize), and treats the Loop as instrumentalization. But the deep question — why a symmetric forced allocation is permissible at all if apertures cannot be aggregated — is answered by reframing rather than by a derivation, and the reframing is exactly as strong as the unsolved aggregation problem allows, which is to say not fully.

The agent-relativity finding is a diagnosis, not yet a theory. Naming that the framework needs an explicit account of how agent-neutral value coexists with agent-relative grounding-firmness is progress, but the account itself is not given here. Until it is, the Carneades verdict (not justified, but excusable) rests on a distinction the framework gestures at rather than grounds.

Several experiments were analyzed at the level of their structure rather than item by item. The Morality Play (19 items), Lifeboat (6 items), Consent (3 principles + 3 cases), and Tit for Tat (8 items) batteries were treated by the machinery their items share, not exhaustively. A fuller pass might find an individual item that separates principles this analysis treated as moving together.


Intellectual lineage


This document records an applied stress-test developed through dialogue. It is exploratory, not peer-reviewed, and is offered as a contribution to thinking about how an aperture-based ethics behaves under thought-experiments designed to detect inconsistency.