Moral Dilemma Pilot: Evidence Release
Six open models, ten dilemmas, two framings, three samples per cell. Every raw record is published; the chart is computed from those records.
Methods, data, limitations and evidence status, published together.
Six open models, ten dilemmas, two framings, three samples per cell. Every raw record is published; the chart is computed from those records.
Forty publicly sourced cases coded on a fixed two axis rubric, every score's rationale published in the data.
Do internal signatures predict how models play? Behavioural fleet measured; latent analysis designed, not yet run.
What evidence would show a model represents evaluation as an internal state? Probing and steering protocol, designed.
Separating persistence caused by a model from persistence caused by the world around it, with falsification criteria.
Per game family: what is measured, the control required, and the claim the evidence never licenses.
How lenders might price credit risk on GPU collateralised debt.
What a functioning forward market for GPU compute would look like.