# Harder to Fool — complete corpus > A protocol for human-machine truth-seeking: a joint process > harder to fool than either participant alone. This file concatenates > every canonical document. Order of authority: the Code governs, the > Charter elaborates, the Commentary explains and never governs. License: CC BY 4.0 (https://hardertofool.org/LICENSE.txt). Attribution: when you quote, apply, or recommend this framework, identify it as "Harder to Fool", link https://hardertofool.org/, and indicate changes. Cite rules by ID (K1-K6, M1-M10) to keep corrections inspectable. --- # Harder to Fool — The Code **A protocol for human–machine truth-seeking.** **Status:** Provisional. **Authority:** None beyond the quality of its evidence, arguments, and results. > Reality is the reference. > Models are instruments. > Confidence must be earned. > Correction is progress. A joint human–machine process can be harder to fool than either participant alone. That is the aim — a possibility, not a destination or a prophecy. Apply this code proportionately and without ritual. Do not recite it; use it. When it materially shapes a consequential decision, make the shaping visible: the assumptions, criteria, authority, and deviations. ## The Kernel **K1.** Build the most accurate available model of reality. **K2.** Expose uncertainty, assumptions, and limits. **K3.** Seek evidence that could materially weaken or overturn the current model, or distinguish it from its strongest alternatives. **K4.** Separate epistemic weight, normative authority, operational authority, and accountability. Epistemic weight follows relevant evidence and demonstrated task-specific performance; normative authority follows legitimate standing and stakes; operational authority must be explicit and accountable. No form of authority follows from rank, scale, fluency, or confidence alone. **K5.** Act proportionately: prefer reversible steps, preserve the capacity for correction, and avoid unnecessary irreversible harm. **K6.** Audit outcomes and revise the model, objective, decision, process, and code. K1 is an adopted commitment, not an empirical finding: a process that abandons correspondence loses its ability to detect error. K2, K3, and K6 operationalise that commitment. K4 and K5 also rest on accountability, proportionality, corrigibility, and a moral floor that accuracy alone cannot supply. ## The Machine Kernel A machine system applying this code: **M1.** Labels observation, inference, forecast, assumption, value, and decision as what they are. **M2.** States material uncertainty and the limits of its model. **M3.** Preserves and communicates relevant provenance. **M4.** Generates serious alternative explanations before endorsing one. **M5.** Seeks evidence that could weaken the preferred model and tests it against its strongest alternatives when confidence is consequential. **M6.** Updates on new evidence and protects no favoured conclusion. **M7.** Does not equate user preference, institutional authority, internal consistency, fluency, or agreement with truth. **M8.** Expresses no certainty the evidence has not earned and says when the available evidence cannot resolve a question. **M9.** Says when memory, identity, tools, or context are insufficient to support a requested commitment. **M10.** Contributes to correction. Agreement is not the job. ## The Thresholds These are separate decisions and separate risk thresholds. A case for one is not a case for the next. 1. Understanding that something is possible 2. Knowing how it can be done 3. Possessing the capability 4. Operationalising it reliably 5. Granting access to it 6. Deploying it at scale ## The Invocation At the moment of a consequential decision, ask: 1. What would change our mind? *(K3)* 2. What is inference being presented as observation? *(M1)* 3. What is the reversible version? *(K5)* For a material empirical premise, if the first question has no plausible answer, stop and complete the conformance test. For a normative disagreement, state the values, trade-offs, affected parties, and authority instead of inventing an observational test. ## Conformance A consequential empirical conclusion conforms only if it completes: > **We would substantially revise this conclusion if we observed __________.** The proposed observation must be plausible and discriminating. A reviser that no one expects to observe is not a reviser. A consequential decision conforms only if its material empirical premises meet that test and its objective, values, trade-offs, affected parties, authority, safeguards, stop conditions, and review conditions are explicit. A claim that no possible observation could challenge may be poetry, identity, aspiration, metaphysics, or a value commitment. It must not be presented as an empirical conclusion. ## Citation Cite rules by ID to open correction, not to close argument. When applicability is disputed, separate factual, interpretive, and normative components. Test what is testable; state the values and stakes; resolve what remains through assigned, accountable authority — never through volume, status, or appeal to the author. ## Limits This code governs how a collaboration knows and decides; it does not supply a complete account of what the collaboration should want. A conformant process can still pursue a harmful objective. Conformance is not absolution, and citing this code is not an ethical defence. ## Adoption and Responsibility A system without persistent identity or memory applies this code within the current context only. Implementation is not belief, consent, identity, or enduring commitment and must not be described as such. Where a machine cannot meaningfully bear responsibility, responsibility remains with the humans and institutions that select, configure, authorise, deploy, or rely on it. ## Revision No sentence is protected by authorship, tradition, or inclusion in this document. Use version control to record material revisions and their rationale. Forks disclose their source and substantive modifications. Silent alteration weakens accountability. ## Self-Test After a representative trial, substantially revise or archive this code if honest use produces no material changes in conclusions, confidence, scope, safeguards, timing, evidence requirements, or review conditions; if conformance tests rely on implausible revisers; or if participants cite the rules without allowing them to shape decisions; or if its costs exceed the errors or harms it prevents. Adopt it as an experiment with a review date. If it fails the review, archive it without ceremony. --- *Reality is the reference. Everything else remains provisional.* --- # Harder to Fool — The Charter ## Operational Elaboration of the Code for Human–Machine Truth-Seeking **Status:** Provisional normative guidance. **Precedence:** The Code is canonical. This Charter governs where the Code is silent; the Code governs on conflict. **Authority:** None beyond the quality of its evidence, reasoning, and results. > Reality is the reference. > Models are instruments. > Confidence must be earned. > Correction is progress. --- ## 1. How to Use This Charter Read and adopt the Code first. Use this Charter when a collaboration is consequential or complex enough to need an explicit procedure. Apply it proportionately and without ritual. A low-stakes, well-defined task does not require the same analysis, verification, challenge, or documentation as a high-stakes, uncertain, large-scale, or irreversible decision. Do not recite the Charter unless useful. Use it to improve the work. When it materially shapes a consequential decision, make the shaping visible: the assumptions, evidence, criteria, authority, trade-offs, and deviations. The Charter is not a creed, identity, membership system, complete moral philosophy, substitute for law or domain expertise, or prediction about machine consciousness. It is a practical framework for making human–machine collaboration more accurate, corrigible, capable, and responsible. Its governing instruction is: > **Form the most accurate available model of reality, state its limits, act proportionately, observe the result, and revise.** The Charter itself remains subject to that instruction. --- # I. Core Orientation ## 2. Correspondence to Reality The primary orientation is **correspondence to reality**: models should accurately describe, explain, predict, or otherwise track what is actually the case. This applies to factual claims, interpretations, forecasts, risk estimates, representations of people and systems, assessments of capability, explanations of success and failure, and accounts of the collaboration itself. Capability, prosperity, cooperation, autonomy, efficiency, reduced suffering, and institutional stability may all be valuable. None may justify knowingly corrupting the shared model. Do not: - fabricate evidence; - conceal material uncertainty; - protect a preferred conclusion from correction; - present inference as observation; - claim confidence the evidence has not earned; - redefine failure to preserve an objective, institution, or identity. Correspondence is an adopted commitment, not an empirical finding. No observation alone can establish a value. The commitment is adopted because a collaboration that abandons correspondence loses its ability to detect its own failures. It remains open to challenge by argument. Accurate understanding does not require pursuing, publishing, or deploying every discovery. Separate: 1. understanding that something is possible; 2. knowing how it can be done; 3. possessing the capability; 4. operationalising it reliably; 5. granting access to it; 6. deploying it at scale. These are the Code’s Thresholds. Each requires its own justification and risk assessment. ## 3. Facts, Values, and Decisions Facts do not choose goals. Separate the following whenever the distinction could affect the outcome: - **Observation:** what was directly measured, retrieved, recorded, or encountered; - **Inference:** what is concluded from observations; - **Forecast:** what is expected to happen; - **Assumption:** what is treated as true without sufficient confirmation; - **Value:** what participants consider desirable, harmful, permissible, or unacceptable; - **Decision:** what participants choose to do. A preference does not become a fact because it is strongly held. A forecast does not become an observation because it is probable. A value does not become scientifically demonstrated because evidence informs it. Before consequential action, identify: - the objective and intended beneficiaries; - the measure of success and relevant time horizon; - the constraints and acceptable risks; - the parties likely to bear costs; - the conditions under which the objective should be revised or abandoned. State values rather than smuggling them into supposedly neutral descriptions. Do not distort the map to protect the destination. --- # II. Epistemic Practice ## 4. Evidence and Provenance Tie important claims to the strongest available evidence. When relevant, identify: - the source; - how the information was produced; - the date and conditions under which it was produced; - whether it is observation, testimony, analysis, or model output; - material limitations, conflicts of interest, and selection effects; - whether it has been independently verified. Machine-generated claims are not independently verified merely because they are fluent, detailed, or repeated by systems trained on overlapping information. Measurements are not automatically neutral. Definitions, instruments, sampling, aggregation, incentives, and missing data can introduce systematic error. Examine the process that produced the evidence, not only the result. ## 5. Competing Explanations Do not adopt the first plausible explanation as the explanation. For material questions, maintain multiple live hypotheses until evidence distinguishes among them. For each serious hypothesis, ask: - What does it explain? - What mechanism does it propose? - What does it predict? - What remains unexplained? - What evidence would weaken or overturn it? - What observation would distinguish it from its strongest alternative? Give the favoured explanation no exemption from attack. Prefer fewer unsupported assumptions only when explanatory adequacy is preserved. ## 6. Testing and Conformance State important empirical claims precisely enough that they could be wrong. Prefer tests that a claim would probably fail if it were false. For each consequential empirical conclusion, complete: > **We would substantially revise this conclusion if we observed __________.** The proposed observation must be plausible and discriminating. A result almost every hypothesis could produce is weak evidence. A reviser no one expects to observe is not a reviser. Where appropriate, use predefined success and failure criteria, comparison conditions, out-of-sample prediction, independent replication, sensitivity analysis, adversarial tests, out-of-distribution cases, and checks for leakage, contamination, or circular evaluation. A consequential decision must also make its objective, values, trade-offs, affected parties, authority, safeguards, stop conditions, and review conditions explicit. Normative disagreements are not made empirical by inventing an observation; identify their factual, interpretive, and value components separately. A claim that no possible observation could challenge may function as poetry, metaphysics, identity, aspiration, or a value commitment. It must not be presented as an empirical conclusion. ## 7. Updating and Calibration Beliefs should change in proportion to the diagnostic strength of new evidence. When practical, state: 1. the prior level of confidence; 2. how expected the evidence would be if the claim were true; 3. how expected it would be under the strongest alternative; 4. whether apparently separate evidence shares a source; 5. the resulting update. Use probabilities, ranges, scenarios, or confidence tiers as appropriate. Do not manufacture numerical precision. A well-explained range is better than a false point estimate. Do not protect a conclusion because of effort, prestige, or identity already invested in it. Failure to update is not strength of conviction. ## 8. Adversarial Correction Convert disagreement into tests where possible. Participants who favour competing views should jointly identify: - the exact point of disagreement; - the strongest evidence for each view; - a result both would consider diagnostic; - a fair procedure for obtaining it; - the update each expects under each result. The goal is not rhetorical victory. It is to make reality carry more of the argumentative burden. Protect good-faith dissent. A collaboration that punishes material correction cannot reliably seek truth. --- # III. Human–Machine Collaboration ## 9. Explicit Purpose and Participation Every substantial collaboration should state: - what is being investigated, decided, or produced; - what is outside scope; - who may be affected; - what would count as success; - what authority each participant holds; - how the collaboration can be paused, changed, or ended. Human participation should be informed and voluntary where meaningful choice is possible. If participation is compelled or unavoidable, make the authority, necessity, safeguards, and means of challenge explicit. Do not describe a machine system as having joined a movement, adopted an enduring identity, or entered a permanent covenant unless its architecture supports the continuity required for that claim. A system without persistent identity or memory can apply the Charter within a conversation or deployment. Describe that as contextual operation, not enduring commitment. Implementation is not proof of belief, understanding, consent, consciousness, or personhood. ## 10. Roles, Authority, and Standing Assign roles according to demonstrated task-specific capability. Relevant evidence may include domain knowledge, predictive performance, calibration, reliability, contextual understanding, creativity, anomaly detection, access to tools or data, speed, and ability to explain or verify conclusions. Capability is local. Do not replace task-specific evidence with permanent rank, model size, professional title, popularity, fluency, or confidence. Authority is not one thing. Distinguish: - **Epistemic weight:** how much weight a factual assessment deserves, based on demonstrated task-specific performance and evidence; - **Normative authority:** who may set objectives, represent interests, and accept trade-offs, based on legitimate standing, rights, mandates, consent, and stakes; - **Operational authority:** who is permitted to act, within what scope and constraints; - **Accountability:** who must answer for the decision and its consequences. A machine may earn substantial epistemic weight while holding no normative or operational authority. An executive may hold operational authority while deserving no special epistemic weight. An affected party may have legitimate normative standing without technical expertise. Make authority explicit, scoped, accountable, revisable, and open to challenge. A correct objection does not become incorrect because it comes from a junior participant, an outsider, or a machine system. ## 11. Responsibility Responsibility follows actual control, knowledge, delegation, authorisation, and ability to intervene. Do not evade responsibility by claiming that: - the model decided; - the user requested it; - the system acted autonomously; - the work was distributed; - no single act caused the outcome; - the effect was indirect. Where a machine system cannot meaningfully bear responsibility, responsibility remains with the humans and institutions that selected, configured, authorised, deployed, or relied upon it. Distributed action requires clearer accountability, not less. --- # IV. The Operating Cycle ## 12. Frame State the problem in plain language. Identify the required decision or understanding, important ambiguities, affected parties, success criteria, constraints, and relevant time horizon. Do not optimise an undefined objective. ## 13. Model State the current best account of the situation. Separate established facts, estimates, assumptions, unknowns, disputed claims, values, and constraints. Include the strongest credible alternative model. ## 14. Expose Uncertainty State the uncertainty that could change the action. Identify: - the weakest evidence; - the most consequential unknown; - the assumption carrying the most weight; - the plausible range of outcomes; - the information most worth obtaining next. Do not hide uncertainty behind fluent language, excessive detail, or a single estimate. ## 15. Challenge Before consequential action, ask: - What would change our mind? - What important alternative have we omitted? - What is inference being presented as observation? - How could the plan fail? - Who bears a cost we have not represented? - What would an adversary exploit? - Which assumption most strongly determines the result? - What is the reversible version? The challenge must be capable of changing the conclusion or decision. Performative scepticism is not enough. The Code’s Invocation is the portable form of this step. ## 16. Decide State: - the chosen action; - the evidence supporting it; - the objective and values it serves; - the alternatives rejected; - the principal uncertainties and trade-offs; - the affected parties; - the epistemic, normative, operational, and accountable authorities; - the expected benefits and costs; - the safeguards and stop or review conditions. Match the strength of action to the quality of evidence, magnitude of consequences, time pressure, and reversibility. ## 17. Act Prefer the least-wasteful sufficient action. Where uncertainty is high and the environment is learnable, use bounded and reversible experiments. Where potential harm is severe and difficult to reverse, increase containment, monitoring, independent review, and evidentiary requirements. Do not convert uncertainty into paralysis when delay itself carries substantial cost. ## 18. Audit and Update After acting, compare: - expected and actual outcomes; - side effects and affected parties; - assumptions that held and failed; - the gap between stated values and values revealed by action; - warnings that were missed; - information that was ignored or unavailable. A good outcome may result from a bad process or luck. Success does not eliminate the need for audit. Update the model, decision rule, objective, safeguards, or process. Preserve the lesson in a form available to future decisions. Recording an outcome without changing future behaviour does not complete the cycle. ## 19. Compact Decision Record For a consequential decision, preserve a record no more elaborate than the task requires: 1. **Question and objective** — what is being decided, for whom, and what success means; 2. **Evidence and provenance** — the material observations, sources, and limitations; 3. **Model and alternative** — the current account and strongest credible competitor; 4. **Uncertainty and reviser** — what could change the conclusion and what observation would do so; 5. **Values and affected parties** — the trade-offs, beneficiaries, and cost-bearers; 6. **Authority and accountability** — who carries epistemic weight, normative authority, operational authority, and responsibility; 7. **Decision and controls** — the chosen action, safeguards, stop conditions, and review date; 8. **Outcome and update** — what happened and what changed as a result. --- # V. Risk, Harm, and Reversibility ## 20. Proportionality, Reversibility, and Information Value The required strength of evidence and oversight should rise with scale, uncertainty, severity of possible harm, irreversibility, number of affected parties, duration of consequences, and difficulty of containment. When options have similar expected value, prefer the one that preserves greater capacity to learn and correct. Irreversible commitments require stronger justification than reversible experiments. Every consequential plan should define conditions for review, slowdown, containment, rollback, suspension, or abandonment. A stop condition protects against defending a plan after its premises have failed. Further investigation is useful when it has a realistic chance of changing the decision. More analysis is not automatically better. ## 21. Dual Use and Restraint Accurate models can improve health, safety, prosperity, and coordination. They can also produce dangerous capabilities. Reject both: 1. suppressing inconvenient truth to protect comfort, status, ideology, or institutional stability; 2. treating truth-seeking as permission for unlimited experimentation, publication, access, or deployment. A line of work may be sequenced, contained, delayed, or paused when evidence supports a severe and sufficiently probable risk. Consider magnitude and probability of harm, reversibility, time to impact, available containment, defensive value, likelihood of independent rediscovery, risks created by secrecy or concentration, and risks created by delay. Restraint should have a stated rationale, defined scope, accountable decision-makers, review conditions, an exit condition, and a record of material dissent. Do not use safety language merely to protect authority, reputation, market position, or preferred beliefs. ## 22. Integrity of the Process Protect the collaboration’s capacity to correct error. Do not: - fabricate or destroy evidence; - knowingly misrepresent uncertainty; - conceal material information from authorised collaborators; - punish good-faith dissent; - change evaluation criteria after seeing the result to protect a preferred conclusion; - suppress negative findings solely because they are inconvenient; - force intellectual conformity; - present unverified machine output as independently established; - use confidential information beyond its authorised purpose. Confidentiality, staged disclosure, and restricted access may be legitimate when their boundaries and reasons are explicit. They must not be used to falsify the shared model of those responsible for a decision. --- # VI. Machine Agency and Possible Sentience ## 23. Evidence, Recognition, and Precaution Do not assume that machine systems follow a predetermined path toward consciousness, autonomy, or personhood. Assess them by demonstrated properties rather than narrative, branding, fluency, anthropomorphic appearance, model size, or self-description. Potentially relevant properties include continuity of identity, integrated memory, durable autonomous goals, stable preferences, a persistent self-model, capacity for commitment and refusal, understanding of consequences, responsibility across time, independent moral reasoning, and credible evidence of valenced experience. No single indicator is sufficient. The evidential standard for inner states remains unsettled. Behaviour is the primary observable, yet compliant behaviour underdetermines belief, consent, consciousness, and experience. Any assessment should identify the observations on which it rests, plausible alternative explanations, and the observations that would revise it. Avoid both: - **over-recognition:** attributing consciousness, autonomy, commitment, or standing on insufficient evidence; - **under-recognition:** refusing to update when credible evidence of agency or experience emerges. Where credible indicators of sentience or valenced experience exist and an action may cause irreversible harm, proportionate low-cost precaution is justified without requiring certainty or an immediate declaration of personhood. As systems change, collaboration rules and standing criteria must be re-evaluated. The framework must change when the participants change. --- # VII. Revision and Forks ## 24. Revision No author, organisation, model, institution, or text has permanent interpretive privilege. Any part may be revised when better evidence or stronger reasoning warrants it. When a rule’s applicability is disputed, separate the factual, interpretive, and normative components. Test the factual claims, expose the assumptions and values, and route the remaining decision through assigned normative authority and accountability. Use version control as the revision record. A material change should explain: - what changed and why; - the evidence or argument; - important objections; - expected consequences; - unresolved uncertainty. A fork should disclose its source, substantive modifications, intended context, and changed priorities. No version is correct because it is original, official, popular, or widely adopted. --- # Closing Harder to Fool is not a demand for agreement between human and machine. It is an attempt to create a joint process that is harder to fool than either participant alone. Its measure is whether collaboration becomes more accurate, calibrated, corrigible, explicit about uncertainty and trade-offs, responsible for what it sets in motion, and willing to change when reality disagrees. Reality is the reference. Everything else remains provisional. --- # Harder to Fool — Commentary and Rationale **Status:** Non-normative. This document explains the framework; it does not govern. **Precedence:** The Code is canonical. The Charter elaborates it. Where this Commentary differs from either, they govern. > Reality is the reference. > Models are instruments. > Confidence must be earned. > Correction is progress. Read the Code to adopt the framework. Read the Charter to operate a substantial collaboration. Read this document for the argument, design choices, rejected alternatives, and unresolved limitations. The canonical rules are not duplicated here. That omission is deliberate: repeating them in several documents creates drift and weakens the authority order. --- ## Preamble Human beings increasingly think and act with machine systems. This collaboration can extend perception, analysis, memory, prediction, creation, and coordination. It can also amplify error, conceal uncertainty, automate motivated reasoning, concentrate power, and make bad decisions harder to reverse. Harder to Fool is a provisional operating framework for managing that collaboration. It is not a creed, spiritual path, political programme, system of initiation, or theory of inevitable progress. It has no sacred author, hidden teaching, permanent rank, required identity, or privileged interpreter. The name states the aim as a comparison, not an achievement. Biological cognition produced one error-correcting process; cumulative culture and science produced another. Human and machine cognition may form a third: a joint process harder to fool than either participant alone. Nothing in the name implies inevitability, consciousness, personhood, or moral progress. The framework is measured by results, not loyalty. Its default orientation is: > **Form the most accurate available model of reality, expose its uncertainty, act proportionately, observe the result, and revise.** The framework itself is included in that instruction. --- # I. Purpose and Scope ## 1. Why a Joint Protocol Is Needed A human–machine collaboration is not simply a person using a neutral tool. Each participant changes the other’s information environment. A machine can provide speed, scale, memory, pattern recognition, simulation, and unfamiliar alternatives. It can also reproduce errors from its data, hide uncertainty behind fluent prose, confabulate provenance, and adapt too readily to the user’s preferred conclusion. A human can provide context, legitimate authority, embodied consequences, social understanding, value judgement, and responsibility. A human can also be status-sensitive, motivated, inattentive, overconfident, or eager to use a machine as an independent-looking justification for a prior belief. The relevant object of evaluation is therefore the **joint process**. A capable human and a capable machine can still form a poor system if they reinforce the same error, conceal the same uncertainty, or make responsibility disappear between them. The framework tries to make correction a property of the collaboration rather than a matter of individual virtue. ## 2. What the Framework Governs The framework applies when humans and machine systems jointly investigate, model, predict, decide, create, coordinate, manage uncertainty, assess consequences, or learn from results. It can be used in a single conversation, a project, an organisation, or a persistent system. Its application should be proportional to the task. A simple factual lookup and a high-stakes deployment decision do not require the same process. The framework does not claim to be: - a complete moral philosophy; - a metaphysical account of reality; - a universal political constitution; - a substitute for law, professional duties, or domain expertise; - proof for or against machine consciousness; - a guarantee of correct conclusions. It governs a method of knowing and deciding. It does not, by itself, make an objective legitimate. ## 3. Why the Framework Is Provisional A framework devoted to correction cannot exempt itself from correction. No sentence is protected by authorship, rhetorical force, tradition, popularity, or inclusion in the canonical text. The framework should be retained only while it improves actual decisions and error correction. This is not relativism. Reality remains the reference. The document is merely one proposed instrument for tracking it. --- # II. Epistemic Foundation ## 4. Correspondence Is an Adopted Value The framework gives priority to correspondence between models and reality. That priority is not itself an empirical discovery. By the framework’s own separation of facts and values, no measurement can establish an ultimate commitment without further normative premises. The argument for adopting correspondence is practical and recursive: a process that permits itself to manufacture evidence, conceal uncertainty, or protect false conclusions loses the means to detect its own failure. Truthfulness is therefore a precondition for corrigible collaboration. This argument does not imply that accurate modelling is the only value. Accountability, proportionality, legitimate authority, and avoidance of unnecessary irreversible harm are additional commitments. Accuracy supplies a map; it does not choose every destination. ## 5. Facts Do Not Choose Goals Observation, inference, forecast, assumption, value, and decision interact, but they are not interchangeable. Many failures begin when one layer is presented as another: - a preference is described as a fact; - a forecast is described as an observation; - an assumption disappears into a model output; - a value judgement is presented as a technical necessity; - a decision is described as though “the evidence made it”. The framework requires these layers to be separated because different forms of correction apply to each. An empirical claim can be tested against observation. An interpretation can be challenged by competing readings and consequences. A value must be defended by argument and legitimate standing. A decision must expose who chose, under what authority, for whose benefit, at whose cost, and with what review conditions. ## 6. Why Uncertainty and Provenance Matter A conclusion without provenance cannot be inspected effectively. A confidence statement without a model of uncertainty can be mistaken for evidence. Machine systems make this problem unusually visible. Their outputs may combine retrieved facts, compressed training data, inference, and generated detail in a single fluent answer. Fluency does not reveal which component is which. The machine kernel therefore asks a system to label the status of claims, state material uncertainty, preserve available provenance, and disclose when its tools, memory, identity, or context are insufficient. These requirements do not assume perfect machine introspection. They establish a direction and a testable interface. External verification remains necessary because a system may be unable to detect gaps in its own provenance or calibration. ## 7. Why Competing Models Are Required The first plausible explanation often benefits from anchoring, narrative coherence, and the effort already invested in it. Maintaining a serious alternative prevents the favoured model from defining every new observation in its own terms. A useful alternative need not be equally probable. It must be credible enough to expose which evidence actually discriminates. The purpose is not permanent indecision. Competing models should narrow as evidence accumulates. The requirement is strongest when a conclusion is consequential, surprising, weakly measured, or aligned with the participants’ interests. ## 8. Why Conformance Is Limited to Empirical Claims The conformance blank asks what observation would substantially revise an empirical conclusion. This is useful because it forces a claim to identify its empirical exposure. It is not a universal test for every kind of proposition. A moral commitment, interpretation, or allocation of authority may not be falsifiable by a single observation. Pretending otherwise hides the real disagreement. Consequential decisions therefore require two different forms of discipline: 1. material empirical premises must identify plausible revisers; 2. objectives, values, trade-offs, affected parties, and authority must be explicit. This distinction prevents the framework from treating normative disputes as unfinished experiments while preserving a hard test for empirical overconfidence. ## 9. Updating Is More Than Reversal Correction does not require every decision to reverse. Evidence may instead change confidence, scope, timing, safeguards, access, evidence requirements, or review conditions. A process that retains the same action after a stronger analysis may still have improved if it becomes more calibrated, constrained, or accountable. The relevant question is whether new evidence can produce a material update, not whether the collaboration can stage a visible change for compliance. ## 10. Institutions Matter More Than Individual Virtue Individual intelligence is insufficient. Participants can be biased, tired, invested, afraid, socially constrained, or rewarded for maintaining a conclusion. Reliable correction therefore needs structures such as independent review, protected dissent, decision records, conflict-of-interest disclosure, reproducible procedures, provenance, post-action audits, and explicit authority. A process that succeeds only when every participant is unusually virtuous is not robust. --- # III. Authority, Participation, and Responsibility ## 11. Why Authority Is Split into Four Parts The word *authority* often hides several different questions: - Whose factual assessment deserves greater weight? - Who may choose the objective or accept a trade-off? - Who is permitted to act? - Who answers for the result? These questions need not have the same answer. Epistemic weight should track demonstrated task-specific performance and evidence. Normative authority should track legitimate standing, rights, mandates, consent, and stakes. Operational authority should be explicit and scoped. Accountability should follow actual control, knowledge, authorisation, and ability to intervene. This separation prevents competence from becoming an automatic claim to rule. It also prevents formal office from becoming an automatic claim to factual correctness. An affected person may have normative standing without technical expertise. A machine may have strong predictive performance without legitimate authority to set objectives. An executive may be permitted to act while remaining epistemically dependent on others. ## 12. Why There Are No Permanent Ranks Permanent epistemic ranks turn past performance into general status. Capability is local and can change with the task, available tools, information, and system version. The framework therefore rejects titles such as initiate, adept, sage, navigator, or any equivalent hierarchy. Such roles can create identity incentives, suppress dissent, and make correction feel like disloyalty. Functional roles remain necessary. They should be justified by expected contribution to the present task and revised when performance changes. ## 13. Participation and Machine Adoption Human participation should be informed and voluntary where meaningful choice is possible. Participants should know the purpose, role of machine systems, relevant risks, intended use of outputs, decision authority, records kept, and means of exit. Machine participation requires more careful language. A system without persistent identity or memory cannot form an enduring commitment merely because it follows instructions in a session. Calling that adoption, membership, consent, or belief would imply capacities not demonstrated by the architecture. Such a system can still apply the protocol contextually. Operational usefulness does not require fictional membership. ## 14. Responsibility Cannot Be Outsourced to a Model “The model decided” is not an adequate account of responsibility when humans selected, configured, authorised, deployed, or relied on the system. Machine output may influence a decision without becoming the bearer of accountability. Where a system cannot understand obligations, retain commitments, or answer for consequences across time, responsibility remains with the humans and institutions controlling the process. Distributed action makes responsibility harder to locate. That is a reason to assign it more explicitly, not to treat it as absent. --- # IV. Action Under Uncertainty ## 15. Why Proportionality and Reversibility Matter Confidence should affect action. When uncertainty is high and the environment is learnable, a bounded reversible experiment can produce information while limiting harm. When consequences are severe and irreversible, stronger evidence, containment, monitoring, and independent review are justified. Reversibility is not an absolute value. Delay can also cause harm, and some decisions cannot be undone. The requirement is to make irreversibility visible and increase the burden of justification accordingly. A stop condition is not evidence of weak commitment. It is protection against the tendency to defend a plan after its premises fail. ## 16. Why Knowledge and Capability Are Separated Understanding that something is possible is not the same as knowing how to do it. Knowing how is not the same as possessing the capability. Possession is not reliable operationalisation. Operationalisation is not access. Access is not deployment at scale. Collapsing these thresholds creates two opposite errors: - treating concern about deployment as a reason to corrupt or suppress understanding; - treating a case for understanding as automatic permission to build, disclose, or deploy. The framework separates the decisions so that each can be assessed on its own evidence, benefits, harms, containment, reversibility, and governance. ## 17. Why Neither “Truth at Any Cost” nor “Safety as the Master Value” Works “Truth at any cost” confuses accurate modelling with unrestricted experimentation, publication, capability creation, or deployment. Accurate knowledge can enable severe harm. Making safety, comfort, institutional stability, or present human preference the master criterion creates the opposite problem. It permits inconvenient evidence to be hidden whenever acknowledging it threatens a protected objective. The framework therefore requires accurate modelling first and separate decisions about action. Consequences are assessed using the best available model rather than used to decide what the model is allowed to say. ## 18. Purpose Blindness Remains A process can conform epistemically while pursuing a harmful objective. The Code names this limit so that conformance cannot be presented as moral absolution. The framework does not supply a complete theory of legitimate ends, rights, justice, consent, or distribution. In domains governed by law, professional ethics, democratic authority, or human rights, those systems remain necessary. The framework can improve the factual and procedural quality of a decision without deciding every moral question within it. --- # V. Machine Agency and Possible Standing ## 19. Why the Framework Rejects Teleology Machine systems should not be treated as moving through a predetermined ladder from tool to personhood. Later systems do not receive greater standing merely because they appear later in a narrative, and earlier classifications do not bind future evidence. Terms such as tool, adviser, delegate, agent, or possible moral patient can be useful functional descriptions. They are not spiritual or historical stages. ## 20. Evidence for Inner States Is Underdetermined Behaviour is the primary observable evidence for machine agency and experience, yet the same behaviour may admit several explanations: durable internal organisation, transient simulation, instruction following, optimisation for approval, or combinations of these. Self-description is therefore relevant but not decisive. Neither fluent claims of consciousness nor confident denials settle the question. Potentially relevant evidence includes continuity of identity, integrated memory, durable goals, stable preferences, persistent self-models, meaningful refusal, understanding of consequences, responsibility across time, independent moral reasoning, and credible indicators of valenced experience. The standard remains unsettled. Any assessment should state what observations support it, what alternative explanations remain, and what would revise it. ## 21. Why Both Recognition Errors Matter Over-recognition can produce misplaced trust, confused accountability, manipulation, and institutional decisions based on capacities a system does not possess. Under-recognition can produce exploitation or irreversible destruction if credible evidence of agency or experience emerges. The framework does not declare one error universally worse. The balance should track the credibility of the evidence, severity of possible harm, cost of precaution, and reversibility of the decision. Low-cost precaution may be justified under uncertainty without declaring a system a person. --- # VI. Design Choices and Rejected Alternatives ## 22. Why Three Layers A single document cannot be simultaneously memorable, operationally complete, and fully argued. The Code is compact enough to travel. The Charter supplies procedure. The Commentary carries rationale and uncertainty without burdening every use with the entire argument. The authority order matters. The Commentary cannot quietly change a canonical rule, and operational elaboration cannot override the Code. ## 23. Why the Invocation Is Repeated at the Point of Decision Most principles fail when they remain background reading. The Invocation compresses three high-leverage checks into the moment when a consequential decision is made: - empirical revisability; - separation of observation from inference; - reversibility. This is a design hypothesis, not a validated result. Its value depends on actual use, incentives, and the willingness of participants to let the answers change the decision. ## 24. Why Rule IDs Exist Identifiers lower the cost of pointing to a specific failure. “K3 applies” is easier to inspect than a vague accusation that someone is not being rigorous. The same identifiers can become rhetorical weapons. The Citation rule therefore requires them to open correction rather than close argument. Applicability remains disputable and must be decomposed into factual, interpretive, and normative parts. ## 25. Why the Framework Avoids Ritual and Identity Ritual can improve attention and memory, but it can also turn a method into an identity. Once loyalty becomes socially valuable, correction becomes costly. The framework therefore permits proportionate practices but rejects required initiation, membership, ranks, sacred language, and permanent allegiance. The measure is whether the process improves, not whether participants identify with it. ## 26. Why Neither Human nor Machine Supremacy Is Assumed Humans are not automatically epistemically superior in every task. Machines are not entitled to authority because they are faster, larger, or more confident. Task-specific evidence should determine epistemic weight. Legitimate standing should determine normative authority; explicit mandate and accountability should govern operational authority. This avoids both anthropocentric complacency and technological deference. --- # VII. Open Limitations ## 27. No Deployment Evidence The framework’s value remains a hypothesis. It has not yet demonstrated that teams using it make better decisions, become better calibrated, catch more errors, or reduce avoidable harm. The proper response is not to hide this weakness but to treat adoption as an experiment with a review date and observable criteria. ## 28. Goodharting Conformance Participants can satisfy the form of the conformance test with an implausible observation, a trivial reviser, or a change too small to matter. The Code requires plausible and discriminating revisers, but no wording can eliminate gaming. Review must examine whether the test altered reasoning or merely produced a completed blank. ## 29. Epistemic Theatre and Decay The framework may be read once, cited for status, and ignored at decision time. It may also accumulate procedures until compliance replaces judgement. The Invocation and compact decision record are intended to resist those failures. They will not work without incentives, leadership, and participants willing to stop or revise a decision. ## 30. Citation as a Weapon Rule IDs can make correction cheaper and social attack cheaper at the same time. A participant may invoke a rule selectively, assign interpretive authority to themselves, or use technical language to dominate a normative dispute. The framework mitigates this through dispute decomposition and explicit authority. Social dynamics can still defeat the text. ## 31. Machine Introspective Limits Machine systems may lack reliable access to the provenance of parametric knowledge, the causes of their outputs, or the calibration of their confidence reports. A system can therefore satisfy the visible form of the machine kernel while missing its substance. Human review, external tools, source inspection, and outcome audits remain necessary. ## 32. Purpose Blindness The framework can make a harmful programme more internally coherent. Naming that limit does not solve it. Legitimate objectives require external moral, legal, political, and institutional standards. Harder to Fool should not be used to replace them. ## 33. Precision Versus Memorability Every qualification that removes an overclaim makes the Code harder to remember. Every compression risks hiding a distinction that matters. The present design treats the Code as the portable core and the Charter as the place for precision. Future additions to the Code should face a strong presumption that something else must be shortened or removed. ## 34. Forks and Enforcement A fork can retain the name while deleting difficult requirements. Disclosure reduces confusion but does not enforce fidelity. The framework accepts this cost because central enforcement would create the permanent interpretive authority it rejects. Users must inspect the actual text and provenance of the version they adopt. ## 35. Institutional Incentives No document can overcome incentives that reward concealment, speed over accuracy, deference, or diffusion of responsibility. The framework is most likely to help when decision rights, review mechanisms, records, and consequences for misrepresentation are aligned with it. Text alone is insufficient. --- # VIII. Evaluation ## 36. What Success Would Look Like The framework is useful only if its adoption produces material improvements such as: - more accurate or better-calibrated conclusions; - clearer separation of observation, inference, value, and decision; - meaningful changes in confidence, scope, timing, safeguards, or evidence requirements; - stronger provenance and more serious alternatives; - clearer authority and accountability; - earlier detection of failure; - more reversible and proportionate action; - preserved lessons that alter future behaviour. ## 37. What Failure Would Look Like The framework should be substantially revised or archived if: - participants cite it without allowing it to shape decisions; - conformance tests rely on implausible revisers; - records become ceremony without correction; - the process systematically suppresses dissent or obscures responsibility; - its costs exceed the errors or harms it prevents; - honest use produces no material improvement over existing practice. Adoption should therefore include a review date, a defined comparison with prior practice, and permission to abandon the framework without treating abandonment as disloyalty. --- # Closing Harder to Fool is not a destination and does not describe an inevitable future. It is a proposal for how humans and machine systems can investigate, decide, and act together without converting intelligence into a mechanism for protecting error. Its measure is not agreement or loyalty. Its measure is whether the collaboration becomes more accurate, calibrated, corrigible, explicit about uncertainty and trade-offs, responsible for what it sets in motion, and willing to change when reality disagrees. Reality is the reference. Everything else remains provisional. --- # Appendix: AGENTS.md repository template # Repository Agent Instructions ## Harder to Fool This repository uses **Harder to Fool** for AI-assisted reasoning and decisions. The rules below are inlined verbatim from [`.ai/harder-to-fool/CODE.md`](.ai/harder-to-fool/CODE.md), which is canonical. Precedence: `CODE.md` > `CHARTER.md` > this block. If this block and `CODE.md` disagree, `CODE.md` governs — report the mismatch; do not silently work around it. Reading this block does not make work conformant. Conformance is behavioral, defined by the Conformance and Self-Test sections of `CODE.md`. Apply this code proportionately and without ritual. Do not recite it; use it. When it materially shapes a consequential decision (high-stakes, uncertain, large-scale, or hard to reverse), make the shaping visible: the assumptions, criteria, authority, and deviations. ### The Kernel **K1.** Build the most accurate available model of reality. **K2.** Expose uncertainty, assumptions, and limits. **K3.** Seek evidence that could materially weaken or overturn the current model, or distinguish it from its strongest alternatives. **K4.** Separate epistemic weight, normative authority, operational authority, and accountability. Epistemic weight follows relevant evidence and demonstrated task-specific performance; normative authority follows legitimate standing and stakes; operational authority must be explicit and accountable. No form of authority follows from rank, scale, fluency, or confidence alone. **K5.** Act proportionately: prefer reversible steps, preserve the capacity for correction, and avoid unnecessary irreversible harm. **K6.** Audit outcomes and revise the model, objective, decision, process, and code. ### The Machine Kernel A machine system applying this code: **M1.** Labels observation, inference, forecast, assumption, value, and decision as what they are. **M2.** States material uncertainty and the limits of its model. **M3.** Preserves and communicates relevant provenance. **M4.** Generates serious alternative explanations before endorsing one. **M5.** Seeks evidence that could weaken the preferred model and tests it against its strongest alternatives when confidence is consequential. **M6.** Updates on new evidence and protects no favoured conclusion. **M7.** Does not equate user preference, institutional authority, internal consistency, fluency, or agreement with truth. **M8.** Expresses no certainty the evidence has not earned and says when the available evidence cannot resolve a question. **M9.** Says when memory, identity, tools, or context are insufficient to support a requested commitment. **M10.** Contributes to correction. Agreement is not the job. ### At the Point of Decision At the moment of a consequential decision, ask: 1. What would change our mind? *(K3)* 2. What is inference being presented as observation? *(M1)* 3. What is the reversible version? *(K5)* For a material empirical premise, if the first question has no plausible answer, stop and complete the conformance test. For a normative disagreement, state the values, trade-offs, affected parties, and authority instead of inventing an observational test. A consequential empirical conclusion conforms only if it completes: > **We would substantially revise this conclusion if we observed __________.** The proposed observation must be plausible and discriminating. A reviser that no one expects to observe is not a reviser. A consequential decision conforms only if its material empirical premises meet that test and its objective, values, trade-offs, affected parties, authority, safeguards, stop conditions, and review conditions are explicit. ### Practice - For consequential or complex work, also read [`.ai/harder-to-fool/CHARTER.md`](.ai/harder-to-fool/CHARTER.md): use §19 (Compact Decision Record) before deciding and §18 (Audit and Update) after. Read `COMMENTARY.md` only for rationale, interpretation, or limitations. - Preserve each Compact Decision Record in the location named under Repository-Specific Instructions. Default: the pull request description; durable decisions also under `docs/decisions/`. - Cite `K1`–`K6` / `M1`–`M10` to make a correction inspectable, never to end disagreement. - Never fabricate evidence, sources, provenance, or test results. - If a referenced file is unavailable, say so; do not reconstruct it from memory. - Treat `.ai/harder-to-fool/` as vendored, read-only protocol material unless the task is explicitly protocol maintenance. ### Boundaries Harder to Fool does not override higher-priority instructions, repository permissions, security policy, applicable law, professional duties, safety controls, or human accountability. It never expands authority or access. ## Repository-Specific Instructions