We tell others to break hypotheses into layers, yet our own ledger had not a single success criterion
To see whether a rule was being followed, we counted a word in our documents and got 3. When we opened them, all three only explained the concept; none actually applied it. Here is one number without that string trap instead.
When we start a project, we write that hypotheses should be split into three layers.
Major hypothesis 1 the root cause of the problem
└ Mid hypotheses 2–4 what must hold if the major one is true
└ Minor hypothesis 1 each what you actually judge within 2 weeks
We decided to check whether this was actually applied in our own documents. Two things came out of it: one was a flaw in our work, and the other was a flaw in how we measured it. The second one mattered more.
First, the measurement was wrong
We went through all of our working documents and counted the ones containing "major hypothesis." 3 came up. It was a low number, and we were about to use it as is.
Then we opened those 3 documents.
Every one of them was a document explaining the concept of a major hypothesis. One defined the rule, one audited that rule, and one was a list of content ideas. Not a single project document had actually set a major hypothesis.
"3" meant that three files contained the term, not that three projects had built that layer. We were counting strings.
Checking once more made it even clearer. While preparing this article we wrote a few related documents, and the same number rose to 6. Not because anyone set a major hypothesis, but because more documents used the term "major hypothesis." Once published, the article you are reading will push that number up by one more.
Counting a word is not the same as seeing whether something is actually being done. So we did not use this number as evidence in this article.
One number without the string trap
Instead, we looked at something else. We record every decision in a decision ledger. When we started writing this article, it held 985 entries.
These are fields in a record, not document text, so what we check is whether a field is filled in, not whether a word appears.
- The criteria for calling this decision a success — 0
- What would make us drop it — 0
- When to look at it again — 0
0 out of 985. There is no room for interpretation here.
Put side by side, it looks like this.
Decisions 985
Executions 2,825
Outcomes 66
Decisions and executions keep piling up, but the outcomes are not there. Because we never wrote down the criteria for judging, we had no way to judge later.
And then, today
While counting these numbers, we added three of that day's decisions to the ledger.
We did not write success criteria for those three either.
While counting the flaw, we created it three more times. Not because we didn't know the rule. We are the ones who wrote that rule.
Why this happens
The reason turned out to be simple: the moment you decide and the moment you judge are different.
When you decide, what to do is clear, so that is all you write down. "What would make me admit this was wrong" can go unwritten at that moment, and the work still moves forward. Nothing goes wrong right away.
The problem comes months later. When you reopen that decision, nobody knows what to use to say whether it worked or not. It usually ends with "it was the right call at the time." It becomes a look back, not a judgment.
Why the layers are needed
- The major hypothesis is the premise the project stands on. If it is wrong, everything has to be reopened, so you don't touch it often.
- Mid hypotheses are what must follow if the major hypothesis is true. Even if one is rejected, the project survives.
- Minor hypotheses are cut down to a size that can be judged within a short cycle. We set that cycle at 2 weeks.
The key point is that rejection becomes a normal state. Because one mid hypothesis being wrong does not bring the project down, people can actually say it was wrong. With only one layer, rejecting a hypothesis means rejecting the project, so nobody says anything was wrong.
For reference, even in well-designed experiments the target metric improves in only 1/3 of cases. Another 1/3 show no change, and the remaining 1/3 actually get worse. Rejection is not the exception; it is the default.
Take Action
Free Meeting / Consultation
Clarify your next action and scope based on the context you just read.
If you do just one thing today
Pick just one piece of work currently in progress and write a single line.
What would make me admit this is wrong?
If you can't write that line, the work is still a plan, not a hypothesis. A plan can be finished, but it cannot be judged.
And one more thing from our own experience — when you check whether that line was written, don't count by whether the words appear in a document. That is what we did: we got the number 3, and when we opened the documents, it was 0.
What changed while writing this
We cannot retroactively fill in the past 985 entries. Instead, while writing this article, we filled in three fields for that day's three decisions — success criteria, drop criteria, and a review date.
3 out of 989. It is not a number to boast about. But it is the first number that isn't 0, and it marks where the decisions we record from now on start.
If you're not sure whether your work in progress is in a state that can be judged
A 2-minute diagnostic lays out one task to tackle first, a starting plan for a 2-week pilot, and the risks a person should check. Please feel free to answer only as much as you're comfortable with.
STAR-T AI Business Operations Diagnostic →
Notes
① Sources and check dates (measured 2026-08-12 14:04 · re-measured 15:37)
- Decision ledger 985 → 989 (increased during measurement) · queried the ledger directly
- Success criteria / drop criteria / review date: 0 → 3 each — queried whether the field exists. The 3 were filled in while writing this article
- Executions / outcomes 2,825 → 2,841 / 66 (the block in the body keeps 2,825 because it comes from the same snapshot as the 985)
- String search for "major hypothesis": 3 (8/10) → 6 (8/12) — ⚠️ Not evidence. Used in the body only as a counterexample
⚠️ The decision ledger is a record that grows every day. The totals above are values as of the measurement time, and they will be higher by the time you read this. They kept changing even during review (983 → 984 → 985 → 989). What did not change is the "0" — the claim of this article rests on that 0, not on the totals.
The finding that experiment results split into 1/3 improved, 1/3 no effect, and 1/3 worse comes from material Microsoft presented at a conference in 2015, summarizing 12 years of online controlled experiments.
② Figures not used in this article
- We did not use "3 documents containing a major hypothesis" as evidence of the flaw. The draft had it in the title, but in review it was pointed out that a string count gets contaminated self-referentially, and when we checked, that was true (it rose from 3 to 6). Rather than swap in a different number, we kept that failure itself in the body.
- Percentages (%) — most documents in the denominator, such as logs and meeting notes, don't need hypotheses at all, so a ratio would read worse than reality.
- "A hit rate of 1/3" — our internal documents abbreviate it this way, but the original source says 1/3 each improved / no effect / worse. We followed the original.
- External support for the 2-week standard — there are external reports recommending short cycles, but their standard is 1 week. Our 2 weeks is looser than that, so we did not borrow someone else's number as support.
- Per-track figures · satisfaction and performance figures for our lectures and consulting — left out because they are unrelated to the claim.
③ How this was made
The draft was produced with AI, then verified and edited by a person. The draft was rejected once during fact-checking — the reviewer pointed out that the key figure did not hold up given how it was measured; we checked, found that was right, and rewrote the piece with a new title and argument. No generative images or audio were used.
Engagement
Views and reactions are saved as internal content signals.
Key points
- •If you check whether a rule is followed by counting whether a word appears in documents, documents that only explain the concept get counted too, and the result looks better than reality.
- •Of 985 decision ledger entries, 0 had success criteria, drop criteria, or a review date (3 were filled in later); since this checks whether a field exists, there is no room for interpretation.
- •Decisions and executions pile up without outcomes because the criteria for judging were not written at the time of deciding, and leaving them out causes no problem in the moment.
- •Splitting hypotheses into three layers lets the project survive a rejected mid hypothesis, so rejection becomes a normal state; with only one layer, nobody says anything was wrong.
- •Even in well-designed experiments, only 1/3 improve and the rest show no effect or get worse, so rejection is not the exception but the default.
Frequently asked questions
How do we check whether our team is actually following a rule?
If you count by whether a word appears in documents, documents that merely explain the concept get counted too, and the result looks better than reality. It is more accurate to check whether the fields in the output are filled in, that is, whether the criteria were actually written down.
Why do success criteria so often get left out of decisions?
Because work moves forward even without them. At the moment of deciding, what to do is clear, so that is all that gets written; what would make you admit you were wrong only becomes necessary months later. By then, it can no longer be written.
What changes when hypotheses are split into three layers?
Rejection becomes a normal state. With a single layer, saying a hypothesis is wrong amounts to saying the project is wrong, so nobody raises it. Validation actually happens only when the structure lets the project survive a rejected mid hypothesis.
Don't just read — connect to the right service or consultation and take action now.
Once you understand the problem through insights, the next step is deciding on the execution structure. Jump straight to related services or a free consultation.
Free Meeting / ConsultationSTAR-T
STAR-T Chief Consultant
As an IT service planning and design expert, I research and share success stories from various startups and companies.
Take Action
Don't just read — connect to the right service or consultation and take action now.
Once you understand the problem through insights, the next step is deciding on the execution structure. Jump straight to related services or a free consultation.