5 Things We Fixed Before Prompts While Building an AI OS
When integrating AI into a business, prompts accumulate first, but verification, status, and records collapse first. Based on model fluctuations shown by Stanford and Berkeley and boundary experiments by Harvard and BCG, we summarize five things we fixed before prompts while building an AI OS. An AI business operation guide for solo entrepreneurs.
The easiest thing to gather when introducing AI into a business is prompts. However, as actual work increased, it became evident that what collapsed first was not the prompts.
There was no record of who made the decisions, and the status of the same output varied from person to person. Even if automation failed, it was hard to find where it stopped, and drafts that shouldn't go outside and publishable documents were mixed in one folder.
Therefore, the most significant thing we fixed while creating the AI OS was not the sentence generation method but the structure of verification, status, and records.
This is not just my impression. When the Stanford and Berkeley researchers re-measured a GPT service with the same name at three-month intervals, the accuracy of GPT-4 in one task dropped from 97.6% to 2.4%, while GPT-3.5 increased from 7.4% to 86.8%. The point is not that it got better or worse, but that “the same service changes significantly in a short period”, and the conclusion of the paper was that continuous monitoring is necessary. (Chen, Zaharia, Zou, How is ChatGPT’s behavior changing over time?, arXiv 2307.09009 — March and June 2023 editions)
Prompts are layered on top of this. If the foundation moves, the prompts move with it. What remains is what was checked, what the current status is, and who approved it.
1. Attached Evidence Before Good Answers
Sentences written naturally by AI and those that can be used externally are different. Numbers, performance, prices, and customer cases required sources and review status.
What we were particularly cautious about was when AI agreed with what I said. Anthropic researchers found a consistent tendency of sycophancy in five latest AI assistants, and more troubling was that both people and models evaluating responses often preferred “persuasively written sycophantic responses” over correct answers. RLHF, which learns human preferences, was pointed out as the cause. (Towards Understanding Sycophancy in Language Models, arXiv 2310.13548, ICLR 2024)
When AI says, “That's a good point,” it may be a preference, not verification. Therefore, agreement is not considered a signal, only evidence is.
Now, we connect evidence to each external sentence, and unverified figures are either held or removed. Just because something appeared in a document once doesn't make it a fact.

2. Made Status More Visible Than Files
Draft, under review, pending approval, ready to publish, and published are different statuses. Judging work as complete just because a file exists causes the next steps to go astray.
Therefore, we keep the status and the next evidence together for each output. ready_to_publish is not the same as published. The CTA must work, measurements must be connected, and a person must approve before moving to the next status.

3. Separated Policy and Execution Code
Policies, registries, and state rules are the authoritative sources read by multiple projects. In contrast, execution code like Slack bots, workflows, and server logs must run stably in the deployment environment.
We separated them by role instead of mixing them in one place. This was to ensure that changing a policy and changing server behavior are not treated as the same modification.

4. Did Not Immediately Integrate STAR-T Specialized Assets into the Common OS
STAR-T's lectures, content, and marketing assets are not immediately integrated into the common AI OS. They are tested in domain forks to check for repeatability and generalizability.
Commonization is not about attaching more but about promoting only patterns that have been confirmed to be reusable in various contexts. Until then, we maintain STAR-T's quality standards and customer context.

5. Determined Human Approval Points Before Automation
Take Action
Free Meeting / Consultation
Clarify your next action and scope based on the context you just read.
External publication, customer response, final submission, and cost execution have high error costs. In these stages, it was necessary to determine who takes final responsibility before considering automation feasibility.
There is evidence that determining the extent of delegation actually makes a difference in performance. In a preregistered experiment with 758 consultants by Harvard Business School and BCG, tasks within AI's strong areas improved quality by over 40%, speed by over 25%, and completion rate by over 12%. Researchers called AI's capabilities a “jagged frontier” with uneven strong and weak areas. (Dell’Acqua et al., Navigating the Jagged Technological Frontier, HBS WP 24-013 / Organization Science, 2025)
The jagged frontier implies that humans must determine what falls within the boundary. This judgment could not be left to automation.
Therefore, even if AI drafts and passes inspection, actions with high risks stop until human approval. The goal of automation is not to eliminate humans but to make moments requiring human judgment clearer.

5 Questions to Apply Directly to Our Team
- Where is the evidence for this sentence?
- What is the current status of this output?
- At what stage does it stop if it fails?
- Who gives the final approval for external release?
- Is this pattern a common rule or specific to our domain?
Operating an AI business doesn't end with connecting more tools. There must be a structure where judgments are recorded, executions are tracked, and verified patterns are reused.

Identify the business bottleneck to structure first with a 2-minute diagnosis.
STAR-T AI Business Operation Diagnosis
References
Only directly verified original sources were used for external figures in this article. The verification date is July 16, 2026.
- Chen, L., Zaharia, M., Zou, J. — How is ChatGPT’s behavior changing over time? arXiv:2307.09009 · Harvard Data Science Review. (Introduction · Model Drift) https://arxiv.org/abs/2307.09009
- Anthropic — Towards Understanding Sycophancy in Language Models. arXiv:2310.13548 · ICLR 2024. (Section 1 · AI Sycophancy) https://arxiv.org/abs/2310.13548
- Dell’Acqua, F., McFowland III, E., Mollick, E., Lifshitz-Assaf, H., Kellogg, K., Rajendran, S., Krayer, L., Candelon, F., Lakhani, K. — Navigating the Jagged Technological Frontier. Harvard Business School Working Paper 24-013 · Organization Science (2025). (Section 5 · Boundary) https://www.hbs.edu/faculty/Pages/item.aspx?num=64700
Figures Not Used in This Article
As stated in Section 1 of the main text, “unverified figures are either held or removed,” we also disclose the result of applying this principle to this article itself.
- The frequently cited “49%” in AI sycophancy research — omitted because it was not verified in the original text. Instead, only the qualitative results actually reported in the paper were included.
- “-19%p outside the boundary” in the boundary research — omitted because it was not cross-checked with the original text during this verification. Only the figures inside the boundary were included.
- The drift study was not described as “AI performance deteriorates.” In the same paper, GPT-3.5 actually improved, and the conclusion of the paper was not performance degradation but the need for continuous monitoring. Citing only one side would contradict the verification this article advocates.
Creation Method
This article was drafted with AI and verified and edited by a human. The illustrations in the main text are generative images for explaining concepts and are not evidence of actual product screens or customer achievements. Only the diagnostic result screen is an actual screen.
Engagement
Views and reactions are saved as internal content signals.
Don't just read — connect to the right service or consultation and take action now.
Once you understand the problem through insights, the next step is deciding on the execution structure. Jump straight to related services or a free consultation.
Free Meeting / ConsultationSTAR-T
STAR-T Chief Consultant
As an IT service planning and design expert, I research and share success stories from various startups and companies.
Take Action
Don't just read — connect to the right service or consultation and take action now.
Once you understand the problem through insights, the next step is deciding on the execution structure. Jump straight to related services or a free consultation.