Simple Memo - for Obsidian ObsidianMedia kitMarkdown tool

Field report · a small iPhone app

What happened when AI ran the daily operations of an app?

We published the runs that shipped, the runs that failed, and the times a human stepped in. Across 23 days, 19 of 28 attempted runs shipped work. A successful workflow status did not always mean useful work was produced.

Published 5 September 2026 · Observation window: 11 August–2 September 2026, Japan time · By the Simple Memo development team

67.9%Shipment rate: 19/28 attempted runs
25.0%Human intervention: 7/28 attempted runs
41Ledger records, including skips and no-run records

This is a historical case study of Simple Memo - for Obsidian, a working iPhone app. The operating loop measures results, chooses a task, makes a change, checks it, ships eligible changes, and records the outcome. It is a study of app operations, not a claim that an AI independently owns or runs a company.

Download all 41 rows (CSV)Snapshot and source (JSON)Media kit and illustrations

How we counted

We selected public ledger records dated 2026-08-11 through 2026-09-02, inclusive. A row is one recorded run. The 13 rows with attempted=false include 11 gate skips and 2 no-run records. They remain in the download but are excluded from the shipment and intervention denominators.

A shipment means outcome=shipped in the ledger. It can be a website article, an update, or an operations change; it is not an App Store release. A human intervention means at least one recorded intervention on an attempted run. It includes infrastructure repair, initial setup, substitution, and an owner request. It does not measure minutes worked.

The snapshot was extracted from a specific revision of the public ledger. Later records do not change this observation window. The live ledger and full Japanese operations page continue to change. The page and all three charts are generated from the same downloadable snapshot.

Results with the denominators intact

Attempted-run outcomeRunsShare of attempted runs
Shipped1967.9%
Failed621.4%
No artifact27.1%
Cancelled13.6%
Of 28 attempted runs: 19 shipped, 6 failed, 2 produced no artifact, and 1 was cancelled.
Shipment rate is an outcome measure. It is not the percentage of the business that is automated.

There were no recorded human edits to the shipped artifacts. This does not mean there was no human involvement: infrastructure repairs occurred in four attempted runs, and initial setup, substitution, and an owner request each appeared in one run.

Human intervention appeared in 7 of 28 attempted runs, infrastructure repair in 4, and artifact editing in 0.
Intervention categories describe work types. Zero artifact edits does not imply zero human work.
17 of 23 calendar days had at least one shipment; 6 had none.
17/23 calendar days had a shipment. Starting a workflow and shipping useful work are separate events.

Three failures that changed what we check

Green status, no artifact

On 22 August, the workflow reported success while the ledger recorded 14 permission denials and no artifact. Run: ap-20260822-actions. A success status alone was insufficient evidence of delivery.

A stale claim blocked the fallback

On 29 August, a fallback claimed the day’s branch but produced neither an article nor a PR. The primary route treated the claim as work in progress and skipped. Run: ap-20260829-ccr0920.

Two routes, one shared limit

On 30–31 August, the primary route hit an account usage limit. The 31 August fallback shared the account and also failed. Different schedules did not provide independence from the same limit.

These are observations from the source records, including later diagnostic annotations. They are not a controlled comparison of AI vendors, and the original cause was not always known at the time of failure.

What humans still decide

The published operating policy separates reversible website work from pricing, contracts, spending, and App Store release decisions. Humans remain responsible for those decisions and for device verification. The complete, changing permission table is on the Japanese operations page.

Task-inventory automation rates, shipment rates, and human working-time savings have different denominators. This report does not infer one from another. It also does not demonstrate automatic production rollback, end-to-end incident recovery, or an increase in revenue.

Limits and how to cite this report

This is one team, one app, a short observation window, and an operator-maintained ledger. Some entries were reconstructed from workflow history. Missing interventions would undercount human involvement. There is no control group and no measured counterfactual human workload.

For coverage, retain the observation dates and the denominator. A suitable description is: “In Simple Memo’s 11 August–2 September 2026 operations ledger, 19 of 28 attempted runs shipped work; 7 of those 28 recorded human intervention.” Link to this report or its snapshot so readers can check the definitions.

To inspect the product behind these records, see the Obsidian workflow, Siri setup, or the product facts and media kit. For corrections or reproducibility questions, contact the developer.