Technical Research · September 11, 2026

Less Fiddling.
Ten Small Fixes.

Prioritized changes to reduce recurring setup, confusing status, and maintenance noise using the existing tools.

Each item is scoped for one agent and an estimated 10–19 minutes, including focused verification. Estimates are judgments, not measured guarantees. Code changes stop at verified local diffs; production releases are separate. This report implements none of the recommendations.

Ranked Recommendations

01

Stop routine telemetry from dirtying Git

10–15 minutes

Two frequently changing skill telemetry files are already in the ignore rules, but Git still tracks them. That means ordinary bot activity can keep creating changes you have to inspect or land. I would preserve both files exactly, remove only their Git tracking, and verify that subsequent writes remain ignored. This is a targeted correction to an existing policy, not a repository cleanup campaign. Completion means both files still exist with matching contents, Git no longer tracks them, and the diff contains only the intended index changes. It should remove a recurring source of maintenance noise immediately.

02

Stop cleanup from deleting installed automation browsers

10–15 minutes

The disk-cleanup script currently deletes everything in Playwright’s browser installation cache. Those files are executable browser dependencies, and Playwright requires compatible installed versions. Removing them can make the next browser task require another download or repair. I would remove this automatic deletion step and its misleading cleanup description, leaving browser removal as a deliberate maintenance operation. Verification would check that the cleanup path cannot target this directory and that the installed browser executable still exists; no destructive cleanup run is necessary. The tradeoff is retaining that disk space, in exchange for fewer broken previews and surprise installation chores.

03

Separate machine health from AI account allowance

15–19 minutes

Your dashboard already keeps account quota separate from host health, but its maintenance evaluator still includes quota in the same success count and attention list. That disagreement can send you investigating a healthy Mac simply because account allowance is low or unavailable. I would introduce explicit host-health and budget results, retaining compatibility fields where an existing caller needs them. The bounded task covers the evaluator, its direct consumer wording, and focused tests. Success means a healthy host stays healthy with missing or exhausted quota, while a separate budget notice remains visible. No quota reset, purchase, or threshold relaxation is involved.

04

Give maintenance one consistent memory-health rule

10–15 minutes

The routine correctly says macOS memory pressure supersedes unused RAM, then still lists raw free RAM as an all-green objective elsewhere. An agent reading both instructions has to reconcile them every run, and you inherit the confusion. I would rewrite the active goal and intervention sections so pressure drives memory decisions, free RAM remains context, and account budget has a separate objective. Historical logs would remain evidence rather than operational instructions. Verification is a targeted consistency review against the current evaluator and dashboard, with stale directives removed. This is a documentation correction, not a promise to consolidate every older collector.

05

Stop expired readings from looking healthy

15–19 minutes

The dashboard already marks old telemetry as stale, but its main dial and summary can retain the previous ALL HEALTHY result. You then have to notice the smaller timestamp to know whether the reassuring status is current. I would make the overall result visibly unknown after the existing freshness window expires, while preserving the last measurements as historical values. Both dashboard copies would receive the same change. A controlled-clock test would cover fresh data, time passing without another fetch, and recovery when new data arrives. This improves trust without changing health thresholds, sampling frequency, or the five-gauge layout.

06

Make sampling-interval changes recover from failure

15–19 minutes

The interval-setting helper validates the old configuration, unloads the sampler, changes the file, and reloads it. If a later step fails, there is no explicit rollback path. A harmless preference change could therefore become a missing-telemetry investigation. I would prepare and validate the replacement before unloading, retain the previous file and load state, and attempt restoration on failure with an explicit recovery error if needed. Success and failure paths would be tested with mocked commands, including a failed reload. This task hardens the helper; it does not require changing your current interval or interrupting the running sampler.

07

Make an empty skill audit fail honestly

10–15 minutes

I ran the existing skill-collision checker with no paths: it successfully returned zero entries and zero collisions. That looks clean even though it checked nothing. Given your earlier command-routing confusion, this is exactly the kind of reassuring output that creates another troubleshooting loop. I would require explicit valid roots, reject empty scans, and report how many files and real locations were inspected. Focused tests would cover no arguments, an empty directory, duplicate names, and multiple symlinks to the same file. This strengthens the existing checker rather than creating another one; it does not claim that filesystem names prove runtime dispatch.

08

Stop stale disk samples from generating cleanup advice

10–15 minutes

The evaluator rejects stale measurements for its green checks, but still calculates exact disk-space shortfalls from those same old numbers. The operating notes say those gaps should be unavailable when evidence is stale. I would gate the actionable gap fields on fresh, valid disk evidence and return an explicit unavailable result when it cannot be trusted. Tests would cover expired readings, missing timestamps, invalid totals, and future timestamps, alongside a valid sample with a known shortfall. This prevents a maintenance agent from treating yesterday’s disk deficit as today’s cleanup target. Valid historical measurements can remain available for reference.

09

Keep the two dashboard copies from drifting

10–15 minutes

The two dashboard files currently match, but the existing tests check selected behaviors separately rather than enforcing that match. Future edits can quietly fix one surface while leaving the other behind, making you wonder why two views disagree. I would add an equality check and a small explicit synchronization command from the documented canonical dashboard. The command would detect unexpected destination edits before overwriting them. Verification would introduce a difference in temporary fixtures, prove the check fails, and prove synchronization restores equality. This gives the current arrangement a simple guard without turning a short task into a shared-component architecture project.

10

Remove old measurements from live dashboard help

10–15 minutes

Some hovercards contain fixed sentences saying the latest check showed particular swap and compression percentages. The main gauges update live, so those sentences can disagree with what you are looking at and invite another investigation. I would replace the baked-in readings with stable explanations, or derive them from the same current sample when the reading is useful. Both copies would stay synchronized. Verification would load two different samples and check that no old percentage is still described as latest. You keep the detailed hovercards and the existing visual design, but the help stops making historical information sound current.

Internal

Current source inspection supplied the strongest evidence. Findings include tracked telemetry covered by ignore rules, browser-binary deletion, quota mixed with host evaluation, stale status handling, and a collision checker that reports success after scanning zero files. Dashboard copies currently match; the proposed parity check is preventive.

Local source anchors, reviewed September 11, 2026:

  1. .gitignore:89
  2. crons/routines/forgebotmini-maintenance/scripts/free-disk.sh:105
  3. apps/forgebot/scripts/green-health.ts:15
  4. crons/routines/forgebotmini-maintenance/ROUTINE.md:21
  5. apps/forgebot/src/pages/mini-dashboard.html:115
  6. apps/forgebot/scripts/configure-mini-health.ts:12
  7. .forgebot/skills/skill-librarian/scripts/hermes-collisions.ts:49
  8. apps/forgebot/scripts/green-health.ts:13
  9. apps/forgebot/scripts/mini-dashboard-contract.test.ts:6
  10. apps/forgebot/src/pages/mini-dashboard.html:178

Workplace coverage: three public Slack searches found direct historical evidence of unattended-operation and Git-maintenance concerns. Fireflies returned one relevant weekly-meeting summary, but it did not establish a technical root cause. No meeting content, staff details, or private message quotations are reproduced here. Historical Computer History context informed scope; current service health was not inferred from it.

External

Primary documentation supports the mechanisms; the ordering is an engineering judgment based on local evidence and recurring interruption risk.

Method And Limits

Research questions: Which interruptions recur? Which code currently causes ambiguity? Which protections already exist? Which fixes fit a short solo task? How can each be verified without disrupting active work?

Ranking favors recurring effort avoided, confirmed gaps, small scope, and low disruption. No invented time-savings percentages were used. Already implemented: single-run protection, bounded history, restart cooldown, subprocess timeouts, and HTTP timeouts. Full background-service migration, fleet-wide threshold consolidation, and unattended reboot recovery were excluded from this timebox.

Coverage includes local code, historical context, public Slack, Fireflies search, official web documentation, and upstream GitHub sources. Video and social commentary were not useful for these code-specific findings. No independent external-model research pass was used; a separate code-review agent checked local findings. Live scheduling and production deployment were not changed.