LOCAL ARTIFACT · /Users/forgebot/forgeapps/artifacts/research/dave-codex/dave-codex-working-model.html
Summary
Recommendations
Keep Dave’s primary workstation for now. Diagnose the reported failures before changing hardware. Keep his morning review short, and start with one useful financial task that leaves the existing records unchanged. Allow more unattended work only when the agent can reliably complete comparable tasks and provide evidence that they work.
Responsibilities
Adam can provide shared infrastructure, technical supervision, and rules for releasing changes. Dave can supply business examples and judge whether the result saves time without disrupting billable work. The agent should implement the agreed change and provide test results.
When Unattended Work Is Appropriate
Not every task needs supervised practice first. It needs clear inputs, checks that can detect mistakes, and limited consequences if it fails. A new agent can prepare read-only research overnight. Changes to financial records still need specific controls, even with an experienced agent.
Limit how much work is waiting for Dave’s review. Producing more changes overnight can slow the project if he cannot review them. Measure useful, accepted improvements against the human time they require.
Source Discussion
Dave’s And Adam’s Priorities
Dave wants reliable tools, less repetitive accounting work, and progress that does not take time from John Deere RWC work. Morning feedback lets him review changes around those commitments. It is not evidence that he is disengaged. [I1]
Adam recommends avoiding extra infrastructure too early, reusing ForgeBooks, keeping Dave involved in feature decisions, working interactively to correct the agent, and eventually delegating longer jobs. His reports of never having a computer crash, spending roughly 80% of AI time on credentials, and running 8–12-hour jobs describe his experience. They are not comparative benchmarks. [I1]
Dave says the app “hangs up and crashes sometimes.” He does not say Windows crashes. Adam’s statement, “I’ve never had Codex crash my computer,” addresses a different failure. An app crash, stalled command, lack of memory, and machine reboot need different investigations.
ForgeBot’s initial answer recommended the backup PC before diagnosing the failure. It supported morning batches without limiting the work Dave would need to review. Test data was a sensible recommendation, but it did not resolve whether Dave needed a second financial application.
Dave’s quoted statement was verified in Adam’s source message. Searches within ForgeBot’s permitted channels did not find a separate original. No inaccessible DM or private source was read.
Public Research
Findings And Limits
1. Where The App, Tools, And Model Run
The computer runs an interface. Tools read files, build and test code, render images, or control applications. The model service generates answers and chooses actions. Using a desktop interface does not mean the model runs on Dave’s GPU. Local tools can still use substantial CPU, RAM, disk, and GPU resources. Cloud execution moves the tools to a provider-hosted environment. [E1] [E2]
OpenAI documents native Windows and WSL environments with different default locations for configuration and authentication. One computer can therefore require more than one environment to be maintained. Keeping one machine reduces some setup work, but does not remove all of it. [E1]
Do not buy a Mac to speed up cloud model responses. Consider a separate computer if a measured local resource problem warrants it, if isolation is useful, or if jobs need to run independently of Dave’s workday. Check platform-specific dependencies and licensing before choosing macOS.
2. Which Tasks Can Run Without Supervision
Anthropic’s work on long-running agents found repeated failures: attempting too much, leaving work unfinished, and claiming completion without end-to-end checks. Its approach uses saved progress records, a feature list, small implementation steps, and testing. Later work separates the agent producing the work from the agent evaluating it because agents tend to praise their own output. These are provider engineering experiments. They do not establish a best workflow for every team. [E3] [E4]
A new agent can inventory screens, reproduce a bug in a disposable workspace, or propose tests while Dave is offline. Supervised practice is not a prerequisite for all such tasks. Changes to rounding rules or payment behavior still require expert review, regardless of how familiar the agent is.
Use a short joint session to resolve unclear business requirements and approve representative examples. The agent can then implement the agreed work without Dave watching each step.
3. Test Results And Business Correctness
METR’s March 2026 study found that roughly half of test-passing SWE-bench patches from the studied mid-2024 through mid/late-2025 agents would not be accepted by maintainers, after adjustment for noisy human decisions. The study used selected repositories and gave agents no opportunity to revise work after feedback. It does not show that half of current Codex work fails. It shows that passing benchmark tests does not establish that a maintainer would accept a change. [E5]
Financial features need both software checks and Dave’s approved business examples. The checks test consistent behavior. The examples test whether the calculation means what the business intends. A screenshot does not prove complete reconciliation or financial correctness.
4. Measure Whether The Work Saves Dave Time
The finding that “AI made developers 19% slower” should not be treated as a current general result. METR’s February 2026 update says its new experiment has selection and measurement problems and likely misses important gains. Its May survey reports increased perceived value. That survey distinguishes value from speed and warns that self-reported benefits may be overstated. [E6] [E7]
Track Dave’s review time, corrections, interruptions, and recurring work removed by accepted features. Use those results to choose his workflow. Building a dashboard faster does not necessarily help if he has little use for it.
5. Credentials And Access Limits
Doppler distributes secrets and supports config-scoped service tokens. Storing a password or token there does not ensure that another provider’s OAuth grant, browser session, device approval, or MFA stays valid. Doppler also warns that cached fallback secrets can remain available after a Doppler token is revoked. [E8]
Give the agent only the access needed for the task, rather than “all the logins and tools it wants.” A development worker should use development credentials. Keep payment authority and unrelated browser accounts inaccessible. A second device with the same broad credentials does not provide meaningful security isolation.
Adam’s reported credential overhead supports reducing unnecessary integrations and separate login setups before copying the environment to another machine. Reuse the smallest setup that reliably completes the intended work.
6. PC Maintenance And Scheduled Jobs
Microsoft recommends checking resource-heavy processes and storage, managing startup applications, and keeping supported software current. It also notes that software tuning may do little for old hardware. This guidance does not establish that a bot can improve a healthy PC’s performance every night. [E9]
Nightly checks should make changes only when measurements identify a problem with a known repair. A check may find that nothing needs changing. Repeated cache deletion can force rebuilds. Uncontrolled dependency updates can break projects. Browser cleanup can remove required logins.
OpenAI says scheduled local-project tasks require the computer to be on and the app running. It warns that unattended tasks with full access carry elevated risk. Check how a job handles sleep, permissions, missing files, and expired sessions before relying on its schedule. [E2]
ForgeFX Research
Internal Findings
Accounting Already Competes With Production Work
Historical production messages describe accounting work taking time from John Deere and other production tasks. They establish a recurring conflict, but do not quantify the current time lost. [I2]
In a separate AI discussion, Dave described a successful pedal-electronics exercise. He explored a simulator with Codex before discovering that the simulator could not support the actual task. He reported substantial time savings, but those estimates were not independently measured. [I3]
Use representative examples when working together. A blinking-LED exercise may not reveal whether a simulator can handle the intended circuit. For ForgeBooks, begin with an actual accounting exception so the team can check whether the proposed approach handles the work Dave needs.
Dave Previously Reported Slow Image Iterations
Dave previously reported slower image iterations in Codex than in browser ChatGPT. Adam challenged ForgeBot’s agreeable explanation and pointed out the cost of switching tools. Dave chose to keep Codex as his default while learning, with the browser as a fallback if problems recurred. [I4]
The earlier discussion shows that Dave has raised performance concerns before and that the team values continuity. It does not establish the cause of that slowdown, a general speed difference, or the cause of the current crashes. ForgeBot’s earlier explanation was speculation, not independent evidence.
ForgeBooks Currently Uses The Workbook As Its Authority
ForgeBooks has reporting code and integrations. Its maintenance instructions make Dave’s budget workbook authoritative and require one-way copying into the app. The P&L hook also allows edits to downstream transaction rows. Those edits could conflict with the routine that restores agreement with the workbook. The source code supports this concern; no live overwrite incident was reproduced. [I5]
Decide whether Dave needs a better view of the workbook, an easier way to enter its inputs, or a replacement for it. Reuse ForgeBooks where it fits. Replacing the workbook would require an explicit change to the current operating rules.
The checked-in Reconciliation page creates a mock result and displays a completion message. A separate maintenance reconciliation pipeline exists. The mock page does not demonstrate that pipeline working. Reporting screens, bank-feed staging, and passing unit tests do not establish a complete general ledger or an accountant-approved migration. [I5]
Before expanding use, verify access controls for the intended user, clear labels for real and demo data, source freshness, approved financial examples, and safe handling of changes. Source review raised control questions. Detailed security observations remain in the local research notes rather than this link-accessible preview. Deployed controls and current test results were not verified.
Meeting Evidence About The Intended Accounting Tool
Direct Fireflies retrieval found a September 8 discussion in which Dave wanted software to replace his complicated spreadsheet. A separate September 8 production-budget meeting still used the spreadsheet. An earlier production-budget discussion had introduced ForgeBooks to staff. [I6]
Dave may want to improve budget work rather than replace tax, payroll, or double-entry bookkeeping. His full intended scope remains unclear. A narrow operational tool could be justified. Recommending ForgeBooks still requires a decision about which system will hold the authoritative records.
Working Approach
How To Organize The Work
Diagnose Failures On The Primary PC
Start with one project, one selected execution environment, and one active coding job. Distinguish an interface crash, command failure, stalled cloud response, and operating-system failure. During an actual incident, record the app version, time, logs, and CPU, RAM, and disk pressure.
Keep the primary PC if normal use is stable. If reproducible resource contention interrupts billable work, try the same workload on the existing backup PC before buying hardware. If failures follow the project or account, investigate those before replacing the computer. This research had no hardware inventory or incident logs from Dave’s PC.
Choose One Financial Task
Extend shared capabilities where they meet the need. The ForgeBooks name alone does not mean it should handle every accounting requirement. Keep the established accounting system authoritative until the responsible owner approves a change.
A read-only reconciliation view or exception report may solve the problem without creating a new ledger. A separate prototype can cheaply test an uncertain workflow with synthetic data. It must not become a second production accounting system without an explicit decision.
Keep Dave’s Review Simple
Keep the morning review. Let Dave give ordinary feedback and have the agent propose the priority order. Dave approves the next outcome. Keep that change’s scope fixed during implementation and save new ideas for later.
For the pilot, allow one active change and one item awaiting review. These are proposed limits, not measured findings. Use a short joint session when business requirements are unclear, and prohibit production financial writes. If Dave misses a morning, the agent can improve tests, reproduce known bugs, or gather evidence within the agreed scope. It should not keep adding features that await his review.
What To Include In The Morning Review
- Result: one sentence explaining what changed for the user.
- Evidence: the exact preview and version, a before-and-after example, and test results.
- Open questions: assumptions, failures, and business decisions still needed.
- Decision: accept, request a specific revision, or defer.
- Next task: one proposed change.
Set Permissions For Each Type Of Work
- Read-only research: allow inventories, searches, comparisons, and draft tests to run without supervision from the start.
- Sandbox implementation: allow scoped edits with test data, reproducible tests, and a preview for review.
- Larger batches: allow these only after comparable work is repeatedly accepted without major correction and rollback has been demonstrated.
- Privileged actions: apply separate explicit controls to production writes, financial rules, credential changes, and operating-system maintenance.
Increase unattended work when the agent repeatedly meets the task’s requirements with little correction. An eight-hour run alone does not demonstrate reliability. It could include useful work, repeated attempts, or unresolved problems.
Define Allowed PC Maintenance
- Monitor: free disk space, abnormal process load, failure logs, scheduled-job results, and required tool availability.
- Allow: approved, reversible cleanup of disposable outputs and known recovery actions, after checking for active work.
- Require separate approval: operating-system or driver changes, broad package upgrades, service changes, restarts during work, and permission changes.
- Protect: repositories, uncommitted work, browser profiles, credentials, security controls, and financial data.
- Verify: save before-and-after observations and check that the original task still works. Report when no action is needed.
Proposed Pilot
Test One Recurring Accounting Problem
Use the requested meeting to choose one recurring accounting problem Dave wants solved. Agree on a representative example and who owns the decision. This research did not send an invite, change a PC, create scheduled tasks, or change accounting records.
Dave’s Input
Bring one sanitized input, the correct expected result, one difficult exception, and a short explanation of what takes time today. Clarify whether the app is for ForgeFX, another business, or personal use. Shared infrastructure does not authorize mixing organizations’ data.
Adam’s Decisions
Confirm the authoritative source, a suitable ForgeBooks module or limited extension, permitted users, the approver for financial logic, and how to roll back a bad change. Delegate routine review where practical so work does not have to wait for Adam unnecessarily.
The Agent’s Deliverable
Implement one small read-only or sandboxed change. Provide reproducible tests, a user-visible preview, and a list of unresolved issues. Save the approved procedure and regression examples in the project for reuse. Ordinary conversations do not retrain model weights.
Measures Of Success
Record human minutes spent directing and reviewing work, repeated corrections, interruptions, recurring work removed, and any financial mismatch. Expand only if the workflow saves more human time than it requires. Otherwise, simplify it before adding automation.
Glossary
The Thread In Plain English
- Codex
- OpenAI’s coding agent, available through different interfaces. The interface, model, and place where tools execute are separate choices.
- Local / cloud / remote
- Local means on the working computer. Cloud means a provider-hosted environment. Remote means another machine reached over a network; it is not necessarily a cloud service.
- Inference
- The model generating an answer or next action. Local file editing does not imply local inference.
- Agent / bot
- A model connected to tools and a control loop so it can do work, not only produce text.
- Synchronous
- Working together in real time, with rapid questions and corrections.
- Asynchronous
- Giving work a clear handoff and reviewing it later rather than staying present throughout.
- Batch / nightly job
- A collection of work done between reviews, often on a schedule. Scheduling does not itself make the work reliable.
- Skill
- A reusable procedure, often saved as instructions plus scripts, examples, and references. A skill is not the same as retraining the underlying model.
- Context / context window
- The information available to the model for a particular turn or session. Older details may be omitted or summarized.
- Harness
- The machinery around the model: tools, permissions, task loop, saved state, tests, and recovery behavior.
- Shared infrastructure
- Common code, hosting, identity, integrations, and operational practices used by multiple features or people.
- Coordination overhead
- Time spent keeping people, agents, machines, and versions aligned instead of doing the primary work. One machine reduces some forms, not all.
- Context switching
- Reconstructing what you were doing after moving between tasks or tools. This can cost attention even when the new tool is faster.
- Doppler / secret
- Doppler is a secrets-management service. A secret is sensitive access material such as an API key or token; it is not always a human password.
- OAuth / MFA / service account
- OAuth grants limited application access without handing over a password. MFA adds another verification factor. A service account is a non-human identity for an application or automation.
- Least privilege / blast radius
- Give only the access needed. Blast radius is how much harm a mistake or compromise could cause.
- Sandbox / worktree / branch
- A sandbox limits what code can reach. A Git worktree is a separate working folder. A branch tracks a line of changes. A worktree alone is not a security boundary.
- Acceptance criteria / regression test
- Acceptance criteria define what “done” means. A regression test checks that previously correct behavior stays correct.
- System of record / ledger / reconciliation
- The authoritative source; the accounting record of transactions; and the process of comparing records and explaining differences.
- Idempotent
- Safe to repeat without creating another copy of the same effect—for example, not importing the same invoice twice.
- Production / staging / rollback
- The real user environment; a separate test environment; and a way to return to a known working state.
- Human approval gate
- A point where a responsible person must approve a consequential action. It need not mean watching every implementation step.
- First penguin
- Adam’s metaphor for being the first to try an uncertain path and absorb the early mistakes so others can take a simpler route.
- One-man band / bang for buck
- A setup with few handoffs; and useful benefit relative to cost. Here, cost includes human attention and maintenance, not just hardware price.
- Billable work / John Deere RWC
- Client work that earns revenue; and the John Deere project/workstream referenced in the thread. The expansion of “RWC” was not verified and is deliberately not guessed.
Evidence Register
Sources, Confidence, And Limits
Research date: September 10, 2026. Product documentation is a live snapshot, not a guarantee about Dave’s installed version or plan. Findings labeled as recommendations are synthesis, not reported test results.
External Sources
- OpenAI: Windows desktop guidance — live documentation, retrieved September 10, 2026. Native/WSL behavior and environment differences. High confidence about documented behavior, not Dave’s configuration.
- OpenAI: Scheduled tasks — live documentation, retrieved September 10, 2026; currently branded ChatGPT Learn. Local-project availability, review, worktrees, and unattended permissions. High.
- Anthropic: Effective harnesses for long-running agents — November 26, 2025. Incremental tasks, saved progress, end-to-end checks. High for the reported experiment; transfer to this team is an inference. Companion implementation was inspected as supporting code provenance, not executed.
- Anthropic: Harness design for long-running application development — March 24, 2026. Generator/evaluator separation and structured handoffs. Vendor experiment, not a head-to-head Codex benchmark.
- METR: Many SWE-bench-passing PRs would not be merged — March 10, 2026. Independent maintainer review; older agents and limited repositories. High for that sample, limited generalization.
- METR: Changing the developer productivity experiment — February 24, 2026. Selection effects undermine a simple contemporary speed estimate. High.
- METR: Self-reported impact of early-2026 AI — May 11, 2026. Value versus speed and self-report limitations. High about survey findings; not causal evidence of Dave’s gains.
- Doppler: Service tokens — live documentation, retrieved September 10, 2026. Restricted config access, expiration, and cached fallback caveat. High.
- Microsoft: Tips to improve PC performance in Windows — live guidance, retrieved September 10, 2026. Resource monitoring and supported maintenance. High; not proof of this PC’s fault.
Internal Sources
- The complete source discussion — September 10, 2026. Parent plus all replies available at collection. Dave’s original statement is present as Adam’s quotation. High confidence in what the thread says; operational success claims remain self-reports.
- Dave: accounting displaces John Deere work — historical production message. Supported by other retrieved production messages, including another explicit scheduling conflict. High for the historical statements, not a current time estimate.
- Dave: pedal-electronics workflow and the simulator limitation. Direct user reports; claimed savings not independently audited.
- Prior Codex image-latency discussion, including Adam’s challenge and Dave’s eventual default. Full text thread read; attached screenshot not used as evidence. High for reported experience; root cause unknown.
I5 · ForgeBooks implementation and operating contract. Read-only source snapshot, September 10, 2026. /Users/forgebot/forgeapps/crons/routines/forgebooks-maintenance/ROUTINE.md, lines 27–46: workbook-first mirror. /Users/forgebot/forgeapps/apps/forgebooks/src/hooks/useProfitAndLossDataWithDateRange.ts, lines 340–444: downstream edit path. /Users/forgebot/forgeapps/apps/forgebooks/src/pages/ReconciliationPage.tsx, lines 35–64: mock completion. High confidence in source contents; deployment state unverified. Local detailed notes: /Users/forgebot/forgeapps/artifacts/research/dave-codex/internal-findings.md.
I6 · Fireflies, directly searched and retrieved. September 8, 2026 discussion, statement at 950.12 seconds about replacing spreadsheet work; September 8 production-budget meeting, opening exchange at 5.2–9.2 seconds confirms spreadsheet use; February 3 production-budget meeting, 2128.45 seconds: ForgeBooks introduced. Source links require appropriate access. Only ordinary workflow meaning is summarized; transcripts, business amounts, and unrelated meeting content are excluded.
Fireflies queries covered ForgeBooks, FreshBooks, bookkeeping, Production Budget, accounting software, and QuickBooks. Access succeeded. Partner-only meetings were excluded, and irrelevant hits were not used as supporting evidence. This was bounded discovery, not an exhaustive meeting archive review.
Coverage And Deliberate Limits
Direct Slack searches covered the current channel, ForgeBooks, Dave’s accounting and Codex references, and the original wording. The connector excludes DMs, group DMs, and channels where ForgeBot is not a member. Broad searches surfaced some unrelated private material; none is reproduced here. Search results are not a workspace-wide census.
External research prioritized official documentation, primary engineering reports, independent studies, and GitHub implementation evidence. Video results were discovered but not used to establish product behavior. Social commentary was not necessary to settle the main claims. No additional paid model-provider research calls were made. Email, calendar, CRM, SharePoint, and production databases were not exhaustively searched; they are not represented as checked.
Unresolved: Dave’s exact PC/app version, incident logs, backup-PC specifications, his app’s repository and requirements, the actual authoritative accounting source for his proposed workflow, and current production acceptance under Dave’s identity. This is a research and decision artifact, not a security audit, accounting certification, or verified performance diagnosis.