FIELD NOTES · AUGUST 2026

Why I Built My Personal AI Around Structured Memory

I use AI assistants for work and daily life, and I kept running into the same failures. They forgot details I had already explained, duplicated code they had written, pulled in context that did not apply, and kept using facts that had expired. I eventually stopped treating these as separate bugs. In my system, they were mostly routing problems: the information had no stable home, or the agent had no reliable path back to it. The design I ended up with has two rules. Each piece of information or behavior has one canonical location, and every task that needs it has an explicit route there. The model can follow this structure. What it still struggles with is changing the structure without duplicating or orphaning things.

Yuchen Zeng Addressing · routing · agent architecture

01

Why I built my own personal assistant

I use Codex, Claude Code, and Copilot CLI every day. Each of them can hold on to a rough picture of who I am for a while. But none of them carries the right rules and details across all the things I ask it to do. I wanted an assistant I could actually hand work to without explaining my preferences, accounts, and past decisions from scratch each time.

1.1 It does not know enough about me

Booking a flight and sending an email are standard assistant demos. Both sound easy until the assistant has to make the small personal choices I normally make without thinking.

That is where the handoff breaks.

Scenario 1. Booking a flight

I ask it to book a flight. It needs to know which cards I hold and which one to pay with, which airline I collect miles on, and what I am optimizing for. It knows none of that. The results are not the flights I would pick, so I end up booking the trip myself.

Scenario 2. Sending an email

I ask it to email someone. It doesn't know which of my accounts to send from, and I keep several for different purposes. It can't find the recipient’s address without being told. The draft comes back in the wrong register, so I specify the tone as well. By the time I have specified all of it, I have basically written the email.

In both cases, the work came back to me. The assistant needed context that existed in my head but nowhere it could reach.

So I started writing down the context and the decision paths needed for each kind of task. The difficult part wasn't calling tools. It was deciding where all that memory should live and making sure the right task could find it.

1.2 It does not maintain the project

Long projects expose a different problem. Most of the work isn't writing the next block of code. It is keeping the new work consistent with what is already there. Agents often miss that. They implement the current request in isolation and quietly create a second version of something the project already owns.

I run into this whenever I write a paper.

Scenario 3. The second experiment

I write down the experiment design and ask it to implement the code. The first experiment comes out fine. The second one is supposed to run the same evaluation. Instead of calling what is already there, it writes its own, and some detail never lines up with the first. The two sets of numbers then cannot go in the same table.

The evaluation should be written once and called by every experiment. The agent won't reliably do that unless the project tells it where the evaluation lives and requires reuse.

The same thing happens when I ask it to save something for later.

Scenario 4. Saving a project idea

While working on a project, I find an idea that will matter later and ask the agent to save it. It creates idea.md in the project and writes the idea there. But it does not add a pointer from the project instructions to the new file. The file is on disk, but the agent never reads it again. It saved the idea without maintaining the path that would make the idea useful later.

After a few rounds of this, the project fills with files that exist but are never read.

The individual output was fine. The project around it was not.

1.3 It brings up what I did not ask about

Memory can also be correct and still be wrong for the task.

Scenario 5. An air fryer for my parents

I ask it which air fryer to buy. It comes back with the small ones, because I live in an apartment and I told it once that my counter space is tight. That is true, and it is not the point. The air fryer is for my parents, and their kitchen is nothing like mine. The ones I would have picked are already off the list, ruled out for the wrong kitchen.

Nothing here is false. That is exactly the problem. The assistant found a real preference, attached it to the wrong person, and produced a confident answer for the wrong kitchen.

1.4 It holds on to what is no longer true

Sometimes the retrieved memory is not merely irrelevant. It has stopped being true.

Scenario 6. A deadline that passed months ago

One week I was up against a deadline and I told it so. That went into memory. Months later the deadline is long gone, and it is still answering as though I am short on time. Cut the scope, ship something that works for now, leave the rest. I never told it the deadline had passed, because it never occurred to me that I would have to.

Saving something takes one sentence. Removing it requires me to notice that the system is still using it. Usually I don't notice until an answer has already been built around the old fact.

Retrieval needs more than similarity. The system has to know who a fact is about, when it applies, and whether it is still current.

1.5 It answers like anyone else

This one is different. Personal memory is not domain expertise. Sometimes the assistant knows my context and still gives me the answer of a median person. If I bring it a career question while I am upset, reassurance isn't useful. I want an answer from someone who understands the field.

Negotiation is the clearest case.

Scenario 7. Pushing back on a rent increase

I ask how to push back on a rent increase. What I get is the advice anyone would give: be polite, explain the situation, ask whether there is any flexibility. It has no read on the person on the other side. What they are worried about, what a specific number does to the conversation once it has been said out loud, when saying nothing is worth more than another sentence of justification.

Those details decide whether the negotiation works. Generic advice does not contain them.

02

Two rules that made the system work

I didn't fix these problems by putting more context into every prompt. I changed where information lived and how the agent reached it. The turning point came from a duplicated set of rules for making figures.

My paper-writing module had a list of rules for making figures: how big the fonts have to be, which colours still work when the paper is printed in black and white, what to do when the legend does not fit. I built that list slowly, by fixing figures that had come out badly. Later I added a module for blog posts, which also needed figures. Instead of reusing the existing rules, the assistant wrote a second list. The two copies looked close enough, so I left them alone. A few months later one of my paper figures came out unreadable, and I added a new rule to the list in the paper skill. The next blog post had the same unreadable figure. The blog skill was reading its own list, and that list had never been updated. I had already solved this problem once, and I had to solve it again.

The assistant had not forgotten anything. Both lists were on disk, and both modules read the file they were told to read. The mistake was storing one idea in two places. Updating either copy could never update the other.

The fix has two parts. Put the figure rules in one file. Then make both modules point to it. A shared file without a route is never opened; two routed copies still drift apart.

These became the two rules for the rest of the system:

  1. One owner for each piece of information. Every fact, rule, and reusable behavior has one canonical home. Not two, and not “probably somewhere”.
  2. An explicit route to that owner. When a task needs the information, the routing layer names the file that holds it. Otherwise the file can sit on disk forever without being read.

The global instruction sits at the root of the system. It sends each request to a specific file, which can name the next file to open. Several routes can converge on the same owner, so the structure is often closer to a graph than a tree. The important part is that the route is written down. The model does not have to guess where the next piece of context lives.

Here is what those routes look like for the seven examples above.

2.1 Case 1: Sending one email

An email needs three kinds of context: the recipient, the sending account, and the tone I use with that person. Those details have different owners, so the request follows several routes before the draft is written.

“email my collaborator about the draft” .copilot/ ├── config.json ├── copilot-instructions.md ├── projects.md ├── references/ │ ├── ai-mcp-skills-tips.md │ ├── coupons.md │ ├── personal/ │ │ ├── finance/ │ │ ├── identity_documents/ │ │ ├── living/ │ │ ├── network/ │ │ │ ├── family/ │ │ │ ├── friends/ │ │ │ ├── README.md │ │ │ ├── relationships/ │ │ │ └── work/ │ │ │ ├── <name>/ │ │ │ │ └── README.md │ │ │ └── … │ │ └── … │ └── voice-typos.md ├── settings.json ├── skills/ │ ├── account-ops/ │ │ ├── assets/ │ │ └── SKILL.md │ ├── email-ops/ │ │ ├── assets/ │ │ ├── credentials/ │ │ ├── references/ │ │ │ ├── gmail/ │ │ │ ├── outlook/ │ │ │ ├── outreach/ │ │ │ ├── sender-routing.md │ │ │ ├── writing-guide.md │ │ │ └── … │ │ └── SKILL.md │ ├── experiment-ops/ │ ├── negotiation/ │ ├── people-ops/ │ │ └── SKILL.md │ ├── travel-ops/ │ ├── workplace-communication/ │ │ ├── references/ │ │ │ ├── difficult-conversations.md │ │ │ ├── recipient-guide.md │ │ │ ├── reddit-boss-communication.md │ │ │ ├── reorg-analysis.md │ │ │ └── … │ │ └── SKILL.md │ └── … ├── vault.json └── … 1 2 3 4 5 6 7 8 9 10
One email, and the eleven files it opens, in the order it opens them. Hover a numbered step to see which two files it connects and why that edge exists. Nothing here is found by searching: every file is named by the one before it. A grey … means the folder holds more than the figure shows.

No search step has to guess what comes next. Each file names the next one.

2.2 Case 2: Booking a flight

A flight search needs my airline preferences, payment options, identity documents, and the procedure for comparing flights. The travel module owns only the procedure. The personal facts stay with their own files.

“book me a flight to the conference” .copilot/ ├── copilot-instructions.md ├── references/ │ ├── ai-mcp-skills-tips.md │ ├── coupons.md │ ├── personal/ │ │ ├── finance/ │ │ │ ├── payment.gpg │ │ │ ├── payment.README.md │ │ │ ├── README.md │ │ │ └── … │ │ ├── identity_documents/ │ │ │ ├── passport.pdf │ │ │ ├── signature.png │ │ │ ├── visa.pdf │ │ │ └── … │ │ ├── living/ │ │ ├── network/ │ │ ├── travel/ │ │ │ ├── packing.md │ │ │ ├── README.md │ │ │ └── trips/ │ │ │ ├── <date>-<trip>/ │ │ │ │ └── README.md │ │ │ └── README.md │ │ └── … │ └── voice-typos.md ├── skills/ │ ├── account-ops/ │ │ ├── assets/ │ │ └── SKILL.md │ ├── email-ops/ │ ├── negotiation/ │ ├── people-ops/ │ ├── purchase-ops/ │ │ └── SKILL.md │ ├── travel-ops/ │ │ ├── flights/ │ │ │ ├── money-saving.md │ │ │ ├── output.md │ │ │ └── README.md │ │ ├── manage-reservation.md │ │ ├── packing.md │ │ ├── parents/ │ │ └── SKILL.md │ └── … ├── vault.json └── … 1 2 3 4 5 6 7 8 9
Booking a flight. The travel module holds the procedure and none of the facts — it says so in its own first line. My travel file carries a table of where each kind of fact is allowed to live, and three of the steps above are simply that table being obeyed. The booking is written back to one place.

A table in the travel module says which file owns each fact the booking procedure needs.

2.3 Case 3: The second experiment

The second experiment failed even though all the necessary code already existed. The agent needed a rule that forced it to inspect and reuse the current implementation before writing anything new.

“now implement the second experiment” .copilot/ ├── config.json ├── copilot-instructions.md ├── projects.md ├── references/ ├── settings.json ├── skills/ │ ├── browser/ │ ├── code-writing/ │ │ ├── references/ │ │ │ ├── bugfix.md │ │ │ └── code-architecture.md │ │ └── SKILL.md │ ├── email-ops/ │ ├── env-control/ │ ├── experiment-ops/ │ │ ├── assets/ │ │ ├── references/ │ │ │ ├── experiment.md │ │ │ └── testing.md │ │ └── SKILL.md │ └── … ├── vault.json └── … 1 2 3 4 5 6
The second experiment. The rule that stops the metric being written twice is not in the experiment module at all — it is in the code module’s architecture file, and it runs before any code is written. The experiment module deliberately hands over and asks for control back.

That rule lives in the coding module, which is opened before any code change begins.

2.4 Case 4: Saving an idea

Saving an idea to a file is not enough. If nothing points to that file, the next session will never know it exists.

“save this idea for later” .copilot/ ├── config.json ├── copilot-instructions.md ├── projects.md ├── references/ ├── settings.json ├── skills/ │ ├── browser/ │ ├── email-ops/ │ ├── people-ops/ │ ├── system-maintenance/ │ │ ├── audit-checklist.md │ │ ├── file-directory.md │ │ ├── persist.md │ │ ├── register-project.md │ │ ├── SKILL.md │ │ └── … │ └── … ├── vault.json └── … 1 2 3 4 5
Saving an idea. One module owns every file that gets created, and it produces two things, not one: the file, and a pointer that makes the file reachable. A file nothing points at is the same as a file that does not exist.

So saving requires two writes: the new file and a pointer from a file the agent already knows to open.

2.5 Case 5: An air fryer for my parents

The air-fryer request should use facts about my parents, not facts about my apartment. Filing preferences under the person they describe makes that distinction explicit.

“which air fryer should I get my parents” .copilot/ ├── copilot-instructions.md ├── references/ │ ├── ai-mcp-skills-tips.md │ ├── coupons.md │ ├── personal/ │ │ ├── finance/ │ │ ├── living/ │ │ ├── network/ │ │ │ ├── family/ │ │ │ │ ├── dad/ │ │ │ │ │ ├── documents/ │ │ │ │ │ └── README.md │ │ │ │ └── mom/ │ │ │ │ ├── documents/ │ │ │ │ └── README.md │ │ │ ├── friends/ │ │ │ ├── README.md │ │ │ ├── relationships/ │ │ │ └── work/ │ │ ├── travel/ │ │ └── … │ └── voice-typos.md ├── skills/ │ ├── email-ops/ │ ├── negotiation/ │ ├── people-ops/ │ │ └── SKILL.md │ ├── purchase-ops/ │ │ └── SKILL.md │ ├── travel-ops/ │ └── … ├── vault.json └── … 1 2 3 4 5
An air fryer for my parents. Look at what stays dark: living/ is my own apartment, and it is exactly the file that would produce a confident answer about the wrong kitchen. Because facts are filed by who they are about, no edge leads there. It is never opened, so it never has to be ranked or ignored.

Nothing on this route leads to my apartment, so its constraints never enter the answer.

2.6 Case 6: A deadline that passed

A deadline is temporary state, not a permanent fact about me. It needs an expiration path from the moment it is stored.

“the form is due at the end of the month” .copilot/ ├── copilot-instructions.md ├── references/ │ ├── ai-mcp-skills-tips.md │ ├── coupons.md │ ├── personal/ │ │ ├── finance/ │ │ ├── identity_documents/ │ │ ├── living/ │ │ ├── network/ │ │ ├── README.md │ │ ├── reminders.json │ │ └── … │ └── voice-typos.md ├── skills/ │ ├── email-ops/ │ ├── people-ops/ │ ├── reporting-ops/ │ │ ├── assets/ │ │ │ ├── crawl-emails.md │ │ │ ├── get-calendar.md │ │ │ ├── manage-reminders.md │ │ │ └── … │ │ ├── references/ │ │ ├── SKILL.md │ │ └── … │ ├── system-maintenance/ │ │ ├── file-directory.md │ │ ├── persist.md │ │ ├── SKILL.md │ │ └── … │ ├── travel-ops/ │ └── … ├── vault.json └── … 1 2 3 4 5 6
Where something I say ends up. Three destinations, chosen by how long it stays true: a dated reminder that stays quiet until its date, a durable file with one canonical home, or nothing at all. The third is deliberate — for facts whose real source can change without telling me, a copy would be worse than no copy.

Deadlines go into the reminder file, where every entry has a date and a status. They don't sit beside long-lived personal facts.

2.7 Case 7: Pushing back on a rent increase

The rent question did not need another fact about me. It needed a negotiation process, evidence about the local market, and a check for weak moves before the draft reached me.

“how do I push back on the rent increase” .copilot/ ├── config.json ├── copilot-instructions.md ├── projects.md ├── references/ ├── settings.json ├── skills/ │ ├── email-ops/ │ ├── negotiation/ │ │ ├── references/ │ │ │ ├── anti-patterns.md │ │ │ ├── books/ │ │ │ │ ├── getting-to-yes/ │ │ │ │ ├── influence/ │ │ │ │ └── never-split-the-difference/ │ │ │ │ ├── chapters/ │ │ │ │ ├── index.md │ │ │ │ └── source.epub │ │ │ ├── framework.md │ │ │ ├── rent.md │ │ │ └── salary.md │ │ └── SKILL.md │ ├── people-ops/ │ ├── purchase-ops/ │ └── … ├── vault.json └── … 1 2 3 4 5
A negotiation question. The route ends in a check rather than an answer: the draft is tested against a list of known bad moves and rewritten until it passes. The advice underneath is not the assistant’s own — every tactic in the framework cites the chapter it came from, and those chapters are sitting in the tree.

The draft is checked against that list and rewritten if it gives away my bargaining position.

03

What the system looks like now

After several years of use, the system has two main layers: long-lived personal information and modules that act on it.

3.1 What it knows about me

I keep long-lived personal information in nine areas. The table shows what each area is for and how many files it contains, but not the files themselves.

Area Files What it holds
living/ 614 apartment, floor plans, car, furniture inventory, subscriptions
network/ 49 one folder per person, with a lookup table at the top; work, friends, family
health/ 13 one folder per test type, reports named by year, summary at the top
identity_documents/ 13 passport, visa, work authorization, student records
career/ 9 employment documents, CV, equity statements, long-range career planning
finance/ 9 one folder per institution; card and account numbers in an encrypted file
aesthetic-care/ 4 service providers: who, where, what was done, what it cost
travel/ 4 stable preferences, a reusable packing baseline, one folder per trip
education/ 3 degree certificates, with only one copy in the system

The areas are intentionally different in size. living/ is about one hundred times larger than education/. The folders give information an address; they are not meant to be balanced. Four flat files, academic.md, career.md, identity.md, and side-business.md, existed before the directory layout. They still sit next to the directories. Two have the same name as a directory, which creates an ambiguity I have not resolved yet.

3.2 What it can do

The system has 72 modules, each responsible for one kind of task. The list below shows 70 of them. I omit two names and use generic labels for the four marked with ∗.

Running the system

  • system-maintenance: the only module that creates, moves, or deletes system parts; owns the address tree.
  • session-history-ops: finds past sessions and transcripts.
  • copilot-config: changes the default model and reasoning level.
  • mcp-ops: installs, removes, and fixes tool servers.
  • sync-all: pushes every tracked project to its remote repository.
  • clean-copilot: runs agent experiments in an isolated container with selected configuration.
  • reporting-ops: runs the daily briefing and weekly review that can trigger system changes.

Communication and people

  • email-ops: reads, searches, drafts, and sends email across four accounts.
  • teams-ops, slack-ops, loop-ops: handle work chat and shared documents.
  • sms-ops, wechat-bot: handle personal messages.
  • apple-notification: sends urgent alerts to a phone and watch.
  • people-ops: writes to the network/ area, with one folder for each person.
  • workplace-communication: helps with difficult work conversations and replies.
  • negotiation: handles rent, salary, offers, and contract terms; this is the module behind Scenario 7.
  • social-scout: collects opinions from online communities when real experiences matter.

Research and writing

  • paper-ops: handles outlines, drafts, figures, rebuttals, and final versions.
  • theory-writing: sets rules for theorems, assumptions, and proofs.
  • notation-management: keeps a list of symbols and checks it before writing formulas.
  • overleaf-ops: syncs projects and fixes compilation, figure, and layout problems.
  • openreview-ops: handles review assignments, submissions, and rebuttals.
  • poster-ops, html-slide-ops, office-docs-ops: create posters, slides, and office documents.
  • pdf-ops: reads PDFs and fills forms without form fields.
  • de-ai-tone: removes writing that sounds machine-generated.
  • academic-profile-ops, download-cv: update the homepage, scholar profile, and CV.
  • add-book: splits a book by chapter and files it under the module that uses it.
  • compliance-process∗: handles internal review for publications and open-source releases.

Compute and experiments

  • env-control: chooses where code runs and blocks experiments on the laptop.
  • experiment-ops: designs, submits, evaluates, and reports experiments.
  • code-writing: the only module that controls code changes, code placement, and function contracts.
  • gpu-cluster-ops∗, chtc-ops: manage two clusters with different schedulers and queue rules.
  • job-scheduler-cli∗: submits jobs, reads logs, and downloads results.
  • azure-ops: stores artifacts that are too large for experiment tracking.
  • wandb-ops: finds experiment runs and downloads results.

Money

  • asset-snapshot: reads live balances and never estimates or rounds them.
  • investment-ops: handles reviews, contributions, vesting, and stock screening.
  • spending-ops: tracks transactions, monthly bills, and card rewards.
  • purchase-ops: handles checkout and asks for confirmation before payment.
  • tax-filing: collects tax documents, checks residency, and tracks deadlines.
  • expense-ops∗: creates work expense reports and matches card charges.
  • adp: downloads payroll documents when needed instead of storing them.
  • finance-ops: keeps old instructions working by routing them to the current module.

Life administration

  • calendar-ops: reads a local snapshot of four calendars before any API call and writes to one calendar.
  • travel-ops: handles flights, points, packing lists, and reservations.
  • address-change-ops: lists and completes every update required after a move.
  • apartment-search: compares rent after adding fees that advertisements leave out.
  • home-renovation-ops: reviews floor plans, cabinet drawings, and contractor work.
  • health-ops: handles health portals, lab reports, and prescriptions.
  • account-ops: stores credentials and runs before any login.
  • legal-advisor: supports every legal conclusion with a primary source.
  • high-stakes: uses a slower read, check, fix, and reread process for important documents.
  • internet-debugging: finds network problems and manages router settings.
  • secondhand-ops, facebook-ops: manage resale listings and buyer messages.
  • browser: starts web work with curl, then uses fetch or a browser when needed.
  • ghostty-ops: manages terminal settings.

Personal

  • personal-development: gives rules for emotional control and personal decisions based on classical sources.
  • relationship-advisor: answers relationship questions using one specific school of thought.
  • beauty-search: researches skincare, treatments, fashion, and gifts.

Public presence

  • xhs: writes posts, creates cover images, and reviews post results.
  • twitter-ops, linkedin-ops: write and score announcement drafts.
  • social-media-development: plans how to grow public accounts.
  • notion-ops: publishes content to a shared workspace.

04

What the system still cannot do

The system follows an architecture well once someone has designed it. It is much less reliable at deciding how that architecture should change. Every new fact, rule, or piece of code raises a placement question. Should it go into an existing file? Does a branch need to split? Is there already an abstraction to reuse? I call this architectural reasoning: making a local change without creating a second owner, breaking an existing route, or leaving a new component disconnected. For memory, the same problem can be described as address-space induction.

I want to know whether a model can learn these decisions from a sequence of changes. Give it an existing system and one new requirement. It must choose what to reuse, what to split, and where the new code or information belongs, then repair every route and caller affected by that choice. The same two rules provide a simple test: one owner for each piece, and a route from every task that needs it. The difficult part is keeping both rules true after the hundredth change, not just the first.