FIELD NOTES · AUGUST 2026
Why I Built My Personal AI Around Structured Memory
I use AI assistants for work and daily life, and I kept running into the same failures. They forgot details I had already explained, duplicated code they had written, pulled in context that did not apply, and kept using facts that had expired. I eventually stopped treating these as separate bugs. In my system, they were mostly routing problems: the information had no stable home, or the agent had no reliable path back to it. The design I ended up with has two rules. Each piece of information or behavior has one canonical location, and every task that needs it has an explicit route there. The model can follow this structure. What it still struggles with is changing the structure without duplicating or orphaning things.
01
Why I built my own personal assistant
I use Codex, Claude Code, and Copilot CLI every day. Each of them can hold on to a rough picture of who I am for a while. But none of them carries the right rules and details across all the things I ask it to do. I wanted an assistant I could actually hand work to without explaining my preferences, accounts, and past decisions from scratch each time.
1.1 It does not know enough about me
Booking a flight and sending an email are standard assistant demos. Both sound easy until the assistant has to make the small personal choices I normally make without thinking.
That is where the handoff breaks.
Scenario 1. Booking a flight
I ask it to book a flight. It needs to know which cards I hold and which one to pay with, which airline I collect miles on, and what I am optimizing for. It knows none of that. The results are not the flights I would pick, so I end up booking the trip myself.
Scenario 2. Sending an email
I ask it to email someone. It doesn't know which of my accounts to send from, and I keep several for different purposes. It can't find the recipient’s address without being told. The draft comes back in the wrong register, so I specify the tone as well. By the time I have specified all of it, I have basically written the email.
In both cases, the work came back to me. The assistant needed context that existed in my head but nowhere it could reach.
So I started writing down the context and the decision paths needed for each kind of task. The difficult part wasn't calling tools. It was deciding where all that memory should live and making sure the right task could find it.
1.2 It does not maintain the project
Long projects expose a different problem. Most of the work isn't writing the next block of code. It is keeping the new work consistent with what is already there. Agents often miss that. They implement the current request in isolation and quietly create a second version of something the project already owns.
I run into this whenever I write a paper.
Scenario 3. The second experiment
I write down the experiment design and ask it to implement the code. The first experiment comes out fine. The second one is supposed to run the same evaluation. Instead of calling what is already there, it writes its own, and some detail never lines up with the first. The two sets of numbers then cannot go in the same table.
The evaluation should be written once and called by every experiment. The agent won't reliably do that unless the project tells it where the evaluation lives and requires reuse.
The same thing happens when I ask it to save something for later.
Scenario 4. Saving a project idea
While working on a project, I find an idea that will matter
later and ask the agent to save it. It creates
idea.md in the project and writes the idea there.
But it does not add a pointer from the project instructions to
the new file. The file is on disk, but the agent never reads it
again. It saved the idea without maintaining the path that
would make the idea useful later.
After a few rounds of this, the project fills with files that exist but are never read.
The individual output was fine. The project around it was not.
1.3 It brings up what I did not ask about
Memory can also be correct and still be wrong for the task.
Scenario 5. An air fryer for my parents
I ask it which air fryer to buy. It comes back with the small ones, because I live in an apartment and I told it once that my counter space is tight. That is true, and it is not the point. The air fryer is for my parents, and their kitchen is nothing like mine. The ones I would have picked are already off the list, ruled out for the wrong kitchen.
Nothing here is false. That is exactly the problem. The assistant found a real preference, attached it to the wrong person, and produced a confident answer for the wrong kitchen.
1.4 It holds on to what is no longer true
Sometimes the retrieved memory is not merely irrelevant. It has stopped being true.
Scenario 6. A deadline that passed months ago
One week I was up against a deadline and I told it so. That went into memory. Months later the deadline is long gone, and it is still answering as though I am short on time. Cut the scope, ship something that works for now, leave the rest. I never told it the deadline had passed, because it never occurred to me that I would have to.
Saving something takes one sentence. Removing it requires me to notice that the system is still using it. Usually I don't notice until an answer has already been built around the old fact.
Retrieval needs more than similarity. The system has to know who a fact is about, when it applies, and whether it is still current.
1.5 It answers like anyone else
This one is different. Personal memory is not domain expertise. Sometimes the assistant knows my context and still gives me the answer of a median person. If I bring it a career question while I am upset, reassurance isn't useful. I want an answer from someone who understands the field.
Negotiation is the clearest case.
Scenario 7. Pushing back on a rent increase
I ask how to push back on a rent increase. What I get is the advice anyone would give: be polite, explain the situation, ask whether there is any flexibility. It has no read on the person on the other side. What they are worried about, what a specific number does to the conversation once it has been said out loud, when saying nothing is worth more than another sentence of justification.
Those details decide whether the negotiation works. Generic advice does not contain them.
02
Two rules that made the system work
I didn't fix these problems by putting more context into every prompt. I changed where information lived and how the agent reached it. The turning point came from a duplicated set of rules for making figures.
My paper-writing module had a list of rules for making figures: how big the fonts have to be, which colours still work when the paper is printed in black and white, what to do when the legend does not fit. I built that list slowly, by fixing figures that had come out badly. Later I added a module for blog posts, which also needed figures. Instead of reusing the existing rules, the assistant wrote a second list. The two copies looked close enough, so I left them alone. A few months later one of my paper figures came out unreadable, and I added a new rule to the list in the paper skill. The next blog post had the same unreadable figure. The blog skill was reading its own list, and that list had never been updated. I had already solved this problem once, and I had to solve it again.
The assistant had not forgotten anything. Both lists were on disk, and both modules read the file they were told to read. The mistake was storing one idea in two places. Updating either copy could never update the other.
The fix has two parts. Put the figure rules in one file. Then make both modules point to it. A shared file without a route is never opened; two routed copies still drift apart.
These became the two rules for the rest of the system:
- One owner for each piece of information. Every fact, rule, and reusable behavior has one canonical home. Not two, and not “probably somewhere”.
- An explicit route to that owner. When a task needs the information, the routing layer names the file that holds it. Otherwise the file can sit on disk forever without being read.
The global instruction sits at the root of the system. It sends each request to a specific file, which can name the next file to open. Several routes can converge on the same owner, so the structure is often closer to a graph than a tree. The important part is that the route is written down. The model does not have to guess where the next piece of context lives.
Here is what those routes look like for the seven examples above.
2.1 Case 1: Sending one email
An email needs three kinds of context: the recipient, the sending account, and the tone I use with that person. Those details have different owners, so the request follows several routes before the draft is written.
No search step has to guess what comes next. Each file names the next one.
2.2 Case 2: Booking a flight
A flight search needs my airline preferences, payment options, identity documents, and the procedure for comparing flights. The travel module owns only the procedure. The personal facts stay with their own files.
A table in the travel module says which file owns each fact the booking procedure needs.
2.3 Case 3: The second experiment
The second experiment failed even though all the necessary code already existed. The agent needed a rule that forced it to inspect and reuse the current implementation before writing anything new.
That rule lives in the coding module, which is opened before any code change begins.
2.4 Case 4: Saving an idea
Saving an idea to a file is not enough. If nothing points to that file, the next session will never know it exists.
So saving requires two writes: the new file and a pointer from a file the agent already knows to open.
2.5 Case 5: An air fryer for my parents
The air-fryer request should use facts about my parents, not facts about my apartment. Filing preferences under the person they describe makes that distinction explicit.
living/ is my own apartment, and it is exactly the file that would produce a confident answer about the wrong kitchen. Because facts are filed by who they are about, no edge leads there. It is never opened, so it never has to be ranked or ignored.Nothing on this route leads to my apartment, so its constraints never enter the answer.
2.6 Case 6: A deadline that passed
A deadline is temporary state, not a permanent fact about me. It needs an expiration path from the moment it is stored.
Deadlines go into the reminder file, where every entry has a date and a status. They don't sit beside long-lived personal facts.
2.7 Case 7: Pushing back on a rent increase
The rent question did not need another fact about me. It needed a negotiation process, evidence about the local market, and a check for weak moves before the draft reached me.
The draft is checked against that list and rewritten if it gives away my bargaining position.
03
What the system looks like now
After several years of use, the system has two main layers: long-lived personal information and modules that act on it.
3.1 What it knows about me
I keep long-lived personal information in nine areas. The table shows what each area is for and how many files it contains, but not the files themselves.
| Area | Files | What it holds |
|---|---|---|
living/ |
614 | apartment, floor plans, car, furniture inventory, subscriptions |
network/ |
49 | one folder per person, with a lookup table at the top; work, friends, family |
health/ |
13 | one folder per test type, reports named by year, summary at the top |
identity_documents/ |
13 | passport, visa, work authorization, student records |
career/ |
9 | employment documents, CV, equity statements, long-range career planning |
finance/ |
9 | one folder per institution; card and account numbers in an encrypted file |
aesthetic-care/ |
4 | service providers: who, where, what was done, what it cost |
travel/ |
4 | stable preferences, a reusable packing baseline, one folder per trip |
education/ |
3 | degree certificates, with only one copy in the system |
The areas are intentionally different in size.
living/ is about one hundred times larger than
education/. The folders give information an
address; they are not meant to be balanced.
Four flat files,
academic.md, career.md,
identity.md, and side-business.md,
existed before the directory layout. They still sit next to the
directories. Two have the same name as a directory, which
creates an ambiguity I have not resolved yet.
3.2 What it can do
The system has 72 modules, each responsible for one kind of task. The list below shows 70 of them. I omit two names and use generic labels for the four marked with ∗.
Running the system
system-maintenance: the only module that creates, moves, or deletes system parts; owns the address tree.session-history-ops: finds past sessions and transcripts.copilot-config: changes the default model and reasoning level.mcp-ops: installs, removes, and fixes tool servers.sync-all: pushes every tracked project to its remote repository.clean-copilot: runs agent experiments in an isolated container with selected configuration.reporting-ops: runs the daily briefing and weekly review that can trigger system changes.
Communication and people
email-ops: reads, searches, drafts, and sends email across four accounts.teams-ops,slack-ops,loop-ops: handle work chat and shared documents.sms-ops,wechat-bot: handle personal messages.apple-notification: sends urgent alerts to a phone and watch.people-ops: writes to thenetwork/area, with one folder for each person.workplace-communication: helps with difficult work conversations and replies.negotiation: handles rent, salary, offers, and contract terms; this is the module behind Scenario 7.social-scout: collects opinions from online communities when real experiences matter.
Research and writing
paper-ops: handles outlines, drafts, figures, rebuttals, and final versions.theory-writing: sets rules for theorems, assumptions, and proofs.notation-management: keeps a list of symbols and checks it before writing formulas.overleaf-ops: syncs projects and fixes compilation, figure, and layout problems.openreview-ops: handles review assignments, submissions, and rebuttals.poster-ops,html-slide-ops,office-docs-ops: create posters, slides, and office documents.pdf-ops: reads PDFs and fills forms without form fields.de-ai-tone: removes writing that sounds machine-generated.academic-profile-ops,download-cv: update the homepage, scholar profile, and CV.add-book: splits a book by chapter and files it under the module that uses it.compliance-process∗: handles internal review for publications and open-source releases.
Compute and experiments
env-control: chooses where code runs and blocks experiments on the laptop.experiment-ops: designs, submits, evaluates, and reports experiments.code-writing: the only module that controls code changes, code placement, and function contracts.gpu-cluster-ops∗,chtc-ops: manage two clusters with different schedulers and queue rules.job-scheduler-cli∗: submits jobs, reads logs, and downloads results.azure-ops: stores artifacts that are too large for experiment tracking.wandb-ops: finds experiment runs and downloads results.
Money
asset-snapshot: reads live balances and never estimates or rounds them.investment-ops: handles reviews, contributions, vesting, and stock screening.spending-ops: tracks transactions, monthly bills, and card rewards.purchase-ops: handles checkout and asks for confirmation before payment.tax-filing: collects tax documents, checks residency, and tracks deadlines.expense-ops∗: creates work expense reports and matches card charges.adp: downloads payroll documents when needed instead of storing them.finance-ops: keeps old instructions working by routing them to the current module.
Life administration
calendar-ops: reads a local snapshot of four calendars before any API call and writes to one calendar.travel-ops: handles flights, points, packing lists, and reservations.address-change-ops: lists and completes every update required after a move.apartment-search: compares rent after adding fees that advertisements leave out.home-renovation-ops: reviews floor plans, cabinet drawings, and contractor work.health-ops: handles health portals, lab reports, and prescriptions.account-ops: stores credentials and runs before any login.legal-advisor: supports every legal conclusion with a primary source.high-stakes: uses a slower read, check, fix, and reread process for important documents.internet-debugging: finds network problems and manages router settings.secondhand-ops,facebook-ops: manage resale listings and buyer messages.browser: starts web work with curl, then uses fetch or a browser when needed.ghostty-ops: manages terminal settings.
Personal
personal-development: gives rules for emotional control and personal decisions based on classical sources.relationship-advisor: answers relationship questions using one specific school of thought.beauty-search: researches skincare, treatments, fashion, and gifts.
Public presence
xhs: writes posts, creates cover images, and reviews post results.twitter-ops,linkedin-ops: write and score announcement drafts.social-media-development: plans how to grow public accounts.notion-ops: publishes content to a shared workspace.
04
What the system still cannot do
The system follows an architecture well once someone has designed it. It is much less reliable at deciding how that architecture should change. Every new fact, rule, or piece of code raises a placement question. Should it go into an existing file? Does a branch need to split? Is there already an abstraction to reuse? I call this architectural reasoning: making a local change without creating a second owner, breaking an existing route, or leaving a new component disconnected. For memory, the same problem can be described as address-space induction.
I want to know whether a model can learn these decisions from a sequence of changes. Give it an existing system and one new requirement. It must choose what to reuse, what to split, and where the new code or information belongs, then repair every route and caller affected by that choice. The same two rules provide a simple test: one owner for each piece, and a route from every task that needs it. The difficult part is keeping both rules true after the hundredth change, not just the first.