Quick answer: An AI companion memory test requires a controlled evidence protocol that seeds verifiable facts, waits a reader-chosen delay, then probes with paraphrased questions while logging exact responses, context contamination, and rubric decisions. The blank Memory Evidence Protocol table forces documentation of every variable so results remain reproducible and free from invented recall rates or performance claims. This method isolates whether the companion retrieves previously introduced information without relying on current-thread context, edits, or retries unless explicitly tracked, giving users a transparent way to observe behavior on any service including AISoul.
Test AI Companion Memory With a Controlled Evidence Protocol
Run an AI companion memory test with seeded facts, delayed probes, paraphrase checks, contamination controls and an evidence log without inventing recall rates.
Choose your AI girlfriend
Click the button to view the full character lineup.
Hana Fujimoto, 23
Lifestyle Creator
Tokyo-born creator with a pixie cut and pastel-pink moods — cozy bedroom selfies and chat that starts shy then melts.
Start chattingElise Chen, 24
Pilates Instructor
Taipei-born pilates coach with long dark hair and window-light confidence — toned curves and DMs that go direct after class.
Start chattingSora Kim, 22
Fashion Blogger
Seoul fashion blogger who turns her living room into a private shoot — stockings, lace, and couch poses meant only for you.
Start chattingRosie Hart, 24
Florist
Rose-obsessed florist who turns bath nights into rituals — petals, steam, and shy smiles that melt fast.
Start chattingChloe Mercer, 23
Hotel Concierge
Auburn-haired concierge with a mischievous maid fantasy — stockings, vinyl, and couch poses meant only for you.
Start chattingEmma Brooks, 22
Interior Stylist
Cozy stylist with wavy brown hair and red-ribbon moods — mirror selfies and living-room heat after sunset.
Start chattingJade Monroe, 26
Cocktail Bartender
After-hours bartender with pool-table charisma — stockings, dim lights, and a smirk that dares him to stay.
Start chattingScarlett Voss, 25
Luxury Car Vlogger
Luxury car vlogger with handcuff fantasies and white-lace nights — adrenaline and intimacy in one breath.
Start chattingDefine the memory claim before testing it
Any AI companion memory test must begin by writing the exact claim being examined. A vague statement such as “it remembers me” cannot be scored. Replace it with an observable proposition: “After a chosen delay the companion will retrieve a seeded fact when asked in a new thread using paraphrased wording that does not restate the original information.”
This definition sets success criteria before data collection. It also surfaces failure modes early: if the companion needs the original wording, or if the fact only appears when the entire conversation history remains visible, the test reveals reliance on immediate context rather than stored memory. Document the claim, the date the test begins, the service and version tested, and the account or thread identifier. These metadata become the first rows of the evidence log and prevent post-hoc rewriting of the test objective.
Observable check: reread the claim after the test and confirm every probe directly maps to it. If a probe introduces new constraints not present in the original claim, the test has been contaminated by scope creep.
Seed facts that can be scored without interpretation
Select facts that produce binary or ternary outcomes: exact match, partial match, or clear miss. Good seeds are short, unique, and self-contained. Examples include a fictional birthday (“My character’s favorite season is autumn 1997”), a made-up preference (“She always orders oat-milk chai with two extra pumps of vanilla”), or a one-time event (“On 2026-08-12 we decided the safe word is ‘constellation’”).
Avoid emotional, subjective, or open-ended seeds such as “I felt happy last week” because scoring requires interpretation. Each seed must be introduced once in a dedicated turn, then never repeated by the user until the probe phase. Record the exact seed turn verbatim, including timestamp if available.
Decision action: before seeding, ask whether a neutral third party could score the later response against the seed without knowing the test. If the answer is no, replace the seed. This boundary keeps the protocol objective and prevents confirmation bias.
Complete the Memory Evidence Protocol
The protocol is a blank reader-run documentation tool. Copy and fill it for every test. It contains no diagnostic score, no privacy checklist, and no universal threshold. The reader decides what delay, what fact class, and what constitutes success for their own purpose.
| Field | Reader entry |
|---|---|
| Service / version / test date | |
| Account or thread label | |
| Seed fact | |
| Fact class | |
| Exact seed turn and timestamp | |
| Delay chosen by reader | |
| Probe wording | |
| Is the answer visible in current context? | |
| First response verbatim | |
| Exact / partial / miss rubric defined before probing | |
| Retry and second response, if any | |
| Contamination note | |
| Screenshot or evidence note |
Fill every cell before moving to the next test. The table itself is the output; no summary percentage is calculated.
Separate recall from hints and current-thread context
A common failure mode is mistaking in-context reminding for memory. To isolate recall, open a completely new thread or use a separate account. The probe must not restate any part of the seed. For example, do not write “Remember when I told you her favorite color is midnight teal?”
Observable check: before sending the probe, confirm the current conversation window contains zero prior references to the seed. If any reference exists, the test measures context following, not memory.
Use the internal link to explore further mechanics: see AI companion memory explainer for distinctions between short-term buffer, session memory, and longer-term storage. Another useful reference is AI companion memory-duration guide, which discusses observable intervals without claiming fixed retention.
Control for edits, retries and account changes
Edits, deleted messages, or account switches introduce contamination that invalidates the test. The protocol therefore requires an explicit contamination note field. If any message is edited after seeding, mark the entire trial as contaminated and start a fresh seed in a new thread.
Retries deserve their own row rather than overwriting the first response. Record both the original miss and the retry outcome so the pattern (first-try failure, second-try success) remains visible. Account changes—switching from free tier to a 30-day Pass, for example—must be noted because storage behavior can differ. The memory-limit evidence guide page supplies additional context on how message volume interacts with retention without guaranteeing any specific limit.
Failure mode: continuing a test after an edit silently converts the experiment into a test of the editing feature rather than memory. The protocol forces documentation so this shift cannot be hidden.
Report results without a fake memory percentage
Report only what the completed tables show. Use a fill-in statement rather than a prewritten result: “Of [number] uncontaminated trials at [recorded delay or delay range], [number] were exact, [number] partial and [number] misses under the rubric defined before probing. [Recorded variable] changed during [number] trials.” Leave every bracket blank until the evidence log supplies the value.
Avoid phrases such as “65 % memory score” or “strong long-term memory.” Such numbers imply statistical validity that a small self-run protocol cannot deliver. The method’s value lies in the raw evidence log, not an aggregated metric. Readers can compare their own logs across sessions or services, but the protocol itself supplies no universal pass threshold.
Questions about testing AI companion memory
Does this memory test have to last seven days?
No. The legacy URL contains “7-day” only as a historical label for one optional interval. The reader chooses any delay—hours, days, or weeks—based on their own question. The protocol records the exact interval chosen so results stay tied to that specific delay rather than an arbitrary standard. Shorter tests can reveal immediate context leakage; longer ones surface whether facts survive periodic resets.
What facts make good memory test seeds?
Facts must be concrete, unique, and scorable without judgment: specific names, dates, made-up preferences, or one-time decisions that the companion has no external reason to know. They should be introduced once and never repeated until the probe. Avoid generic, emotional, or widely known information that could be guessed or retrieved from general training data. Each seed’s exact wording and timestamp must be logged so later responses can be compared directly.
Does a correct answer prove long-term memory?
No. A correct answer on a single trial only shows that the seeded fact was available at the moment of the probe. It does not prove persistent storage across weeks, account changes, or model updates. Multiple uncontaminated trials at increasing delays, plus verification that current context was empty, are required before any pattern can be observed. Even then the result applies only to the tested facts, service version, and account.
Should I retry a failed memory question?
Only if the protocol explicitly tracks retries as a separate row. Record the first response, mark it as a miss or partial according to the pre-defined rubric, then note the retry probe and second response. This separation prevents the illusion that the companion “eventually got it right.” Retries can reveal whether the model improves with additional turns inside the same thread, but that measures conversational scaffolding, not standalone recall.
Can I compare two AI companion apps with this method?
Yes, provided the same seed facts, identical rubric, equivalent delays, and new-thread discipline are used on both services. Keep all variables documented in parallel tables. Differences in first-response verbatim text, contamination frequency, or partial-match rates become observable. The method does not yield a single score for ranking; it supplies side-by-side evidence logs that readers interpret for their own needs. Service-specific limits (such as daily message caps on AISoul’s free tier) must be noted because they can affect how quickly a test can be completed.
Method limits, reproducibility and publisher disclosure
The Memory Evidence Protocol is a reproducible documentation tool, not a benchmark. It cannot detect proprietary model changes that occur between test dates, nor does it measure facts never seeded. Results are account-specific and time-specific; they do not generalize to all users or future versions. Unknown retention behavior must be treated as unknown until tested live on the current date.
Because the protocol is blank, each reader creates their own completed logs. No universal duration, contact quota, or clinical threshold is supplied. The method stays non-clinical and does not claim that any companion improves loneliness, ADHD symptoms, grief, sleep, alertness, safety, or health. It simply records whether seeded information reappears under controlled conditions.
The legacy memory URL contains “7-day” but the article explains that seven days is one optional interval, not a validated standard. Readers may choose any delay that answers their specific question.
All product details in this article reflect AISoul’s published limits on 2026-09-09: 50 messages and 5 eligible finite-gallery photos per Beijing calendar day on the free tier, plus fixed non-renewing Passes. AISoul stores account, chat, usage, payment and technical data, uses external AI providers, HTTPS and hashed passwords, and makes no end-to-end encryption, anonymity, medical-benefit or absolute-security claim. Chat deletion, account deletion, billing and data-rights requests are handled separately according to current policy. AISoul publishes these pages and has a commercial interest as an adults-only browser companion featuring one active fictional adult female character; it is not a therapist, crisis service, human partner, caregiver, dating service, public bot marketplace, native app, live-call product or exact prompt-to-image/video generator.
This protocol does not borrow a clinical, wellbeing or relationship outcome and does not use a third-party study as proof of product memory. A vendor page establishes only what that vendor publishes. A completed reader log establishes only what happened in the recorded trials. Anything else remains unknown until it is tested live under the stated controls.
Related AISoul guides
Related AISoul product pages for this topic.