Machine learning looks new. A few years ago we were impressed by spell-check; now large language models write code and argue about philosophy.
The full story is two and a half centuries long. It starts with a pastor reasoning about uncertainty, runs through chess champions and cat videos, and ends on a problem most people still ignore:
Models keep getting smarter. They still don’t remember you.
This is the history of machine learning, and why we think the next step is a memory layer that follows you across the apps, agents and devices you use. We are building that layer, and its next release is Ditto v2.
1763: The Probability Revolution
Long before computers, Thomas Bayes set out an idea machine learning still depends on. His theorem gave us a way to update beliefs as we see new evidence.
That idea, start with a guess, then refine it with data, is still central to modern AI. Spam filters, recommendation engines and language models all rely on some form of probabilistic reasoning.
Machine learning began with reasoning about uncertainty, long before neural networks.
1950s: Machines That Learn
Alan Turing asked the question first: can a machine learn?
In 1950 he proposed a “learning machine” that could become intelligent through experience. A few years later, Arthur Samuel built checkers programs that improved by playing themselves and coined the term “machine learning” in 1959.
Frank Rosenblatt’s 1957 perceptron was the first trainable neural network. It was simple and limited. In 1969, Marvin Minsky and Seymour Papert published Perceptrons, which set out those limits and helped trigger the first AI winter.
Early results did not guarantee useful systems.
1980s: Backpropagation Brings Neural Nets Back
Research picked up again in 1986 when researchers made backpropagation practical. For the first time, multi-layer neural networks could be trained efficiently.
Most of modern AI depends on this step. The algorithms later used for image recognition, speech synthesis and language models were now possible in principle. They still needed more data, more compute and a few decades of engineering.
1990s to 2000s: From Knowledge to Data
Machine learning moved from hand-coded rules to data-driven models. Support vector machines, random forests and kernel methods led the field.
In 1997, IBM’s Deep Blue beat Garry Kasparov at chess. It was a milestone. It also showed that brute-force search and specialized hardware could look intelligent without understanding much.
The bigger change came from data. In 2009, Fei-Fei Li’s team released ImageNet, a large labeled image dataset. Researchers finally had enough labeled data to train large models.
2012: Deep Learning Arrives
AlexNet won the ImageNet competition by a wide margin. It used deep neural networks trained on GPUs. Within a few years, deep learning went from an academic topic to the industry standard.
Machines soon recognized faces, translated languages and drove cars. These systems were still pattern matchers. They processed input, produced output and kept nothing between one input and the next.
2017: Attention Changes Everything
Google Brain’s transformer architecture made parallel training on sequential data practical. The paper Attention Is All You Need became one of the most cited in AI history.
Transformers led to GPT, BERT and the foundation models that followed. They made long-context conversation possible. A longer context is still short-term: the model sees more of the current chat, and none of it carries over to the next one.
2022: ChatGPT Makes AI Mainstream
When OpenAI released ChatGPT, conversational AI became a household product. Millions of people found good answers, and found that the chat kept nothing once they left it.
At launch, you closed the tab and started again. You could upload a file, explain your project and build context over an hour, then lose it, because the model had no persistent memory.
That was a design limit of the architecture, not a fault in one product.
The Gap History Keeps Ignoring
Each major chapter in ML history improved one thing: the model’s ability to recognize patterns in the current input.
- Bayes gave us inference.
- Neural networks gave us function approximation.
- Deep learning gave us rich representations.
- Transformers gave us coherent long-context generation.
None of them solved continuity. Most AI tools start a new chat without your context, or keep memory only inside that one tool. Each agent works alone. Your context sits in a dozen separate chat histories that never share anything.
A context window is wide and temporary. Memory grows, lasts between sessions, and connects ideas across everything you have discussed.
The next step is to give models a memory that lasts across sessions, models and tools, rather than to make the model larger.
Why Memory Comes Next
Real memory is more than a list of facts in a settings panel. It has to be:
- Continuous: it follows you across conversations, not just within one.
- Connected: it links ideas, projects and people into a knowledge graph.
- Visible: you can see what is remembered, search it and correct it.
- Portable: it moves with you, independent of any single model provider.
- Agent-aware: every agent and sub-agent you use can draw from the same context.
Hey Ditto is building this.
Ditto: The Memory Layer for Your Agents
Ditto is a memory layer between you and the models you use.
You talk to GPT-5, Claude, Gemini or Ditto’s own agents. Ditto saves the important parts, organizes them into a personal knowledge graph and brings back the right context when you need it. Threads keep project context separate. Sub-agents start with that context. MCP connectors link your memory to Gmail, Slack, GitHub, Google Workspace and more.
You can switch models and keep the same memory.
Your Memory, Visualized
Ditto stores conversations and also extracts subjects (people, projects, technologies, ideas) and connects them into an interactive knowledge graph.
It does this through what we call the dreaming pipeline: a memory-consolidation process, loosely modeled on how the brain turns the day into long-term memory during sleep. Ditto replays new conversations, extracts the subjects that matter, deduplicates them against what it already knows, and links everything into the graph.
You can search and explore the graph, and any connected agent can use it through MCP. Ask about “OAuth” and Ditto knows it is linked to “Firebase Auth” and “API Security” in your graph, so it pulls the related context rather than every mention of the word.
Writing the Next Chapter of ML History
The field that gave us Bayes, backprop and transformers is still moving, and Ditto publishes research in it.
Storing memories is easy. Finding the right one from a vague question is the hard research problem. When you ask Ditto something indirect, like “what did I decide about that thing a while back?”, there is no keyword to match. The query is fuzzy, and it runs against everything you have told Ditto.
So we ran the experiment. Our Seed Memories v4 research pairs a subject-graph signal with a tiny per-user adapter, a small network that learns the shape of your memory. Together they lift Recall@1 by 7.6 points, train on a CPU in seconds, and stay under a megabyte per user.
As in 1763, the gain came from better use of evidence, not from scale. Personalized retrieval on a shared encoder lets an assistant find the right memory, not only store it.
Memory That Does Things
Ditto can also act on what it remembers.
With Ditto Code, an agent uses what it already knows about you and your projects to build: web apps, dashboards and documents, deployed to a live URL, with code committed back to your GitHub. The agent shares your memory, so you don’t re-explain your stack each time. It picks up where you left off.
Agents now act on your context and remember what they did, where earlier models only recognized patterns.
Memory You Own
History points to one more requirement: continuity you control.
Models will keep changing, and providers come and go. Your memory can’t be locked inside any one of them. Ditto’s Memory Passport lets you export and inspect your memory and take it to other tools.
We’re taking that further with DittoBench and the open-source Ditto Harness, the scoring core of a Bittensor subnet (SN118) where independent miners compete to build the best agent-and-memory harness. The aim is a memory layer that does not depend on any single model lab.
What This Means for You
If you are a developer, you stop re-explaining your stack to every new chat. If you are a researcher, you stop losing track of papers, hypotheses and dead ends. If you are a founder, you stop rebuilding investor context every week. If you are a neurodivergent knowledge worker, you stop spending effort on managing context yourself.
AI should remember the routine details so you can think about the hard problems.
Ditto v2: Where It All Comes Together
Each chapter of this story added one piece: inference, representation, generation, continuity, personalization, action and ownership.
Ditto v2 combines those pieces into an agentic operating system. It runs across your apps, surfaces and devices, and every agent draws from the same persistent memory. It brings together the knowledge graph, the dreaming pipeline, per-user retrieval, agents that build, MCP connectors, a portable Memory Passport and a model-agnostic harness.
For two and a half centuries, each advance made the model smarter. Ditto v2 gives it a memory, and makes that memory available wherever you work.
Models will keep changing. Your memory, and everything built on it, stays with you.
Want an AI that remembers your context? Try Ditto free. It works with the models you already use.