@DavidD_Chapman Isn't it the same for US as well. Where are the native Americans and red-indians there ? . Just because the white had a couple of boats to go to US, doesn't mean they own it
Agent Memory That Works Like Human Memory!
Hindsight is an agent memory system built to create smarter agents that actually learn over time.
Here's the biggest problem:
Existing open-source memory solutions rely heavily on RAG, vector databases, and knowledge graphs. These are great for searching for context, but they do not facilitate deep learning from past experiences.
Vectorize solves this with a fundamentally different approach, building a system that mirrors how humans form and strengthen long-term memory.
Here's the Key Architectural Difference:
Hindsight uses two core techniques to mimic human memory structure:
• TEMPR (Temporal Entity Memory Priming Retrieval): A mechanism focused on context-aware memory recall.
• CARA (Coherent Adaptive Reasoning Agents): A method dedicated to agent-specific reflection and reasoning over past actions.
When new information enters the system, Hindsight deconstructs it into multiple representations: time series, entity-relationship, semantic, and keyword structures.
This enables the system to retrieve the most relevant memories in the most efficient way possible.
Hindsight is the first system to surpass 90% accuracy on LongMemEval, the industry standard benchmark for evaluating long-term memory in AI systems, achieving a score of 91.4%.
It's 100% Open Source.
Link to the Github Repo in the comments!
- you are
- a random CS grad with 0 clue how LLMs work
- get tired of people gatekeeping with big words and tiny GPUs
- decide to go full monk mode
- 2 years later i can explain attention mechanisms at parties and ruin them
- here’s the forbidden knowledge map
- top to bottom, how LLMs *actually* work
- start at the beginning
- text → tokens
- tokens → embeddings
- you are now a floating point number in 4D space
- vibe accordingly
- positional embeddings:
- absolute: “i am position 5”
- rotary (RoPE): “i am a sine wave”
- alibi: “i scale attention by distance like a hater”
- attention is all you need
- self-attention: “who am i allowed to pay attention to?”
- multihead: “what if i do that 8 times in parallel?”
- QKV: query, key, value
- sounds like a crypto scam
- actually the core of intelligence
- transformers:
- take your inputs
- smash them through attention layers
- normalize, activate, repeat
- dump the logits
- congratulations, you just inferred a token
- sampling tricks for the final output:
- temperature: how chaotic you want to be
- top-k: only sample from the top K options
- top-p: sample from the smallest group of tokens whose probabilities sum to p
- beam search? never ask about beam search
- kv cache = cheat code
- saves past keys & values
- lets you skip reprocessing old tokens
- turns a 90B model from “help me I’m melting” to “real-time genius”
- long context hacks:
- sliding window: move the attention like a scanner
- infini attention: attend sparsely, like a laser sniper
- memory layers: store thoughts like a diary with read access
- mixture of experts (MoE):
- not all weights matter
- route tokens to different sub-networks
- only activate ~3B params out of 80B
- “only the experts reply” energy
- grouped query attention (GQA):
- fewer keys/values than queries
- improves inference speed
- “i want to be fast without being dumb”
- normalization & activations:
- layernorm, RMSnorm
- gelu, silu, relu
- they all sound like failed Pokémon
- but they make the network stable and smooth
- training goals:
- causal LM: guess the next word
- masked LM: guess the missing word
- span prediction, fill-in-the-middle, etc
- LLMs trained on the art of guessing and got good at it
- tuning flavors:
- finetuning: new weights
- instruction tuning: “please act helpful”
- rlhf: reinforcement from vibes and clickbait prompts
- dpo: direct preference optimization — basically “do what humans upvote”
- scaling laws:
- more data, more parameters, more compute
- loss goes down predictably
- intelligence is now a budget line item
- bonus round:
- quantization:
- post-training quantization (PTQ)
- quant-aware training (QAT)
- models shrink, inference gets cheaper
- gguf, awq, gptq — all just zip files with extra spice
- training vs inference stacks:
- deepspeed, megatron, fschat — for pain
- vllm, tgi, tensorRT-LLM — for speed
- everyone has a repo
- nobody reads the docs
- synthetic data:
- generate your own training set
- model teaches itself
- feedback loop of knowledge and hallucination
- welcome to the ouroboros era
- final boss secret:
- you can learn *all of this* in ~2 years
- no PhD
- no 10x compute
- just relentless curiosity, good bookmarks, and late nights
- the elite don’t want you to know this
- but now that you do
- choose to act
- start now
- build the models
A layered overview of key Agentic AI concepts.
Let’s understand it layer by layer.
1) LLMs (foundation layer)
At the core, you have LLMs like GPT, DeepSeek, etc.
Core ideas here:
- Tokenization & inference parameters: how text is broken into tokens and processed by the model.
- Prompt engineering: designing inputs to get better outputs.
- LLM APIs: programmatic interfaces to interact with the model.
This is the engine that powers everything else.
2) AI Agents (built on LLMs)
Agents wrap around LLMs to give them the ability to act autonomously.
Key responsibilities:
- Tool usage & function calling: connecting the LLM to external APIs/tools.
- Agent reasoning: reasoning methods like ReAct (reasoning + act) or Chain-of-Thought.
- Task planning & decomposition: breaking a big task into smaller ones.
- Memory management: keeping track of history, context, and long-term info.
Agents are the brains that make LLMs useful in real-world workflows.
3) Agentic systems (multi-agent systems)
When you combine multiple agents, you get agentic systems.
Features:
- Inter-Agent communication: agents talking to each other, making use of protocols like ACP, A2A if needed.
- Routing & scheduling: deciding which agent handles what, and when.
- State coordination: ensuring consistency when multiple agents collaborate.
- Multi-Agent RAG: using retrieval-augmented generation across agents.
- Agent roles & specialization: Agents with unique purposes
- Orchestration frameworks: tools (like CrewAI, etc.) to build workflows.
This layer is about collaboration and coordination among agents.
4) Agentic Infrastructure
The top layer ensures these systems are robust, scalable, and safe.
This includes:
- Observability & logging: tracking performance and outputs (using frameworks like DeepEval).
- Error handling & retries: resilience against failures.
- Security & access control: ensuring agents don’t overstep.
- Rate limiting & cost management: controlling resource usage.
- Workflow automation: integrating agents into broader pipelines.
- Human-in-the-loop controls: allowing human oversight and intervention.
This layer ensures trust, safety, and scalability for enterprise/production environments.
Agentic AI, as a whole, involves a stacked architecture, where each outer layer adds reliability, coordination, and governance over the inner layers.
RAG is not Memory for AI Agents.
5 AI memory engines to build agents that maintain long-term context and learn continuously:
(Last 2 released just this month)
1. Zep builds and queries temporally-aware knowledge graphs that evolve with every interaction.
100% Opensource.
Dear @whirlpool_india , I haven't seen a much worse after sales repair service than yours. Months since I placed fridge repair request, spent more money than an original fridge and still left with a garbage box. Service number : VAR20102468334
@whirlpool_india Already did. Escalated it multiple times on call. The CS representative said the service guy will reach out in next two hours and that was 4 days ago. Anyways, I have pinged you my number on dm. If you refer the service number posted in tweet, that would be great
Learning machine learning has never been easier.
Here’s a step-by-step guide to kickstart your journey:
1️⃣ Start with a solid foundation: Grab the free book "An Introduction to Statistical Learning." It’s an excellent resource for beginners.
2️⃣ Read & Implement: As you read each chapter, implement what you’ve learned in your favorite programming language (Python, R, etc.). This will help reinforce your understanding.
3️⃣ Leverage prebuilt libraries: Once you grasp the concept, use prebuilt libraries like Scikit-learn (Python) or others to simplify your workflow. You can often import these techniques with just one line of code!
4️⃣ Practice with real datasets: Test your skills with two key datasets:
• 🏠 California Housing for regression tasks
• 🔢 MNIST for classification tasks
5️⃣ Rinse & Repeat: Consistently practice, refine, and repeat this process with each technique you learn.
@tnishanthr The so-called protector of "constitution" forgets that this is not a federal subject and the constitution has specific provision of tax devolution based on finance comission report. Uneducated much?