The longer the task, the dumber Claude Code gets — goals drift, specs vanish from context.
So I built hplan: a skill that persists plans to the file system and progressively discloses only what Claude needs at each step.
https://t.co/GUBUSsADJG
#ClaudeCode#Agents#Skills
hplan is now adapted for OpenClaw — integrates with OpenClaw's memory system and uses structured behavioral rules to keep the agent on track across sessions.
ClawHub: https://t.co/4PTIzWAa0w
#OpenClaw#Agents#Skills
The longer the task, the dumber Claude Code gets — goals drift, specs vanish from context.
So I built hplan: a skill that persists plans to the file system and progressively discloses only what Claude needs at each step.
https://t.co/GUBUSsADJG
#ClaudeCode#Agents#Skills
# RLHF is just barely RL
Reinforcement Learning from Human Feedback (RLHF) is the third (and last) major stage of training an LLM, after pretraining and supervised finetuning (SFT). My rant on RLHF is that it is just barely RL, in a way that I think is not too widely appreciated. RL is powerful. RLHF is not. Let's take a look at the example of AlphaGo. AlphaGo was trained with actual RL. The computer played games of Go and trained on rollouts that maximized the reward function (winning the game), eventually surpassing the best human players at Go. AlphaGo was not trained with RLHF. If it were, it would not have worked nearly as well.
What would it look like to train AlphaGo with RLHF? Well first, you'd give human labelers two board states from Go, and ask them which one they like better:
Then you'd collect say 100,000 comparisons like this, and you'd train a "Reward Model" (RM) neural network to imitate this human "vibe check" of the board state. You'd train it to agree with the human judgement on average. Once we have a Reward Model vibe check, you run RL with respect to it, learning to play the moves that lead to good vibes. Clearly, this would not have led anywhere too interesting in Go. There are two fundamental, separate reasons for this:
1. The vibes could be misleading - this is not the actual reward (winning the game). This is a crappy proxy objective. But much worse,
2. You'd find that your RL optimization goes off rails as it quickly discovers board states that are adversarial examples to the Reward Model. Remember the RM is a massive neural net with billions of parameters imitating the vibe. There are board states are "out of distribution" to its training data, which are not actually good states, yet by chance they get a very high reward from the RM.
For the exact same reasons, sometimes I'm a bit surprised RLHF works for LLMs at all. The RM we train for LLMs is just a vibe check in the exact same way. It gives high scores to the kinds of assistant responses that human raters statistically seem to like. It's not the "actual" objective of correctly solving problems, it's a proxy objective of what looks good to humans. Second, you can't even run RLHF for too long because your model quickly learns to respond in ways that game the reward model. These predictions can look really weird, e.g. you'll see that your LLM Assistant starts to respond with something non-sensical like "The the the the the the" to many prompts. Which looks ridiculous to you but then you look at the RM vibe check and see that for some reason the RM thinks these look excellent. Your LLM found an adversarial example. It's out of domain w.r.t. the RM's training data, in an undefined territory. Yes you can mitigate this by repeatedly adding these specific examples into the training set, but you'll find other adversarial examples next time around. For this reason, you can't even run RLHF for too many steps of optimization. You do a few hundred/thousand steps and then you have to call it because your optimization will start to game the RM. This is not RL like AlphaGo was.
And yet, RLHF is a net helpful step of building an LLM Assistant. I think there's a few subtle reasons but my favorite one to point to is that through it, the LLM Assistant benefits from the generator-discriminator gap. That is, for many problem types, it is a significantly easier task for a human labeler to select the best of few candidate answers, instead of writing the ideal answer from scratch. A good example is a prompt like "Generate a poem about paperclips" or something like that. An average human labeler will struggle to write a good poem from scratch as an SFT example, but they could select a good looking poem given a few candidates. So RLHF is a kind of way to benefit from this gap of "easiness" of human supervision. There's a few other reasons, e.g. RLHF is also helpful in mitigating hallucinations because if the RM is a strong enough model to catch the LLM making stuff up during training, it can learn to penalize this with a low reward, teaching the model an aversion to risking factual knowledge when it's not sure. But a satisfying treatment of hallucinations and their mitigations is a whole different post so I digress. All to say that RLHF *is* net useful, but it's not RL.
No production-grade *actual* RL on an LLM has so far been convincingly achieved and demonstrated in an open domain, at scale. And intuitively, this is because getting actual rewards (i.e. the equivalent of win the game) is really difficult in the open-ended problem solving tasks. It's all fun and games in a closed, game-like environment like Go where the dynamics are constrained and the reward function is cheap to evaluate and impossible to game. But how do you give an objective reward for summarizing an article? Or answering a slightly ambiguous question about some pip install issue? Or telling a joke? Or re-writing some Java code to Python? Going towards this is not in principle impossible but it's also not trivial and it requires some creative thinking. But whoever convincingly cracks this problem will be able to run actual RL. The kind of RL that led to AlphaGo beating humans in Go. Except this LLM would have a real shot of beating humans in open-domain problem solving.
@Yang_ML_Estate Similar to autonomous driving, it may become a hot topic, but if large-scale application is desired, many ethical and legal issues need to be resolved. It will take a considerable amount of time to reach maturity.
AGI is still quite far from human capabilities, requiring persistent investment and sufficient patience. Unlike visual models, killer applications for LLM have yet to emerge. It’s not only related to their capabilities, but also to their extremely high computational requirements, energy consumption, and hefty costs. During a period of global economic fluctuation and downgraded consumption, a business model that doesn't show clear improvement and has a steep cost curve can no longer convince consumers to pay.
Excited to share our latest work on discrete flow matching! A new framework that achieves SOTA non-autoregressive generation. For example, pass@1 on HumanEval is 6.7/11.6 and 6.7/13.1 on MBPP.
Paper: https://t.co/oJLgsBiyys
[1/n]
Tech progress will impact or eliminate some jobs. If it can't create new jobs at the same level, should we plan transitions or push ahead regardless? Does this tech advancement equal social progress?🙅♂️
MobileLLM: nice paper from @AIatMeta about running sub-billion LLMs on smartphones and other edge devices.
TL;DR: more depth, not width; shared matrices for token->embedding and embedding->token; shared weights between multiple transformer blocks;
Paper: https://t.co/TDWQWdZeIy
[Deep RNN] by Hand ✍️
A Deep Recurrent Neural Network (RNN) extends a basic single-layer RNN into multiple layers of hidden states, effectively incorporating deep learning into the RNN architecture.
How does a Deep RNN work?
[1] Given
↳ A sequence of four inputs X1, X2, X3, X4 ⬛️
↳ Recurrent weights and biases for hidden layers a 🟩, b 🟧, c 🟪, and the output layer y 🟦.
[2] Initialize Hidden States
↳ Set a0, b0, c0 to zeros
— Process X1 (t = 1)—
[3] First Hidden Layer (a) 🟩: a0 → a1
↳ The transformation matrix is horizontal concatenation of input weights, hidden state weights and biases, visualized as [⬛️ | 🟩 | ⬜️] .
↳ The state matrix is vertical concatenation of input X1, previous hidden state a0, and an extra 1, visualized as [⬛️ ; 🟩 ; 1].
↳ Multiply the two matrices to obtain new hidden state a1 = [0 ; 1].
[4] Second Hidden Layer (b) 🟪: b0 → b1
↳ First layer a1 🟩 becomes the input.
↳ The transformation matrix is visualized as [🟩 | 🟪 | ⬜️].
↳ The state matrix is the combination of a1, b0, and 1, visualized as [🟩; 🟪 ; 1].
↳ Multiply the two matrices to obtain new hidden state b1 = [1; -1].
[5] Third Hidden Layer (c) 🟧: c0 → c1
↳ Second layer b 🟪 becomes the input.
↳ The transformation matrix is visualized as [🟪 | 🟧 | ⬜️].
↳ The state matrix is the combination of a1, b0, and 1, visualized as [🟪; 🟧; 1].
↳ Multiply the two matrices to obtain new hidden state b1 = [1; -1].
[6] Output Layer (Y) 🟦
↳ The transformation matrix is visualized as [🟧 | ⬜️].
↳ The state matrix is the combination of c0 and , visualized as [🟧; 1].
↳ Multiply the two matrices to obtain output Y1 = [3; 0; 3].
— Process X2 (t = 2)—
[7] Previous Hidden States
↳ Copy the values of a1, b1, c1.
[8] Hidden 🟩🟪🟧 + Output 🟦
↳ Repeat [3]-[6] to obtain output Y2 = [5; 0; 4]
— Process X3 (t = 3)—
[9] Previous Hidden States
↳ Copy the values of a2, b2, c2.
[10] Hidden 🟩🟪🟧 + Output 🟦
↳ Repeat [3]-[6] to obtain output Y3 = [13; -1; 9]
— Process X4 (t = 4)—
[11] Previous Hidden States
↳ Copy the values of a3, b3, c3.
[12] Hidden 🟩🟪🟧 + Output 🟦
↳ Repeat [3]-[6] to obtain output Y4 = [15; 7; 2]
[CLIP] by Hand ✍️
The CLIP (Contrastive Language–Image Pre-training) model, a groundbreaking work by OpenAI, redefines the intersection of computer vision and natural language processing. It is the basis of all the multi-modal foundation models we see today.
How does CLIP work?
Goal: 🟨 Learn a shared embedding space for text and image
[1] Given
↳ A mini batch of 3 text-image pairs
↳ OpenAI used 400 million text-image pairs to train its original CLIP model.
Process 1st pair: "big table"
[2] 🟪 Text → 2 Vectors (3D)
↳ Look up word embedding vectors using word2vec.
[3] 🟩 Image → 2 Vectors (4D)
↳ Divide the image into two patches.
↳ Flatten each patch
[4] Process other pairs
↳ Repeat [2]-[3]
[5] 🟪 Text Encoder & 🟩 Image Encoder
↳ Encode input vectors into feature vectors
↳ Here, both encoders are simple one layer perceptron (linear + ReLU)
↳ In practice, the encoders are usually transformer models.
[6] 🟪 🟩 Mean Pooling: 2 → 1 vector
↳ Average 2 feature vectors into a single vector by averaging across the columns
↳ The goal is to have one vector to represent each image or text
[7] 🟪 🟩 -> 🟨 Projection
↳ Note that the text and image feature vectors from the encoders have different dimensions (3D vs. 4D).
↳ Use a linear layer to project image and text vectors to a 2D shared embedding space.
🏋️ Contrastive Pre-training 🏋️
[8] Prepare for MatMul
↳ Copy text vectors (T1,T2,T3)
↳ Copy the transpose of image vectors (I1,I2,I3)
↳ They are all in the 2D shared embedding space.
[9] 🟦 MatMul
↳ Multiply T and I matrices.
↳ This is equivalent to taking dot product between every pair of image and text vectors.
↳ The purpose is to use dot product to estimate the similarity between a pair of image-text.
[10] 🟦 Softmax: e^x
↳ Raise e to the power of the number in each cell
↳ To simplify hand calculation, we approximate e^□ with 3^□.
[11] 🟦 Softmax: ∑
↳ Sum each row for 🟩 image→🟪 text
↳ Sum each column for 🟪 text→ 🟩 image
[12] 🟦 Softmax: 1 / sum
↳ Divide each element by the column sum to obtain a similarity matrix for 🟪 text→🟩 image
↳ Divide each element by the row sum to obtain a similarity matrix for 🟩 image→🟪 text
[13] 🟥 Loss Gradients
↳ The "Targets" for the similarity matrices are Identity Matrices.
↳ Why? If I and T come from the same pair (i=j), we want the highest value, which is 1, and 0 otherwise.
↳ Apply the simple equation of [Similarity - Target] to compute gradients of for both directions.
↳ Why so simple? Because when Softmax and Cross-Entropy Loss are used together, the math magically works out that way.
↳ These gradients kick off the backpropagation process to update weights and biases of the encoders and projection layers (red borders).
We’ve trained a model, CriticGPT, to catch bugs in GPT-4’s code. We’re starting to integrate such models into our RLHF alignment pipeline to help humans supervise AI on difficult tasks: https://t.co/5oQYfrpVBu