🧬 Introducing the AI/ML for Genomics & Genetic Counseling newsletter! 🧬
Stay updated on how AI and ML are transforming clinical genomics.
Join us! Subscribe here -
https://t.co/ewphki0Civ - 1st issue out next week.
#Genomics#AI#MachineLearning#GeneticCounseling
Google just proved that pasting your exact prompt twice beats every advanced prompting technique.
One model jumped from 21% to 97% accuracy with zero extra effort.
Here's the science behind AI's simplest hack:
LLMs like GPT-4o and Claude read your prompt left to right, one token at a time. Early words can't "see" later words on the first pass.
When information appears in an awkward order, like answer choices before the actual question, the model struggles to connect the pieces.
By repeating the prompt, every word in the first copy becomes visible context for the second copy. The model gets a "second read" with perfect attention across your entire input.
Like reading an exam question twice before answering. Same effort, better comprehension.
The results were staggering ↓
Google tested Gemini 2.0 Flash, GPT-4o, GPT-4o-mini, Claude 3 Haiku, Claude 3 Sonnet, and DeepSeek V3 across seven benchmarks.
Across all 70 combinations, prompt repetition delivered 47 statistically significant wins with zero losses.
But here's the result that broke the scale:
Google gave Gemini 2.0 Flash-Lite a task where the AI had to find a specific detail buried inside a massive prompt.
Without repetition, it got the right answer just 21% of the time. With one copy-paste, accuracy jumped to 97%.
Same model. Same task. The only difference was reading it twice.
Skeptics had a theory: maybe longer inputs just help in general?
Google tested that too. They padded prompts with periods to match the length, and it made zero difference. Only meaningful repetition works.
And the best part?
It adds almost no extra processing time because the repeated input runs in parallel.
But there's one scenario where it doesn't help ↓
Chain-of-thought prompts like "Think step by step" already re-read and rephrase your question internally.
They're doing the repetition for you. Out of 28 reasoning tests, only 5 showed improvement.
This means prompt repetition is most powerful for:
• Classification and multiple choice tasks
• Data extraction and short-answer Q&As
• Direct-answer tasks where order sensitivity matters
The practical takeaway:
1) Take your exact prompt
2) Paste it twice as one message
3) Send it and change nothing else
This works across every major model family. Gemini, GPT, Claude, DeepSeek. All of them.
The simplest trick in AI history: if you want a better answer, just ask twice.
—
Thanks for reading!
Enjoyed this post?
Follow Big Brain AI for more content like this.
"Our findings reveal a link between oral [periodontal disease] microbes and breast cancer risk and progression, particularly in genetically susceptible [BRCA1 variant positive] individuals"
https://t.co/wbybDzWkEX
Google just made a $2,000 AI education free (and most people will still choose to stay illiterate).
They quietly launched Google Skills, a hub with 3,000 technical modules that replaces "prompting" fluff with actual DeepMind research workflows. You can now access the exact curriculum used to train their internal teams on transformer architecture for $0.
If you don’t use this, you’ll eventually complain about the people who did.
Andrej Karpathy literally built the neural networks running inside coding assistants.
He taught the world deep learning at Stanford. He ran AI at Tesla.
If he feels “dramatically behind” as a programmer… that tells you everything about where we are.
The confession here is that raw intelligence and deep technical knowledge no longer guarantee mastery. The new stack isn’t about understanding transformers or writing elegant algorithms. It’s about orchestrating a zoo of stochastic systems that nobody fully controls.
Karpathy’s list is revealing: agents, subagents, prompts, contexts, memory, modes, permissions, tools, plugins, skills, hooks, MCP, LSP, slash commands, workflows, IDE integrations. That’s 15+ new primitives that didn’t exist 18 months ago. Each one evolving weekly.
The mental model problem is real. Traditional engineering gives you deterministic systems. You write code, it does exactly what you wrote. Now you’re managing entities that are “fundamentally stochastic, fallible, unintelligible and changing.”
His “alien tool with no manual” framing is exactly right. We’re all reverse-engineering capabilities in real-time. The documentation is always out of date. The best practices from 3 months ago are already wrong.
The magnitude 9 earthquake isn’t coming. It already hit. The aftershocks are the new normal.
Unaffected sperm donor with gonadal mosaicism for a TP53 variant in 20% of his sperm cells (which causes the cancer syndrome Li Fraumeni syndrome) found to have donated sperm to over 197 families across Europe
https://t.co/qeg6TmT7mq
Good discussion on the ethics and potential social risks and benefits of polygenic embryo screening.
Prenatal genetics people will have to power through the intro because some inaccuracies are corrected later.
https://t.co/hTP6GB7kz4
"...Preventive was considering using the United Arab Emirates to conduct tests, as it is a country where embryo editing is legal."
https://t.co/8XBTZLNQvT
A new preprint reports a rare case of a boy with X-linked Fanconi anemia in whom the hematopoietic phenotype (bone marrow failure) was completely rescued by full triploid mosaicism (69, XXY) 🤯
Fanconi anemia (FA) is an inherited genetic disorder caused by mutations in DNA repair genes leading to bone marrow failure, congenital abnormalities, and increased cancer susceptibility.
The one of subtypes of FA is an X-linked disease caused by mutations in FANCB.
In a hospital in Melbourne, doctors one day encountered a remarkable case of a boy with FA. Initially he presented with clinical features suggestive of FA, along with positive family history with X-linked inheritance. Diagnostic work up confirmed a loss of function mutation in FANCB.
As any other child with FA, the doctors expected his clinical course to worsen with time, ultimately leading to bone marrow failure and requiring a bone marrow transplant. But then they noticed something strange!
As doctors monitored his bone marrow over time, they noticed an unusual shift in his cells' chromosomes:
At diagnosis: 99% of bone marrow cells were normal diploid (46,XY) - but FA-deficient
15 months later: 10% had become triploid (69,XXY)
25 months later: 95% triploid
28 months later: 100% triploid
Puzzlingly, his blood counts did not worsen during follow ups, and he didn't go into aplastic anemia.
A research team, led by Wayne Crismani at St Vincent's Institute in Melbourne, investigated the case and discovered something extraordinary.
Here's what happened:
The egg (that became this boy) carrying a defective FANCB gene on its X chromosome was fertilized by a normal Y-bearing sperm, creating a diploid (46, XY) embryo destined for FA.
But then something extremely rare happened: the second polar body from meiosis II that usually gets discarded during fertilization somehow got incorporated into early embryonic cells.
This polar body carried the mother's other X chromosome with a normal FANCB gene. Fusion of this polar body resulted in a mosaic embryo with two cell populations:
- diploid cells (46,XY) with a single copy of defective FANCB gene
- triploid cells (69,XXY) with two copies of FANCB gene--one defective and the other normal.
In most tissues, the triploid cells remained a minority. For example, his skin cells were 96% diploid. But in the bone marrow - where cells divide constantly and accumulate DNA damage - the triploid cells had a massive competitive advantage.
Over just two years, the triploid lineage spread through his entire bone marrow, replacing every cell with full triploid (69, XXY) genome.
The genomic fingerprint clearly showed that two X chromosomes in the triploid cells came from the mother: identical sequences near centromeres (where no cross-overs happen) and different sequences near telomeres (where cross-overs happen)
This pattern could only come from sister chromatids that separated during the mother's egg formation - one went into the egg, the other into the second polar body that was supposed to disappear but instead became this boy's lifeline.
This case represents an extraordinary case of "natural gene therapy" - a developmental anomaly injecting extra DNA inside cells treating a fatal genetic disease.
Sharp, Harris, et al. medRxiv 2025
https://t.co/9PMcJ5UHUF
yesterday, Hugging Face dropped a 214-page MASTERCLASS on how to train LLMs
> it’s called The Smol Training Playbook
> and if want to learn how to train LLMs,
> this GIFT is for you
> this training bible walks you through the ENTIRE pipeline
> covers every concept that matters from why you train,
> to what you train, to how you actually pull it off
> from pre-training, to mid-training, to post-training
> it turns vague buzzwords into step-by-step decisions
> architecture, tokenization, data strategy, and infra
> highlights the real-world gotchas
> instabilities, scaling headaches, debugging nightmares
> distills lessons from building actual
> state-of-the-art LLMs, not just toy models
how modern transformer models are actually built
> tokenization: the secret foundation of every LLM
> tokenizer fundamentals
> vocabulary size
> byte pair encoding
> custom vs existing tokenizers
> all the modern attention mechanisms are here
> multi-head attention
> multi-query attention
> grouped-query attention
> multi-latent attention
> every positional encoding trick in the book
> absolute position embedding
> rotary position embedding
> yaRN (yet another rotary network)
> ablate-by-frequency positional encoding
> no position embedding
> randomized no position embedding
> stability hacks that actually work
> z-loss regularization
> query-key normalization
> removing weight decay from embedding layers
> sparse scaling, handled
> mixture-of-experts scaling
> activation ratio tuning
> choosing the right granularity
> sharing experts between layers
> load balancing across experts
> long-context handling via ssm
> hybrid models: transformer plus state space models
data curation = most of your real model quality
> data curation is the main driver of your model’s actual quality
> architecture alone won’t save you
> building the right data mixture is an art,
> not just dumping in more web scrapes
> curriculum learning, adaptive mixes, ablate everything
> you need curriculum learning:
> design data mixes hat evolve as training progresses
> use adaptive mixtures that shift emphasis
> based on model stage and performance
> ablate everything: run experiments to systematically
> test how each data source or filter impacts results
> smollm3 data
> the smollm3 recipe: balanced english web data,
> broad multilingual sources, high-quality code, and diverse math datasets
> without the right data pipeline,
> even the best architecture will underperform
the training marathon
> do your preflight checklist or die
> check your infrastructure,
> validate your evaluation pipelines,
> set up logging, and configure alerts
> so you don’t miss silent failures
> scaling surprises are inevitable
> things will break at scale in ways they never did in testing
> vanishing throughput? that usually means
> you’ve got a hidden shape mismatch or
> batch dimension bug killing your GPU utilization
> sudden drops in throughput?
> check your software stack for inefficiencies,
> resource leaks, or bad dataloader code
> seeing noisy, spiky loss values?
> your data shuffling is probably broken,
> and the model is seeing repeated or ordered data
> performance worse than expected?
> look for subtle parallelism bugs
> tensor parallel, data parallel,
> or pipeline parallel gone rogue
> monitor like your GPUs depend on it (because they do)
> watch every metric, track utilization, spot anomalies fast
> mid-training is not autopilot
> swap in higher-quality data to improve learning,
> extend the context window if you want bigger inputs,
> and use multi-stage training curricula to maximize gains
> the difference between a good model and a failed run is
> almost always vigilance and relentless debugging during this marathon
post-training
> post-training is where your raw base model
> actually becomes a useful assistant
> always start with supervised fine-tuning (sft)
> use high-quality, well-structured chat data and
> pick a solid template for consistent turns
> sft gives you a stable, cost-effective baseline
> don’t skip it, even if you plan to go deeper
> next, optimize for user preferences
> direct preference optimization (dpo),
> or its variants like kernelized (kto),
> online (orpo), or adversarial (apo)
> these methods actually teach the model
> what “better” looks like beyond simple mimicry
> once you’ve got preference alignment,go on-policy:
> reinforcement learning from human feedback (rlhf)
> or on-policy distillation, which lets your model learn
> from real interactions or stronger models
> this is how you get reliability and sharper behaviors
> the post-training pipeline is where
> assistants are truly sculpted;
> skipping steps means leaving performance,
> safety, and steerability on the table
infra is the boss fight
> this is where most teams lose time,
> money, and sanity if they’re not careful
> inside every gpu
> you’ve got tensor cores and cuda cores for the heavy math,
> plus a memory hierarchy (registers, shared memory, hbm)
> that decides how fast you can feed data to the compute units
> outside the gpu, your interconnects matter
> pcie for gpu-to-cpu,
> nvlink for ultra-fast gpu-to-gpu within a node,
> infiniband or roce for communication between nodes,
> and gpudirect storage for feeding massive datasets
> straight from disk to gpu memory
> make your infra resilient:
> checkpoint your training constantly,
> because something will crash;
> monitor node health so you can kill or restart
> sick nodes before they poison your run
> scaling isn’t just “add more gpus”
> you have to pick and tune the right parallelism:
> data parallelism (dp), pipeline parallelism (pp), tensor parallelism (tp),
> or fully sharded data parallel (fsdp);
> the right combo can double your throughput,
> the wrong one can bottleneck you instantly
to recap
> always start with WHY
> define the core reason you’re training a model
> is it research, a custom production need, or to fill an open-source gap?
> spec what you need: architecture, model size, data mix, assistant type
> transformer or hybrid
> set your model size
> design the right data mixture
> decide what kind of assistant or
> use case you’re targeting
> build infra for the job, plan for chaos, pick your stability tricks
> build infrastructure that matches your goals
> choose the right GPUs
> set up reliable storage
> and plan for network bottlenecks
> expect failures, weird bugs,
> and sudden bottlenecks at scale
> select your stability tricks in advance:
> know which techniques you’ll use to fight loss spikes,
> unstable gradients, and hardware hiccups
closing notes
> the pace of LLM development is relentless,
> but the underlying principles never go out of style
> and this PDF covers what actually matters
> no matter how fast the field changes
> systematic experimentation is everything
> run controlled tests, change one variable at a time, and document every step
> sharp debugging instincts will save you
> more time (and compute budget) than any paper or library
> deep knowledge of both your software stack
> and your hardware is the ultimate unfair advantage;
> know your code, know your chips
> in the end, success comes from relentless curiosity,
> tight feedback loops, and a willingness to question everything
> even your own assumptions
if i had this two years ago, it would have saved me so much time
> if you’re building llms,
> read this before you burn gpu months
happy hacking
In @eLife, our OpenSpliceAI paper, led by @KuanHaoChao, is now 'official' though it's been online since July. If you want to enjoy the reviewers' comments and our responses, check it out at:
https://t.co/Jv7DfmClBa
My pleasure to come on Dwarkesh last week, I thought the questions and conversation were really good.
I re-watched the pod just now too. First of all, yes I know, and I'm sorry that I speak so fast :). It's to my detriment because sometimes my speaking thread out-executes my thinking thread, so I think I botched a few explanations due to that, and sometimes I was also nervous that I'm going too much on a tangent or too deep into something relatively spurious. Anyway, a few notes/pointers:
AGI timelines. My comments on AGI timelines looks to be the most trending part of the early response. This is the "decade of agents" is a reference to this earlier tweet https://t.co/NiSn6jftqq Basically my AI timelines are about 5-10X pessimistic w.r.t. what you'll find in your neighborhood SF AI house party or on your twitter timeline, but still quite optimistic w.r.t. a rising tide of AI deniers and skeptics. The apparent conflict is not: imo we simultaneously 1) saw a huge amount of progress in recent years with LLMs while 2) there is still a lot of work remaining (grunt work, integration work, sensors and actuators to the physical world, societal work, safety and security work (jailbreaks, poisoning, etc.)) and also research to get done before we have an entity that you'd prefer to hire over a person for an arbitrary job in the world. I think that overall, 10 years should otherwise be a very bullish timeline for AGI, it's only in contrast to present hype that it doesn't feel that way.
Animals vs Ghosts. My earlier writeup on Sutton's podcast https://t.co/rSp1noyGBr . I am suspicious that there is a single simple algorithm you can let loose on the world and it learns everything from scratch. If someone builds such a thing, I will be wrong and it will be the most incredible breakthrough in AI. In my mind, animals are not an example of this at all - they are prepackaged with a ton of intelligence by evolution and the learning they do is quite minimal overall (example: Zebra at birth). Putting our engineering hats on, we're not going to redo evolution. But with LLMs we have stumbled by an alternative approach to "prepackage" a ton of intelligence in a neural network - not by evolution, but by predicting the next token over the internet. This approach leads to a different kind of entity in the intelligence space. Distinct from animals, more like ghosts or spirits. But we can (and should) make them more animal like over time and in some ways that's what a lot of frontier work is about.
On RL. I've critiqued RL a few times already, e.g. https://t.co/mYrMFVdVDW . First, you're "sucking supervision through a straw", so I think the signal/flop is very bad. RL is also very noisy because a completion might have lots of errors that might get encourages (if you happen to stumble to the right answer), and conversely brilliant insight tokens that might get discouraged (if you happen to screw up later). Process supervision and LLM judges have issues too. I think we'll see alternative learning paradigms. I am long "agentic interaction" but short "reinforcement learning" https://t.co/2L7FiaoKsw. I've seen a number of papers pop up recently that are imo barking up the right tree along the lines of what I called "system prompt learning" https://t.co/df5mJDdN3C , but I think there is also a gap between ideas on arxiv and actual, at scale implementation at an LLM frontier lab that works in a general way. I am overall quite optimistic that we'll see good progress on this dimension of remaining work quite soon, and e.g. I'd even say ChatGPT memory and so on are primordial deployed examples of new learning paradigms.
Cognitive core. My earlier post on "cognitive core": https://t.co/q2s1ihGy0T , the idea of stripping down LLMs, of making it harder for them to memorize, or actively stripping away their memory, to make them better at generalization. Otherwise they lean too hard on what they've memorized. Humans can't memorize so easily, which now looks more like a feature than a bug by contrast. Maybe the inability to memorize is a kind of regularization. Also my post from a while back on how the trend in model size is "backwards" and why "the models have to first get larger before they can get smaller" https://t.co/6k0FZRGXsb
Time travel to Yann LeCun 1989. This is the post that I did a very hasty/bad job of describing on the pod: https://t.co/fQgqaXPyp6 . Basically - how much could you improve Yann LeCun's results with the knowledge of 33 years of algorithmic progress? How constrained were the results by each of algorithms, data, and compute? Case study there of.
nanochat. My end-to-end implementation of the ChatGPT training/inference pipeline (the bare essentials) https://t.co/SIetgyoKWN
On LLM agents. My critique of the industry is more in overshooting the tooling w.r.t. present capability. I live in what I view as an intermediate world where I want to collaborate with LLMs and where our pros/cons are matched up. The industry lives in a future where fully autonomous entities collaborate in parallel to write all the code and humans are useless. For example, I don't want an Agent that goes off for 20 minutes and comes back with 1,000 lines of code. I certainly don't feel ready to supervise a team of 10 of them. I'd like to go in chunks that I can keep in my head, where an LLM explains the code that it is writing. I'd like it to prove to me that what it did is correct, I want it to pull the API docs and show me that it used things correctly. I want it to make fewer assumptions and ask/collaborate with me when not sure about something. I want to learn along the way and become better as a programmer, not just get served mountains of code that I'm told works. I just think the tools should be more realistic w.r.t. their capability and how they fit into the industry today, and I fear that if this isn't done well we might end up with mountains of slop accumulating across software, and an increase in vulnerabilities, security breaches and etc. https://t.co/8556ESSpyY
Job automation. How the radiologists are doing great https://t.co/FVUI872dkD and what jobs are more susceptible to automation and why.
Physics. Children should learn physics in early education not because they go on to do physics, but because it is the subject that best boots up a brain. Physicists are the intellectual embryonic stem cell https://t.co/p72Elk8lPV I have a longer post that has been half-written in my drafts for ~year, which I hope to finish soon.
Thanks again Dwarkesh for having me over!
We’re excited to introduce Chai-2, a major breakthrough in molecular design.
Chai-2 enables zero-shot antibody discovery in a 24-well plate, exceeding previous SOTA by >100x.
Thread👇
The American Academy of Pediatrics has just released guidelines (https://t.co/2EkrtlwDPi) endorsing exome and genome sequencing as first-tier tests for children with intellectual disability and global developmental delay. Ambry applauds this important milestone. For families looking for answers, exome testing (https://t.co/0Shvh1daxH) and proactive reanalysis (https://t.co/ocQjJqKDiM) can help end the diagnostic odyssey.
#exome #CNS #PatientforLife #MoreAnswersforMorePatients
Happy to introduce AlphaGenome, @GoogleDeepMind's new AI model for genomics.
AlphaGenome offers a comprehensive view of the human non-coding genome by predicting the impact of DNA variations. It will deepen our understanding of disease biology and open new avenues of research.
🧠 "Meet your new medical ethicist: ChatGPT" by Daniel Sokol
Can a chatbot outperform clinicians on an ethics test? In this blog post, Daniel Sokol reports that ChatGPT scored 43/44 on a situational judgement test designed by medical ethics professors, beating most students, doctors, and legal advisers. Whilst fallible, ChatGPT demonstrated principlist reasoning, legal awareness, and professional insight, raising provocative questions about its emerging role as a 24/7 ethical advisor in clinical settings #medicalethics #artificialintelligence #bioethics #chatgpt
🔗 Read more here: https://t.co/QJbfGSSEGN
Type 1 diabetes is often called “insulin-dependent”💉
But now, a single infusion of allogeneic stem cell–derived islet cells led to insulin independence after a year in 10 of 12 patients with type 1 diabetes in an early-phase trial by $VRTX.
This is a massive win for biotech❗️
@michaelmina_lab@Dr_J_Eggington I should've said genetic/genomic tests, of which there are about 40k in the US, and I can only comment on these. But if half of these have questionable quality that's still a significant problem assuming the rest are flawless or only have "few" problems. Did you open the link?
Unfortunately many people misunderstand the impact this will have on patients longterm.
See @Dr_J_Eggington 's great explanations for more details -
https://t.co/IhfCpPYLrD
Huge win for the US people regarding FDA
A Federal Court just massively “struck down” an @US_FDA attempted power grab to oversee Lab Developed Tests and thus try to oversee the individual actions of laboratory physicians.
The LDT Final Rule -
(**a massive attempt at a power grab for FDA - during the last administration - IMO born from greed of the previous leaders of FDA’s diagnostics group, CDRH, after having a “taste of power” to regulate labs and physicians during COVID**)
- would have been a disaster for equitable access to healthcare and diagnostics across the U.S., a disaster to innovation and would have markedly driven up costs to consumers, as their money would have gone, essentially, to pay hefty FDA fees, and would have bankrupted some of the most innovative labs across the U.S.
For me, the most important statements in the courts scathing conclusion are:
“FDA's creative attempt to expand its jurisdiction under the FDCA fails... the Court will not go down that road... the plain text does not support FDA's assertion that laboratory developed test services are "devices" subject to the FDCA."
And
"At no time [since 1967] did Congress suggest that FDA could regulate such laboratories... FDA's strained reading of the FDCA flouts, rather than effectuates, Congress's intent"
This is a big victory for health. Laboratories have been helping keeping people healthier for decades and will keep doing so.