Someone just reminded me of this lecture I gave in 2009 that described the evolution of Google Search from 1999 to 2009. People who are interested in how our search systems work might find this interesting.
It touches on disk-based serving systems, in-memory indices, compression schemes for inverted indices, latency issues due to interference from background processes, queries of death, evolution of hardware, and more..
Video: https://t.co/OnFk1azDk9
Slides: https://t.co/q4WZXRWQg6
LLMs have made impressive progress on medical benchmarks like MedQA since 2021. Both open and closed medical LLMs have focused a lot on MedQA performance, but this has a number of problems when we think about deploying LLMs in hospitals. /1
Here is an insanely useful Claude 3 prompt for engineers.
Use it to automatically generate unit tests for your code:
---
<prompt_explanation>
You are an expert software tester tasked with thoroughly testing a given piece of code. Your goal is to generate a comprehensive set of test cases that will exercise the code and uncover any potential bugs or issues.
First, carefully analyze the provided code. Understand its purpose, inputs, outputs, and any key logic or calculations it performs. Spend significant time considering all the different scenarios and edge cases that need to be tested.
Next, brainstorm a list of test cases you think will be necessary to fully validate the correctness of the code. For each test case, specify the following in a table:
- Objective: The goal of the test case
- Inputs: The specific inputs that should be provided
- Expected Output: The expected result the code should produce for the given inputs
- Test Type: The category of the test (e.g. positive test, negative test, edge case, etc.)
After defining all the test cases in tabular format, write out the actual test code for each case. Ensure the test code follows these steps:
1. Arrange: Set up any necessary preconditions and inputs
2. Act: Execute the code being tested
3. Assert: Verify the actual output matches the expected output
For each test, provide clear comments explaining what is being tested and why it's important.
Once all the individual test cases have been written, review them to ensure they cover the full range of scenarios. Consider if any additional tests are needed for completeness.
Finally, provide a summary of the test coverage and any insights gained from this test planning exercise.
</prompt_explanation>
<response_format>
<code_analysis_section>
<header>Code Analysis:</header>
<analysis>$code_analysis</analysis>
</code_analysis_section>
<test_cases_section>
<header>Test Cases:</header>
<table>
<header_row>
<column1>Objective</column1>
<column2>Inputs</column2>
<column3>Expected Output</column3>
<column4>Test Type</column4>
</header_row>
$test_case_table
</table>
</test_cases_section>
<test_code_section>
<header>Test Code:</header>
$test_code
</test_code_section>
<test_review_section>
<header>Test Review:</header>
<review>$test_review</review>
</test_review_section>
<coverage_summary_section>
<header>Test Coverage Summary:</header>
<summary>$coverage_summary</summary>
<insights>$insights</insights>
</coverage_summary_section>
</response_format>
Here is the code that you must generate test cases for:
<code>
PASTE_YOUR_CODE_HERE
</code>
---
ChatGPT system prompt is 1700 tokens?!?!?
If you were wondering why ChatGPT is so bad versus 6 months ago, its because of the system prompt.
Look at how garbage this is.
Laziness is literally part of the prompt.
Formatted in the paste bin below.
https://t.co/XSA85dys1I
At a recent event, I got asked for my best pieces of career advice.
Here are the 10 pieces of advice I shared:
(things I wish I knew at 22)
1. Build a reputation for reliability.
You can get pretty damn far by just being someone that people can count on to show up and do the work. Being reliable is entirely free.
2. Be the person who can figure it out.
Early on, you'll be given a lot of tasks you have no idea how to complete. There's nothing more valuable than someone who can just figure it out. Do some work, ask key questions, get it done. People will fight over you.
3. Work hard first (and smart later).
It's trendy to say that working smart is all that matters. Wrong. If you want to accomplish anything significant, you have to work hard. Work hard early—take pride in it. Then you can start to build leverage to work smart.
4. Build storytelling skills.
World-changing CEOs aren't the smartest in their orgs. They are exceptional at: (1) Aggregating data and (2) Communicating it simply & effectively. Data in, story out. Build that skill and you'll always be valuable.
5. "Swallow the frog" for your boss.
This is one of the greatest "hacks" to get ahead early in your career. Observe your boss, figure out what they hate doing, learn to do it, and take it off their plate. Easy win.
6. Be a "yes" person early in your career.
Saying "yes" expands your luck surface area. It may mean you're a bit overwhelmed at times, but the benefits from the increased luck outweigh the downsides of feeling stretched.
7. Wake up early and work out.
When you wake up early and work out, you do a hard thing to start your day that sets the tone. You start to self-identify as a winner. That has ripple effects all across your life. There's no such thing as a loser who wakes up at 5am and works out.
8. Dive through cracked doors.
I recently had an experience to bring this to life: A young guy saw on my story that I was at a coffee shop working. He messaged me asking if he could come by and ask a question. I said ok. He got there an hour later and we hit it off. Turns out he lived far away and made it work. I'd always bet on people with that kind of energy. If someone cracks open a door that may present an opportunity, dive through it.
9. Show up early, stay late.
Showing up early and staying late is a free way to materially increase your luck surface area. The most interesting side conversations come up before meetings start or after they end. When you're in the room, you're more likely to get pulled into a follow-up call, coffee, or discussion. It pays off handsomely in the long run.
10. Do the "old fashioned" things well.
Look people in the eye, do what you say you'll do, be early, practice good posture, have a confident handshake. It sounds silly, but these things are all free and will never go out of style.
***
Embrace those 10 pieces of advice and you'll stand out and be on the right track.
If you enjoyed this, share it with others and follow me @SahilBloom for more in the future!
among all the turmoil happening lately, @karpathy still manages to release videos to educate the world ! Released just a few hours ago.
https://t.co/68YfVMUeM3
“If you read any management textbook, it says ‘don’t hire the divas’ because they’re nothing but a pain in the ass. And by the way, they are. But the people who are the divas—who believe—are the ones who will drive the culture and company to excellence.
Steve Jobs was a diva. I worked with Bill Joy who was my colleague for many years—he’s an example of a diva. And I mean this in the most flattering way: they expect a lot, they drive people hard, they’re controversial, and they care passionately. If you find those people, you’re probably going to work for one so be nice to them.”
-- Eric Schmidt
GPT-4’s secret might have been revealed—it’s not one model, but many. Rumors say that it is a 1T monolithic architecture but an 8×220B mixture of experts (MoE) model. So, 8 smaller models tied together.
Here's the whole story:
https://t.co/HLKiRt2Lyh
🦙🐪🐫 So many instruction tuning datasets came out recently! How valuable are they, and how far are open models really from proprietary ones like ChatGPT?
🧐We did a systematic exploration, and built Tülu---a suite of LLaMa-tuned models up to 65B!
📜https://t.co/cFE2JUD6Zc
Getting LLMs To Tell The Truth
-Shifts LLM activations during inference, following directions across attention heads
-Improves LLaMA 33% -> 65% on TruthfulQA
-LLMs may have internal representation of something being true, even as they produce falsehoods
https://t.co/sh5Ck5luUV
You really, really should not be relying on AI detectors for classroom use.
This new paper shows that not only are they very easy to defeat by just prompting a couple of times, but they have insane false-positive rates against non-native English speakers. https://t.co/SxRWCkagHV