@forgebitz@levelsio Vector databases aren't useful because they store vectors. They're useful for implementing HNSW indices for fast fuzzy lookup of vectors.
I made a small node setup that is similar to 9-slice sprites but for 3D meshes. It prevents the 'corners' of the object from distorting when scaled.
Also, you can assign vertex groups to prevent some parts of the object from scaling.
Get here for free:
https://t.co/k069fmEVkl
@minchoi Because R1 hardware is something everyone is carrying in their pocket already and thus could be an app. All your other examples the hardware offers a unique i/o device or sensor to justify the existence of a new physical device.
Llama 3 crushed my financial metrics tests
I tested both the 70B and 8B models.
Both aced the metric calculation tasks.
The result from today’s tests indicate an emergence of 4 distinct LLM tiers:
• thropughput tier
• workhorse tier
• intelligence tier
• groq tier
Groq tier: These are models served by @GroqInc. Extremely fast, excellent pricing, and open source. Groq lives in its own strata.
Throughput tier: Pricing is very competitive. Great for quick, non-critical tasks.
Workhorse tier: Mid-tier pricing, but stronger at complexity than throughput tier. Great for majority of tasks.
Intelligence tier: Premium complexity at higher price and lower speed. Largest models, best in the industry. Great for most critical tasks.
For my tests, I created a financial metrics calculation task.
Given financials, calculate:
• net profit margin
• debt-to-assets
• free cash flow
The tasks require a mixture of text extraction, math, instruction following, and reasoning.
Given that newer models are passing the tests, I need to make the tests tougher.
To date, DCF calculation remains the toughest task, with only opus having passed it.
A lot of the insider knowledge on how to build an LLM has gone underground in the last 24 months.
We are going to build #SnowflakeArctic in the open
Model arch ablations, training and inference system performance, dataset and data composition ablations, post-training fun, big run stability, decontamination tricks, metric subtleties.
To start with the first two blogs in our cookbook, keep reading the 🧵below ...
@Duderichy This is just the vesting though, the annual new stock grant is more interesting to look re. long term compensation. If the new stock grant is unaffected and has this new vesting schedule, this is strictly a win for employees.
YOLOv9
Learning What You Want to Learn Using Programmable Gradient Information
Today's deep learning methods focus on how to design the most appropriate objective functions so that the prediction results of the model can be closest to the ground truth. Meanwhile, an appropriate architecture that can facilitate acquisition of enough information for prediction has to be designed. Existing methods ignore a fact that when input data undergoes layer-by-layer feature extraction and spatial transformation, large amount of information will be lost. This paper will delve into the important issues of data loss when data is transmitted through deep networks, namely information bottleneck and reversible functions. We proposed the concept of programmable gradient information (PGI) to cope with the various changes required by deep networks to achieve multiple objectives. PGI can provide complete input information for the target task to calculate objective function, so that reliable gradient information can be obtained to update network weights. In addition, a new lightweight network architecture -- Generalized Efficient Layer Aggregation Network (GELAN), based on gradient path planning is designed. GELAN's architecture confirms that PGI has gained superior results on lightweight models. We verified the proposed GELAN and PGI on MS COCO dataset based object detection. The results show that GELAN only uses conventional convolution operators to achieve better parameter utilization than the state-of-the-art methods developed based on depth-wise convolution. PGI can be used for variety of models from lightweight to large. It can be used to obtain complete information, so that train-from-scratch models can achieve better results than state-of-the-art models pre-trained using large datasets, the comparison results are shown in Figure 1.
Please wishlist my new game KingMakers on Steam. We've been working really hard on this for 5 years, and can finally unveil it to the public today. Hope you guys like it 🤞
Stability releases Stable Cascade
demo: https://t.co/eychPLlXNS
github: https://t.co/lvzwadisIY
a new text to image model building upon the Würstchen architecture
"a giant cathedral is completely filled with cats. there are cats everywhere you look. a man enters the cathedral and bows before the giant cat king sitting on a throne."
Video generated by Sora.
I built a natural language CLI.
It generates Python scripts to answer your question, then auto-executes them in the cwd. You will not believe how capable this simple pattern is. Rawdogging gpt-4 from the command line.
Rawdog.
1/
@amasad This is my experience as well. In YouTube's case it seems to be because they value new content far too much in their recommendations. I haven't seen the old content because I was busy, not because I don't care about it.
🚨 New Research Alert! After a long journey, our research with @Iciralabama is out! 📚
Key findings:
✅Repeated exposure to myths within corrections increased perceived familiarity.✅This effect heightened misinformation credibility, even among those with low prior beliefs.