If you are a software engineer "experiencing some degree of mental health crisis", now hear this, because I've been coding for 50 years since the days of punched cards and I have a salutary kick in your ass to deliver.
Get over yourself. Every previous "programming is obsolete" panic has been a bust, and this one's going to be too.
The fundamental problem of mismatch between the intentions in human minds and the specifications that a computer can interpret hasn't gone away just because now you can do a lot of your programming in natural language to an LLM.
Systems are still complicated. This shit is still difficult. The need for people who specialize in bridging that gap isn't going to go away.
As usual, the answer is: upskill yourself and adapt. If a crusty old fart like me can do it, you can too.
Padding in our non-AR sequence models? Yuck. 🙅
👉 Instead of unmasking, our new work *Edit Flows* perform iterative refinements via position-relative inserts and deletes, operations naturally suited for variable-length sequence generation. Easily better than using mask tokens.
Introducing All-atom Diffusion Transformers
— towards Foundation Models for generative chemistry, from my internship with the FAIR Chemistry team @OpenCatalyst@AIatMeta
There are a couple ML ideas which I think are new and exciting in here 👇
V happy to see this out - reprocessing 5million SARS-CoV-2 genomes to remove systematic errors (+thereby improve the phylogeny) and improve representation of genomes from the Global South. Huge amount of work esp. from Martin Hunt, Angie Hinrich+collabs
https://t.co/lKiTabLU6p
wow. The level of fraud from Didier Raoult is far larger than I imagined. 50+ papers retracted or with editorial notices. And in 2023, "the largest unethical study performed for years—in France, maybe in the world. … It’s incredible" https://t.co/qIYx1QlXxX
C'est l'histoire, folle, du Parquet National Financer qui perquisitionne une éminente scientifique française (@DgCostagliola) en se basant sur... Du vent. Ou plutôt une plainte d'une association regroupant parmi les pires figures de la complosphère. https://t.co/bmluOeuFig
@giulio_pibiri@daniel_c0deb0t Intel tried something in that vein: https://t.co/lYT3na4DtC
A SSD faster than NAND flash and denser than DRAM. Also less sensible to wear, so it make sense to use them as write caches even today.
Our latest preprint is out!
TLDR:
We propose a novel model, the Tinted de Bruijn graph that associate to each kmer all the reads that contains it
We developed the K2R index that scale nicely on high coverage human datasets
(1/n)
https://t.co/35F22nWAHV
@nomad421@RayanChikhi Pugz didn't get a flexible enough interface for non-sequential reads. Supporting a variety of consumers and access patterns seemed fairly difficult in C++, but rapidgzip did it!
I wonder how much rayon and crossbeam could ease the implementation of rapidgzip architecture 😉
@nomad421@RayanChikhi With pugz, we got started by running libdeflate on every bit-string suffixes. We optimized and specialized the code, but pugz' block finder retains this behavior. Rapidgzip improves a lot on that with stricter filtering and caching results to LUTs.
Working more on minimizers now with @giulio_pibiri. So much fun!
We can now do better than miniception for both small k (green) and large k (purple). This also breaks Schleimer's lower bound for random schemes.
1/3
I ended up writing a Linux scheduler in Rust using sched-ext during Christmas break, just for fun. I'm pretty shocked to see that it doesn't just work, but it can even outperform the default Linux scheduler (EEVDF) with certain workloads (i.e., gaming): https://t.co/3DyVvUHUQc
@CamilleMrcht Grandpa: sounds like a recommendation system.
However, Jean Noli discarded compressed sensing and sparse matrix factorisation, writing «Des extrêmes se rencontrent malgré les lois formelles de la géométrie.»
Mikaël Salson @m_salson obtained his HDR yesterday, with a defense entitled Méthodes sans alignement et indexation
pour l'analyse de données nucléiques massive (alignment free methods for massive nucleic sequences analysis).
@awmayhall@CNC_Kitchen In principle, all should work, so I'm curious as to why 4A is optimal. I can think of surface area, Keq and competition with other gases, but water should adsorb the most strongly.
I have ethanol to dry so opted for 3A in the end.