Eigentlich ist es in Deutschland verboten, Dokumente aus laufenden Strafverfahren zu veröffentlichen. Doch es gibt Dokumente, die gehören an die Öffentlichkeit.
Hier sind die Gerichtsbeschlüsse zu Durchsuchungen & Abhörmaßnahmen bei der @AufstandLastGen https://t.co/Q2kb2BVDxk
I cannot POSSIBLY be any clearer (as someone who worked in IT managing some of this country's biggest databases)
UNDER NO CIRCUMSTANCES SHOULD ANYONE PROVIDE PICTURES OF THEIR PASSPORTS OR DRIVING LICENCES
ESPECIALLY NOT
to databases outside the GDPR
THREAD:
Even OpenAI CEO Sam Altman was skeptical a few weeks ago: "I probably trust the answers that come out of ChatGPT the least of anybody on Earth." https://t.co/k0kx2JF77R
Interesting post -> Generative AI could be a dud, so we shouldn't build around the premise that the tech is world-changing
"If hallucinations aren’t fixable, generative AI probably isn’t going to make a trillion dollars a year." https://t.co/cs1UvcrxFr
I'm really not that interested in this story, but as the facts are revealed, I can't help reminding us all about the frantic speculation last week, "Electric cars start fire!!! AAaaarrrrgh!!!
It was NOT the electric cars.
https://t.co/KqGMSOicG8
Das Erstarken der AfD weckt bei Gerhart Baum Erinnerungen an Traumata der Nachkriegszeit. Die Gefahr werde auch vom Kanzler verharmlost, schreibt der ehemalige Innenminister in seinem Gastbeitrag. https://t.co/UkTrIavtoa
Three years ago, new international rules took effect limiting sulfur in the heavy fuels used by ships.
Practically overnight, maritime sulfur pollution dropped 85%. This is good for humans, as sulfur pollution is toxic, but probably had unintended climate consequences.
A 🧵
Machine Learning for SEO
Are you ready for this?🤯
I just generated top ten internal link recommendations in a dataset of 100,000 pages that don't link to each other but should. Rather than using traditional means of matching pages such as TF-IDF, LDA, n-grams and keyword frequency I use machine learning.
Embeddings
Scraped, pre-processed and tokenised data is embedded to a 768-dimensional vector space using Google's language-agnostic BERT sentence embedding model.
Page > Page
Combined and averaged sentence vectors for each URL are used to form a page similarity matrix in numpy which is then used to build a list of semantically related pages through cosine similarity.
That was the easy part.
Now onto the tricky business of anchor text recommendation by selecting a portion of existing text on the proposed linking page that is also relevant to the target page.
Sentence > Page
Iterating over all sentence vectors nested within each identified link candidate page, once again, looking for the one with the shortest cosine distance in respect to the page-level embedding of the target URL as a whole. Once the most suitable sentence embedding has been identified, it's then possible to, on the fly, reconstitute that sentence by mapping the sentence ID and URL which are stored as raw text key column at the end of the vector values.
Word > Page
Regenerated raw text content is used to produce word IDs + word-level embeddings which are cycled through one final similarity lookup. Again, looking for the shortest cosine distance between the selected word vectors of the source page and page-level vectors of the target page. It's now possible to re-build the keyword text (otherwise lost in vector data) by mapping word ID keys kept in the program memory during this temporary cycle same as we did earlier with sentences.
Word = Anchor Text
All identified keywords are then recorded in the "anchor_text" column of the link recommendation csv (source,target,anchor_text) in descending order of relevance.
Note: The reason word-level vector embeddings are generated on the fly and not during sentence-level embeddings is due to it being computationally demanding and requiring hundreds of gigabytes of storage for a dataframe of 100,000 scraped pages.
Tech Stack
For word embeddings I use SentenceTransformers [1], LaBSE [2], PyTorch [3] and Nvidia's CUDA [4] on a Tesla T4 16GB GPU. LaBSE is a good choice for websites with many herflang silos due to it being completely language agnostic. You can of course use vanilla BERT if you like, but I would recommend BERT-large-cased.
Pro Tips:
1. My tests show significantly better outcomes when using L2 normalised [5] LaBSE embeddings.
2. Don't strip full stops during tokenisation process otherwise you'll lose sentence boundary data.
Useful Links:
[1] https://t.co/fvqNza50VV
[2] https://t.co/TSbXCLhMai
[3] https://t.co/Aiv07gk4B4
[4] https://t.co/0u75qnk4pC
[5] https://t.co/i4DwpvESkS
Es wäre sicher nicht uninteressant die ganz besondere Liebe dieser Bundesregierung für #Wasserstoff inklusive der entsprechenden personellen Verbindungen auch im Gebäudesektor anzuschauen 🧐 https://t.co/WWiABHSvci
Ich möchte diesen Tweet von @HubertAiwanger zum Anlass nehmen und einen 🧵 über CO2e-Rucksäcke von Lebensmitteln zu schreiben.
Viele überschätzen die Bedeutung des Transports. Darum ist es wichtig, sich die Bilanzen anzuschauen und nicht auf sein Bauchgefühl zu vertrauen. ➡️
🔥Fahrzeuge mit Verbrennungsmotor und Hybridantrieb brennen deutlich häufiger als reine E-Autos.
So haben pro 100.000 verkauften Autos 3475 mit Hybridantrieb gebrannt, 1.530 mit Verbrennungsmotor und nur 25 #eautos. https://t.co/kvV8PsS9uE via @riffreporter#FremantleHighway
@StefanRoock@thepumpapp ah Mist (ja bin auf iOS unterwegs). Bei ABRP gibts eine Android Version, aber damit hab ich bisher nur Offline Planung gemacht.
Wir waren gerade mit der Familie am Gardasee auf einem Campingplatz. Alle 500 Autos sind komplett zerstört, die Mobilhomes sind kaputt, die Scheiben zum Teil eingeschlagen, Menschen zum Glück nicht schwer verletzt. Es braucht ein Krisenwarnsystem u. ein Umdenken. @Eurocamp_Urlaub