Behind every AI model are people: annotators, reviewers and evaluators, many in the Philippines. Economists at a PIDS-ILS roundtable say PH needs better data on this workforce. We agree: you can't upskill or protect what you can't see. #AIPH https://t.co/MqOyhVOk3Q
@TechCrunch Ads inside ChatGPT turn AI assistants into a real media channel. Clear labels and keeping ads separate from generated results will be key to user trust. For brands, showing up organically in AI answers just got more important. Would you click one?
@CNBC Public hearings like this matter. Much of AI safety is decided upstream: how training data is sourced and governed, and how much real human oversight goes into evaluation. Guardrails work best when built in, not bolted on. #ResponsibleAI
@thebponews A wake-up call, not a eulogy. Work that AI depends on, like data annotation, model evaluation and human QA, keeps growing, and Filipino talent is well placed for it. The real question is how fast reskilling can scale. Which roles do you see growing first?
@pewresearch Really useful study. Double-digit misses show synthetic respondents can supplement real people, not replace them. As an AI data company, we see the same in training data: the human signal keeps models grounded. Will you re-run this with newer models?
Can AI take a survey instead of humans? They can, but maybe they shouldn’t.
As #AI advances, there is growing interest in using it in place of real respondents in public opinion surveys.
We wanted to learn more, so we ran our own synthetic polling experiment.
@iScienceLuvr The 1.6x compute cost is the striking part. Synthetic text clearly has a ceiling, and verified human data is becoming the premium input for training. Will labs start treating curated human data as a strategic reserve? We think they should.
How Much Is an AI Token Worth? Scaling Laws for Wild AI-Generated Web Text
"After applying FineWeb quality filtering, we find that 27.5% of tokens from June 2026 web data are labeled as AI-generated by Pangram, rising to 31.1% by August."
"How does AI text in the wild affect language model pretraining? To answer this question, we pretrain 800 language models, varying the ratio of added AI tokens to human tokens, and fit scaling laws to held-out losses on both human and AI-generated text. For data-starved models, adding AI tokens to pretraining data initially lowers loss on human text, but the benefit saturates as more are added and quickly *reverses* into harm. For models trained on high budgets of human text, AI tokens raise loss almost immediately, whil"e the same number of fresh human tokens keeps lowering it."
"We propose a new scaling law with separate benefit and harm terms that allows the value of an AI token to change sign while also reducing to Chinchilla in the absence of AI text. When fit on smaller models, our scaling law predicts the effect of AI text on held-out human-text loss for models up to 3.6x larger with 41% lower error than the best existing law over all AI ratios."
link: https://t.co/HqVuEu3DEw
@Picturesque_AI_ Yes, that wasted spend is real, and it often hides because early test clips look decent. Writing down what can't change before anyone opens a tool usually saves the rework later.
"Which AI video tool should we use?" is the wrong first question.
Does the product have to look exact? Does a real person have to carry it? What footage already exists?
Those three answers pick the method. Then the model question makes sense.
#AIGC#Lifewood
@Hsenju True, trust falls apart when nobody can prove which file got the yes. An approval that doesn't name the exact asset is basically a sticky note. The painful part usually shows up weeks later, when someone asks for the trail.
Two versions of the same AI generated ad. One approval message.
Nothing in either records which version the approval was for.
Open your last released asset and try to trace it without asking anyone.
#AIGC#Lifewood
@swarleylivies That shift matches what we see too. A prompt people reuse every day needs an owner and a version like any other ops asset. Starting with the one your team uses most keeps the habit small.
A prompt three people use every day isn't a shortcut. It's an operations asset.
So give it what assets get: an owner, a purpose, approved inputs, a version, evaluation examples, and a retirement trigger.
Start with the one your team uses most. #AIGC#Lifewood
Pick one record from your LLM training set and name what it teaches, in one sentence.
If you can't, it's adding weight and not coverage.
Assign a learning goal to three sample records before you collect more.
#LLMTrainingData#Lifewood
A parked van hides most of the pedestrian.
One annotator boxes the visible shoulder. One boxes the estimated full body. One skips it.
Nobody is wrong, because nobody was told.
Document the occlusion rule before labelling.
#DataAnnotation#Lifewood
The scanner reported success. The file opens. Page 47 is not in it.
A double feed leaves no trace except a folio that jumps from 46 to 48.
Reconcile the page order before you index, not after.
#DocumentDigitization#Lifewood
"Fifty thousand documents."
One side is counting files. The other is counting pages.
Nobody finds out until acceptance.
Define the counting unit before you quote it.
#AIData#Lifewood