Thêm vài nền tảng cho anh em kiếm tiền.
Các nền tảng này đều cần tiếng Anh ở mức cơ bản. Nhưng có vẻ không dành cho thị trường Việt Nam. Anh em quốc tế thì chắc cơ hội nhiều hơn rồi. Nhưng thôi cứ đăng kí đi, cơ hội thì luôn nằm ở đó.
Những nền tảng này mình đã tìm hiểu kĩ, có bằng chứng trả lương hậu hĩnh, không phải chỉ lên hỏi ChatGPT là xong, vì ChatGPT nó liệt kê rất lan man, đăng kí chỉ mất thời gian.
1. Prolific
https://t.co/2090dBE8tW
Nền tảng này trả khá hậu hĩnh, nhưng nhiều quốc gia không nằm trong chính sách của họ, trong đó có Việt Nam. Đăng ký thì có thể được đưa vào waitlist, nên cứ tạo tài khoản trước. Hình như trước đây cũng từng có thời điểm một số anh em Việt Nam vào được.
2. uTest
https://t.co/W4UKm2QPQg
Nền tảng này chủ yếu để test app, web, game, thiết bị... Mình đang làm trên này, có gì sau này sẽ cập nhật thêm.
Ban đầu nhìn hơi rối và cũng cần học cách report bug cho đúng. Trong nền tảng có Academy, anh em cứ học dần hết. Job được phân phối theo profile, thiết bị, rating và lịch sử làm việc nên không phải lúc nào cũng có.
3. UserTesting
https://t.co/1I1EOkIvrY
Mình cũng vừa làm bài onboarding gần đây.
Công việc chủ yếu là vào website/app, sử dụng thử rồi vừa thao tác vừa nói suy nghĩ của mình. Vì vậy cần speaking tiếng Anh tương đối ổn, nhưng không nhất thiết phải nói như người bản xứ.
Theo mình cái khó nhất là nói liên tục và giải thích được tại sao mình thích/không thích một thứ gì đó
Tóm lại mình thấy khả thi nhất đối với cá nhân mình là
uTest > UserTesting > Prolific
Anh em nghiên cứu thử nhé. Có nền tảng nào mới mình cập nhật tiếp. Hôm nay sẽ ngồi học academy bên uTest
How many Chinese AI researchers have you learned the names of recently?
Liang Wenfeng, Zhilin Yang, Jie Tang, Fuli Luo. I keep finding people whose work I should have paid attention to much earlier.
Zhilin Yang, the founder of Kimi, coauthored Transformer-XL and XLNet in 2019. He completed his PhD at Carnegie Mellon that same year, advised by Ruslan Salakhutdinov and William Cohen.
His undergraduate adviser was Jie Tang (@jietang), the cofounder of Zai, whose work on AMiner goes back to 2006. Both names are on the original GLM paper. I knew Kimi and GLM before I knew about that connection.
Fuli Luo (@_LuoFuli) coauthored VECO at Alibaba, worked on DeepSeek, and is now building MiMo at Xiaomi. Her recent post about scaling RL has over 3M views. People clearly want to hear about the work itself, and I'd love to see more of that.
Liang came through quantitative trading and High-Flyer before DeepSeek. MiniMax founder Junjie Yan has an AI PhD and spent over 6 years at SenseTime. Daxin Jiang (@DaxinJiang) worked on search at Microsoft before founding StepFun.
I'd happily read a long interview with any of them. How they chose their research, what failed, who taught them, how they found the people they wanted to work with.
When I think about US AI coverage, Sam, Elon and Dario still come to mind first. I know about Ilya, Karpathy and Schulman too, and Dario himself came through research. I want to keep discovering people behind the models on both sides.
With open models, I can try the work myself and read about how it was done. That makes me want to understand the years of research behind a release, and the people who kept working on it.
I want these teams to succeed. Getting to use their work and hear directly from some of the people who built it is a big part of why I find AI so exciting.
Jev Founder, Diogo Amogo, just released a PDF on building a Jev Harness for coding agents
this is a blueprint on how to make your coding agents 200× faster and 400× cheaper
Send this PDF and the article below to your Claude
Code or Codex instance and start shipping 200× faster 👇
LLMs vs. Jev, clearly explained!
LLMs are great, and the ceiling is one you can watch scroll past:
an LLM writes the answer one token at a time.
give it a failed deploy and four decisions, and it produces a small JSON object where every token depends on the one before it.
token nine cannot exist until token eight does, so four decisions that had nothing to do with each other just stood in a queue.
then your code parses it, validates the shape, and retries when the shape is wrong.
Jev fixes this without being a smaller or faster model: it removes the order.
one turn on that deploy has to know:
→ whether the incident is urgent
→ which team owns it
→ whether the next command is risky
→ whether the task is actually done
you declare the questions and the answer type upfront, and all four come back together, typed, with a probability on each.
three primitives cover almost every fork in an agent:
1. **Choice** picks one of up to 255 options you define, like engineering, billing or sales.
2. **Score** places the state on an ordered scale you define, like low, medium or high risk.
3. **Noul** returns the probability that a yes-or-no condition is true.
here is the sentence that resolves the whole confusion:
text is a line you have to walk. an answer space is a room you see all of at once.
↳ generation: one order you cannot change, one string at the end, a shape you hope holds
↳ evaluation: no order at all, typed answers, a probability on every option
Prompts → Agents → Loops → Graphs → Jev
the probabilities matter more than the answer.
↳ engineering at 0.91 against billing at 0.09 is a route you can automate
↳ 0.52 against 0.46 is a coin flip wearing a label, and the label alone never told you which one you got
that last one catches careful people. an LLM would have said "engineering" in a confident sentence and given you no way to know the race was that close.
thresholds live in your code, one per action, scaled to what being wrong costs.
it works when the options are known and the call depends on meaning. it is not for writing, code, arithmetic, or anything where question two needs the answer to question one.
and the one that eats whole nights: type safety prevents malformed output, not incorrect judgment.
Jev cannot return an option outside your schema, and it can still pick the wrong valid one with confidence. a schema-valid mistake refunds the wrong customer just as fast.
an LLM writes new language when the answer space is open. Jev evaluates known paths when the answer space is closed.
below i have quoted my full breakdown on Jev. it covers the three primitives, the parallel battery, the thresholds, and where it does not belong.
save this and read it below ↓
Layers of observability in AI systems, explained visually:
If an LLM app is serving real users, its input and output are not enough to debug it.
Consider a RAG pipeline where a query passes through embedding, retrieval, context assembly, and generation.
Every operation adds latency, may call a paid API, and can fail while still producing a valid-looking response.
Traces and spans provide visibility.
- A trace records the full path of one request. The Trace column runs from query to response.
- A span records one operation within that trace. The colored boxes are spans.
Each span captures:
> Query span
The input, timestamp, session identifier, and request metadata.
> Embedding span
The model, input size, latency, retries, and rate-limit errors.
> Retrieval span
The retrieved chunks, document IDs, relevance scores, filters, top-k value, and latency. Many RAG failures originate here. Without these fields, there is no evidence that retrieval selected the wrong documents.
> Context span
The context assembled from retrieved chunks, instructions, and conversation history. This catches truncated documents, duplicated chunks, missing citations, and prompts exceeding the token budget.
> Generation span
The model, token counts, time to first token, total latency, finish reason, retries, and estimated cost.
With these details, a bad response can now be traced to retrieval, context assembly, or generation.
To use this in practice, Opik already implements this observability infrastructure for LLM apps and is open source.
It captures traces and spans across LLM calls, retrieval steps, and tool executions, with latency, token usage, and cost attached to each operation.
GitHub repo: https://t.co/vahjkkfJCt
(don't forget to star it ⭐)
In Opik, every operation belonging to one request carries the same Trace ID. If the app processes 1,000 requests, it creates 1,000 traces, each containing its own spans.
This makes cost analysis more useful. Instead of aggregate spend, teams can identify the model calls, retries, or oversized prompts responsible.
Over time, changes in retrieval scores, embedding latency, or context size become visible before they turn into broader quality problems.
That said, observability is one of eight areas I would learn for building production LLM systems.
I covered all eight in the 2026 LLM Engineering Roadmap, with free and open-source resources for each one.
Read it below.
Introducing Bolt Forge. Free until Oct 14th:
- Up to 50x more usage
- The new frontier: GLM, DeepSeek, Kimi
- Zero usage charges
Live now in your model picker on https://t.co/UH6gFfHvbp
And one more thing... 👇
Switzerland is spending more than $11 million to move away from Microsoft 365 and switch to open-source software.
The Swiss government plans to spend 9 million Swiss francs to replace Microsoft apps such as Outlook and Teams with the open-source openDesk suite.
The first step will cover 3,000 government computers, with the switch expected to be completed by the end of 2027.
Switzerland first tested openDesk with 172 employees and it worked well for everyday tasks.
The plan could grow much larger. Switzerland has about 54,000 government computers and may eventually switch all of them.
Switzerland want give the government more control over its software and data.
Kiro students just went global 🌎
Expanding to 121 new universities across 16 countries.
Same deal for everyone: a full year free, 1,000 credits/month, premium models + Kiro Web included.
Sign up and learn more here 👉 https://t.co/nO5LCZxVE7
The most important skills for using AI coding agents effectively. Presenting the AI Engineering Skills Map for using coding agents. https://t.co/GrEw7wG5Wz
Introducing Docs7 - the best way to serve documentation to agents.
◆ Automatically serve clean, agent-readable docs
◆ Native WebMCP and skills support
◆ Deep insights into AI traffic & searches
We make agents love your product 👇
Trending repository of the day 📈
archify
Agent skill for beautiful, verifiable architecture, workflow, sequence, data-flow, and lifecycle diagrams—self-contained HTML with motion and crisp export.
Last 24h: 3,902 ⭐
Total: 32,898 ⭐️
https://t.co/ggvj9dnkoQ
Everyone has done it: You build a new application with coding agents and love the toy version so much that you wind up deploying it without first evaluating trade-offs like latency, uptime, and compute costs. And we all know vibe coders who do this every day and only discover the production consequences later.
When a developer lacks the proper grounding in software engineering fundamentals, they can’t steer an agent to make the structural decisions needed for building great applications.
In this week's letter, Andrew Ng outlines why understanding the full software development stack, from data management to user interfaces, remains essential for AI engineering.
Here’s Andrew’s list of the key technical capabilities requiring a knowledgeable human’s guidance:
🏗️ Full-Stack Development: Planning out API design, session management, caching strategies, and asynchronous processing.
🗄️ Data Lifecycle Management: Picking data models and storage infrastructure to maintain consistency and clean feeds for downstream AI systems.
📐 System Architecture Design: Choosing everything from the monolith versus microservice dilemma to load balancing.
🔒 Security and Reliability: Building in smart failure-handling policies to ensure graceful degradation, unit and integration testing, and security and testing — early and often.
🚀 Production Operations: Configuring CI/CD pipelines, managing databases, and creating observability tools for understanding workloads.
Read Andrew’s full letter to explore how software fundamentals shape effective AI engineering, and how those fundamentals fit into our overall map of AI Engineering Skills: https://t.co/QE4DsmAJzy
How have software engineering fundamentals changed with agentic coding? Here is our AI Engineering Skills map for software engineering fundamentals. https://t.co/cnRLj43DLs