음성 에이전트한테 가짜 손님 전화를 잔뜩 걸어서, 어디서 깨지는지 배포 전에 찾아 줌. LiveKit이 Agent Simulations를 10월 8일 정식 출시했음.
핵심:
① LLM이 연기하는 가상 사용자가 시나리오를 대화로 끝까지 풀고, LLM 심사관이 대화록을 보고 통과·실패와 이유를 매김
② 시나리오는 에이전트 코드를 읽어 자동으로 만들어 주고(lk agent simulate), 실제 운영에서 실패한 대화를 테스트로 바꿔 넣을 수도 있음
③ 텍스트 모드는 싸서 PR마다 돌리고, 오디오 모드는 배경 소음·끼어들기·맞장구까지 넣어 지연과 전사 정확도를 잼. 실패하면 CLI가 빌드를 깨뜨림
Serve AI가 에이전트를 통째로 다시 짠 뒤 실제 통화 100건으로 먼저 검증했다고 소개함. 10월 한 달은 무료임.
한 가지 주의할 점은 무료여도 시뮬레이션이 실제 에이전트를 돌린다는 것. LLM 호출, 오디오 모드면 STT·TTS 비용까지 평소 요금대로 나감.
데모 영상은 약 11분임.
공식 영상↓
Know exactly what your agent can do, catch gaps before customers do, and test any model against your own scenarios before you switch. Your agent has a QA team now.
Try LiveKit Simulations free through October: https://t.co/dx8DzqXb0s
27B 모델을 8GB 밑으로 줄였는데,
도구 호출은 원본보다 더 잘 맞춤. Underdog이 Saluki 27B를 공개함.
Underdog Bench 120문항 중 88개 통과
같은 하네스로 돌린 원본 Qwen3.8-27B는 84개임. 파일 크기는 7.89GB, 원본 BF16은 54GB.
핵심:
① Qwen3.8-27B를 2비트 혼합 정밀도 GGUF로 줄인 모델임. ISTA-DASLab의 GSQ-RCO 양자화 파일 위에 Underdog이 도구 호출을 지키는 쪽으로 한 번 더 손봄
② 병렬 도구 호출 100문항에선 42 대 35로 원본보다 높게 나옴. 9개 벤치 평균 유지율은 96%라고 함
③ 별도 포크 없이 기본 llama.cpp에서 돌아감. 이미지가 필요하면 629MB나 928MB짜리 비전 애드온을 붙이면 됨
④ 라이선스는 Apache 2.0
Hugging Face에 모델 카드와 파일이 올라와 있고, 로컬 에이전트용으로 쓰라고 나옴.
한 가지 주의할 점은 수학·추론이 꽤 깎인다는 것. AIME 2025는 79.2로 원본 공개 점수 96.7보다 낮고, 병렬 호출 응답 다섯 개 중 하나꼴로 작은 형식 실수가 난다고 모델 카드에 적혀 있음.
원문↓
https://t.co/XgZ7PeehUt
600B짜리 코딩 모델을 일주일 동안 공짜로 돌려볼 수 있음.
쓰던 코딩 에이전트에서 모델만 Step 5 Preview로 바꾸면 됨.
DeepSWE v1.1 67.7%
StepFun이 직접 공개한 점수임. Kimi K3(67.5%), GLM-5.3(66.9%)보다 조금 높고 GPT-6 Astra(74.1%), Claude Opus 5(74.0%)와는 아직 차이가 있음.
핵심:
① 중국 StepFun의 주력 모델. 전체 600B 중 토큰마다 27B만 쓰는 MoE이고 컨텍스트는 100만 토큰, 이미지와 영상도 입력으로 받음
② 처음 공개된 건 9월 20일임. 이번에 바뀐 건 10월 8일부터 OpenRouter에 올라왔고 OpenCode, Cline, Kilo Code, Nous Portal(Hermes Agent)에서 일주일 무료가 풀렸다는 것
③ OpenRouter 가격은 100만 토큰당 입력 $1, 출력 $2.70
④ StepFun은 10월 15일에 가중치를 공개하겠다고 밝힘
무료 기간은 툴마다 일주일이고, 끝나면 OpenRouter나 StepFun API에서 유료로 쓰면 됨.
한 가지 주의할 점은 에이전트 작업에선 약점도 보인다는 것. Artificial Analysis는 지능 지수 44점으로 Kimi K3(max)와 같다고 하면서도 에이전트 평가는 경쟁 모델보다 뒤처진다고 적었고, StepFun 표에서도 Terminal-Bench v4는 33.3%로 GLM-5.3(41.9%)보다 낮음.
원문↓
https://t.co/Iro76a7t4s
글만 주면 코드로 애니메이션을 짜고, 그걸 영상으로 뽑아 줌. MiniMax Design이 Opus 5.5로 돌리는 “Code Your Next Video” 데모를 공식 계정에 올렸음.
핵심:
① JavaScript 애니메이션, 모션 그래픽, 설명 영상, 웹·제품 데모까지 코드→모션으로 이어지는 흐름을 보여 줌
② 언어 모델이 비주얼을 해석하고 결과물을 계속 다듬는다고 소개함
③ 제품 사이트는 텍스트·이미지·영상을 한 캔버스에서 에이전트가 나눠 맡는 멀티모달 제작 툴로 설명함
데모 길이는 약 68초임.
한 가지 주의할 점은 이 영상이 제품 소개 클립이라는 것. 벤치 점수나 요금은 글에 없고, Opus 5.5가 어떤 체크포인트인지도 영상 캡션 표기만으로 확인됨.
공식 영상↓
Opus 5.5 on #MiniMaxDesign | Code Your Next Video
Turn text into video—from code to motion.
✨ JavaScript Animation
✨ Motion Graphics
✨ Explainer Videos & Visual Storytelling
✨ Web & Product Demos
More than a language model—it can interpret visuals and keep refining your creations.🎨
12B짜리 코딩 모델인데, 저장소를 직접 훑고 파일을 고치고 테스트까지 돌리게 바뀜.
JetBrains가 Mellum2.1을 공개함.
SWE-bench Verified 47.0%
같은 에이전트 하네스(Pi)로 잰 JetBrains 자체 점수임. Mellum2 Thinking은 2.0%였음.
핵심:
① 구조는 Mellum2와 같음. 전체 12B 중 토큰마다 2.5B만 쓰는 MoE, 컨텍스트 131,072토큰, Apache 2.0
② 바뀐 건 학습 방식임. 짧은 후처리가 아니라 실제 저장소 안에서 셸·파일 편집으로 수백만 번 돌려 RL을 돌렸고, 테스트가 통과하면 보상을 줌
③ LiveCodeBench v6은 82.0%, HumanEval+ 91.5%. 단일 요청에선 MTP로 약 1.6배, 부하가 클 땐 Qwen3.5-9B보다 토큰을 거의 2배 처리한다고 함
④ Hugging Face에 Thinking 체크포인트가 올라와 있고, GGUF·MTP 헤드는 곧 나온다고 적혀 있음
로컬이나 자체 인프라에 올려 코딩 에이전트·서브 에이전트로 쓰라고 나옴.
한 가지 주의할 점은 점수가 전부 JetBrains 자체 평가라는 것. 같은 표에서 Qwen3.5-9B는 SWE-bench Verified 50.0%, Terminal-Bench 2.1은 21.7%로 Mellum2.1(17.4%)보다 높음.
원문↓
https://t.co/BCOIcgQvCn
@iptton@ishaan_gpt Agree — once it's multi-file / multi-step, the IDE-native loop matters more than which model autocomplete uses. Routing helps cost, but workflow fit is the real gap.
@ishaan_gpt That JetBrains 39% vs 21% figure is useful context — thanks. The piece is about Copilot's local/cloud task routing, not a claim it'll beat Claude Code on that survey. Hard to say if auto-routing closes the gap until it's out.
Claude가 회사 데이터로 대시보드를 만들고, 보고서를 짧은 애니메이션으로 바꿔 줌. Claude Dashboards와 Claude Motion이 베타로 나왔음.
달라진 점:
① Dashboards는 BigQuery, Databricks, Snowflake, Redshift, ClickHouse 같은 데이터 플랫폼이나 Salesforce를 연결해 두고 말로 물으면 대시보드를 만들어 줌. 데이터가 바뀌면 대시보드도 같이 갱신됨
② 숫자를 누르면 그 값을 뽑은 쿼리가 보이고 차트마다 마지막 갱신 시각이 붙음. 더 깊게 볼 땐 Amplitude, Grafana, Hex, Mixpanel 같은 분석 도구로 바로 넘김
③ Motion은 분기 보고서나 차트를 전사 회의용 설명 영상처럼 움직이게 만들어 줌. 영상 생성 모델 없이 코드로 그려서 단어 하나, 숫자 하나, 타이밍만 따로 고칠 수 있고 MP4로 내려받음
④ 같은 날 Docs, Slides, Design은 베타를 끝내고 Free까지 모든 요금제에 열림. 지금까지 Claude로 만든 문서, 덱, 디자인이 4,500만 개를 넘었다고 함
Dashboards는 유료 요금제 전체, Motion은 Team과 Enterprise에서 베타로 쓸 수 있음. Enterprise는 관리자가 직접 켜야 함.
한 가지 주의할 점은 Motion이 실사 영상이나 사람을 만들지 않는다는 것. 글자, 차트, 도형, 이미지를 움직이는 용도이고 길고 복잡할수록 사용량 한도를 더 씀.
공식 영상↓
Cline has launched Cloud Agents. Hand it a task in the browser and it writes and tests code in a cloud sandbox, then opens a PR on your GitHub repo. Nothing to install.
What's new:
① Each session gets an isolated sandbox and its own cline/ branch, and you can run up to 10 in parallel
② Closing the laptop does not stop it. It commits and pushes to the branch as it works, so progress stays on GitHub even if a session stops mid-way
③ Progress, approvals, and new tasks also work from a phone
④ Pick a model at start and switch mid-session. Cline's launch thread says free models such as DeepSeek-V4.1-Flash are available to try, and ClinePass (open-weight subscription) is described as about 5x discounted access
To start: sign in to a Cline account, connect GitHub, open the Agents tab, pick a repo, and describe the task.
One caveat: the result comes back as a PR. Cline says each session's GitHub token is scoped to the single repo you picked, but a human review before merge is still required.
Official video↓
1/ Introducing Cline Cloud Agents.
You can now use Cline straight from your browser. Give it a task and it runs in a secure cloud sandbox, writing code, testing its work, then opening a PR on your GitHub repo.
Same experiment, different GPU count - and the score moves.
DatologyAI open-sourced Zephon, a data loader meant to stop that.
0.82 → 0.05
DCLM Core v1 score swing when the same 1B model is trained for 20B tokens on 8 to 64 GPUs. DatologyAI reports up to 0.82 points with a standard loader, and 0.05 with Zephon (FineWeb suite: 0.64 → 0.011). That 0.82 swing is about four times the 0.2-point bar the DCLM authors used when comparing curation methods.
Key points:
① It tokenizes, packs, and mixes data during training, and still keeps the same global batch order when the GPU count changes
② Resume a 16-GPU run from a mid-checkpoint on 8, 32, or 64 GPUs, and every global batch's SHA-256 hash still matches the original run
③ You do not have to pre-materialize the full processed corpus on disk, so swapping the pipeline does not force a full rebuild each time
It is on GitHub under Apache-2.0, with TorchTitan and Megatron integration code plus a paper.
One caveat: it is not a score booster. DatologyAI notes that on DCLM the Zephon data order happened to score lower; what shrinks is the swing caused by GPU count.
Source↓
https://t.co/gnzm9ltzqA
Anthropic has released Claude Haiku 5.5. It's the small model meant for high-volume, repetitive work - summaries, classification, database queries - and as a subagent under Opus and Sonnet 5.5.
What's new:
① Price is $0.10 input / $0.50 output per million tokens (for prompts up to 100k tokens). Anthropic says that averages about 75% cheaper than Haiku 4.5
② Context is 1M tokens, with up to 128k output. Haiku 4.5 was 200k and 64k
③ First Haiku-class model with an effort setting, so you can trade cost and quality per task
④ On Anthropic's numbers, OSWorld 2.1 (offline subset) hits 72.4%. Haiku 4.5 was 15.7%
There's more in the same launch. Sonnet 5.5 cache reads dropped from $0.20 to $0.10, cutting most agentic work by about 20%. Max and Team subscribers start getting monthly API credits this week (Max 5x $100, Max 20x $200, Team up to $500).
Model ID is claude-haiku-5-5, live on the Claude API plus AWS, Google Cloud and Microsoft Azure.
One caveat: the tokenizer changed. The same text costs about 30% more tokens than on Haiku 4.5, so token-based budgets need a recheck. Anthropic still says complex agentic coding is better on Sonnet or Opus.
Official video↓
Introducing Claude Haiku 5.5: the cheapest, fastest, and most capable small model we’ve ever released.
On average, it costs around 75% less to run than Claude Haiku 4.5.
An open model pushing toward a trillion parameters.
They split it across 11 PCs - no data-center GPUs.
60.29 tokens/sec
Peak aggregate decode when Cascadia ran Thinking Machines' Inkling (975B) across 11 Intel Core Ultra PCs, at 88 concurrent requests. Cascadia posted the figure on its own blog.
Key points:
① The 66 layers were packed in sets of six, one shard per PC. Each PC had 64GB of memory (704GB total), linked over ordinary Gigabit Ethernet
② On each machine the CPU handled attention and routing, and the integrated GPU (Arc B390) ran the expert compute. Expert weights were cut down to INT4
③ The two always-on dense layers were split into eight expert-sized slices that reuse the same fused ops, cutting per-layer time from 8.1ms to 4.5ms
It was a joint experiment with Intel, and the Inkling support code is open source, being upstreamed into the Cascadia repo (Apache-2.0).
One caveat: it's slow if you're the only user. A single request gets 7.96 tokens/sec, and at 88 concurrent requests the median time to first token is 34.61 seconds. Cascadia itself says this setup fits multi-user sharing or background work better.
Source↓
https://t.co/GkPODKt4SI