Claude Code Projects의 메인 채팅이 Low여도 코드 분석과 계획을 맡는 스레드의 effort는 따로 설정할 수 있네요.
@lydiahallie 설명에 따르면 각 스레드는 독립된 세션이고, 메인 채팅은 작업을 배분하고 결과를 전달합니다.
‘Project settings > General’에서 Coordinator와 Thread의 모델·effort를 각각 바꿀 수 있습니다.
Pro·Max에 순차 제공 중인 공개 베타라, 계정에 Projects가 표시되어야 쓸 수 있습니다.
If you're using Claude Code Projects and had to bump the main chat's effort, I'd love to know why! Low is the default since it mostly just coordinates threads, but curious where that falls short
You can override the defaults in Project settings. Sonnet 5.5 is in there too 👀
Databricks 내부 비교에선 GPT-6 Luna의 작업당 비용이 Opus 5.5의 20분의 1 이하였다고 합니다.
Patrick Wendell이 개발자 2,400명의 실제 사용과 별도 평가를 묶어 공개한 결과입니다.
그런데 회사에서 일상적인 코딩에 추천한 모델은 Opus 5.5였어요.
Luna의 품질 평가는 아직 진행 중이고, 전체 벤치마크도 공개 전입니다.
Crazy few weeks for model releases! Our findings @Databricks show several new models meaningfully advance the pareto frontier. Results below (online workload analysis of N=2,400 engineers, plus offline evals):
1. Two of the three models released last week clearly expand the cost/quality frontier: Opus 5.5 and GPT-6 Luna.
2. Opus 5.5 is now the highest quality mid-tier model. It is better than all prior Opus models, better than GPT-6 Sol, and better than GPT-5.6 Sol.
3. Opus 5.5 reduces same-task costs consistently by 20% in both offline and online analysis. This is against a baseline of Opus 4.8, the prior least-cost Opus model (Opus 5.0 was a bit of a dud with high costs and barely noticeable quality improvements).
4. Due to best-in-class quality and lower costs, Opus 5.5 is a strong candidate as an “every day default” model for coding, and we are now encouraging it for this purpose at Databricks.
5. GPT-6 Luna is very, very, very cheap. It was at least 20 times cheaper per-task than Opus 5.5 in every offline benchmark we tested and in observed online use.
6. GPT-6 Luna is surprisingly capable given how cheap it is. On one of our most difficult evaluation suites it roughly matches Opus 4.6 performance, while being 99.3% cheaper per-task than Opus 4.6 was at that time. That's a 100X cost reduction in ~9 months! This finding is preliminary and we are still evaluating Luna quality on a broader set of offline and online tests.
Our production setup: Unity Gateway to route workloads across models and trace agentic interactions. A mix of end-user harnesses including: Omingent (meta-harness), Claude Code, Codex, and Cursor.
Claude Code로 프롬프트를 고쳐도, 따로 남긴 시험 사례에서 개선이 없으면 되돌리는 방식이네요.
Anthropic이 공개한 평가 자동화 가이드입니다. 기본 claude-api 스킬에서 /claude-api build-eval로 평가를 만들고, 사람이 사례와 채점 기준을 확인합니다.
/claude-api hillclimb은 한 번에 하나씩 바꾸며, 성능 개선이나 품질을 유지한 비용 절감을 목표로 합니다.
결과가 측정 오차 안에 있으면 병합하지 말라고 보고하는 흐름까지 담았습니다.
Claude can now help you build evaluations and hillclimb on them.
In this article, we share guidance on eval design & skills that Claude Code can use to improve your applications.
https://t.co/PgKFC2DWth
가구를 옮기기 전에 거실 사진 5장으로 배치를 바꿔보는 Claude Code 실험입니다.
@Skylartkitchen이 스킬을 공개했습니다. 방 사진과 새 가구의 치수(또는 상품 링크)를 주고, 2~3개 배치안에서 문이 막히거나 가구가 겹치는지 확인하는 흐름입니다.
방 치수는 사진 속 사물을 기준으로 추정합니다. 작성자도 댓글에서 실측 치수나 평면도를 함께 주는 방법을 제안했습니다.
Took 5 phone photos of my living room and Claude Code rebuilt it in 3D. Then I turned it into a skill.
Now I can try different layouts and drag furniture around before moving anything heavy. It even flags when a chair gets boxed in.
Inspired by @scheemunai's room planner
블로그 글을 Claude·ChatGPT가 찾아 읽게 하는 ‘요미타스’가 공개됐습니다.
@tkashiwazaki2의 서비스로, 사이트 등록 → 소유 확인 태그 삽입 → 수집 → MCP 주소 연결 순서입니다.
무료 수집은 요금 안내 기준 사이트당 300페이지. ChatGPT는 개발자 모드 사용 권한이 필요합니다.
앱에 Codex·Claude Code를 붙일 때 파일 전달·진행 표시·작업 취소를 같은 API로 묶을 수 있네요.
Avi Chawla가 소개한 HarnessRouter입니다. Docker로 띄우고 모델 제공사의 API 키를 연결한 뒤, 요청의 metadata.harness_id로 실행기를 고릅니다.
취소 전에 보낸 이메일처럼 이미 끝난 동작까지 되돌리지는 못한다고 작성자가 덧붙였습니다.
System 1 vs. System 2 Agent Harnesses, clearly explained:
(bookmark this)
There are two different ways to put AI inside an application.
A System 1 harness asks the model to make a bounded judgment.
The app provides a request and relevant state. The model selects a constrained result such as a category, score, route, or extracted field.
Code checks confidence and policy before executing the result or escalating the case.
The model does not own the workflow. It selects from choices supplied by code.
This works well for intent classification, model routing, risk gates, reranking, and other high-volume decisions where the possible outputs are known in advance.
Jev is designed for this System 1 role to make fast, bounded judgments while application code controls the workflow.
A System 2 harness handles work whose path cannot be specified upfront.
The LLM receives a goal, constraints, and current state. It plans the next step, while a guarded router controls which tools it can invoke.
Tool results return to the harness, which updates state and verifies progress. Each cycle can end in a final action, a question for the user, or another plan.
The LLM helps drive the workflow, but it should not have authority. Tool permissions, budgets, stop conditions, and final execution remain in application code.
This pattern is useful for coding, research, incident investigation, and other tasks involving several dependent actions.
So this is not primarily a small-model versus large-model distinction but rather a difference in control flow.
Most production apps need both: System 1 for fast, bounded judgments and System 2 for ambiguous or multi-step work.
The next engineering problem is running both patterns without maintaining a separate execution layer for every harness.
And the solution to this is now actually open-source and implemented in HarnessRouter.
It provides the infrastructure layer between an application and the harnesses that execute its work.
- Its System One base can run Jev and other decision models that select finite actions using typed outputs and confidence gates.
- Its agent-harness bases support runtimes such as Codex, Claude Code, and Hermes, which can plan, call tools, manage files, and complete open-ended tasks.
GitHub repo: https://t.co/4y6zJShKSu
(don't forget to star it ⭐)
HarnessRouter does not make these systems reason identically. Instead, it only standardizes the infrastructure surrounding their execution.
Through the Unified Harness Protocol, an application gets one contract for starting tasks, streaming progress, continuing sessions, exchanging files, cancelling work, and reporting structured failures.
This lets an application select the appropriate harness without rebuilding its task lifecycle around every runtime.
If you want to dive deeper, Akshay wrote a detailed article explaining how HarnessRouter and the Unified Harness Protocol provide this infrastructure layer.
It covers why model routing is not harness routing, how tasks and sessions work across runtimes, the local setup, and a complete API call.
Read it below.