https://t.co/J6eqNoarkn 넷플릭스에서 볼 영화를 고르는 일이 너무 지겨운 사람들을 위한 따끈따끈한 신작 웹사이트, ‘빈지버스터’입니다. 구독 중인 스트리밍 서비스를 고르면, 그곳에서 볼 수 있는 영화들이 3D 비디오 대여점에 진열됩니다. 가게를 돌아다니며 케이스를 하나씩 집어 들고, 마음에 들면 예고편을 구경하고, 내키면 곧장 영화를 보러 갈 수도 있어요. 역시 영화는 선반에서 골라야 제맛이죠.
Opus 5는 왜 함께 작업하기 더 불편하게 느껴지는가?
4.7·4.8보다 역량과 벤치마크 점수는 높지만, 모호한 의도에 질문하기보다 가정하고 실행해 세밀한 감독이 필요
보상 강화학습은 과감하게 가정하는 모델을 선호하지만, 코딩에는 추측보다 질문하는 에이전트가 더 적합함
https://t.co/TcH3nOCRA0
오늘 학생분의 장례식에 다녀왔다. 림 랜지 같은 감독이 되고 싶어했다. 평론을 잘 썼고 철학을 좋아했다. 인테리어를 좋아해서 짙은 나무와 지브리로 꾸민 방 사진을 종종 보내주었다. 유도부 주장이어서 끈기 하나는 자신있다고 했다. 영화를 정말 하고싶어했다. 영화과에 입학한지 고작 한 학기가 지났다. 성실하고 끈기있어서, 24시간 동안 촬영을 하는 극한의 스케쥴을 소화하다 사고를 당했다.
I do not personally read most of the code my agents write. I occasionally inspect specific implementations, but line-by-line code review is not my default workflow.
I spend my attention on specifications, architecture, invariants, documentation, test results, and high-risk boundaries. My AGENTS.md explicitly requires documentation and implementation to remain synchronized, while automated Codex reviews inspect diffs for local mistakes, violated constraints, and code–documentation drift. I read code selectively when those systems surface uncertainty or when a change crosses an important boundary.
I have also built a system in which agents decompose larger tasks into appropriately small, bounded increments before implementing them. The resulting PRs are usually small enough to reason about, test, revert, and review automatically, so oversized PRs rarely become a problem.
This is why “your problems are basic” misses the point. Good engineering deliberately turns complex problems into small, constrained, independently verifiable ones. “Basic” is not necessarily a property of the original problem; it can be the product of good decomposition, architecture, and tooling.
A model adding a cargo-culted 700ms delay proves that models need verification. It does not prove that humans need to read every generated line. My goal is not full, blind autonomy. It is to build a process in which most code does not require my direct attention—and where mistakes are cheap, visible, and contained. I’m not trying to become better at reading AI-generated code. I’m trying to build a development process where reading most of it is unnecessary.
codex tip: I always use define-goal instead of /goal. It's basically like having a master level genius AI write your goal better than you could. And it's an official OpenAI skill.
You can get your prompts to run longer and your output better. Ask a high effort Sol to research + plan it. Sol then spawns a new thread for a workhorse model (Luna Max) to execute the work and return a result to Sol for review. LLM-as-judge to determine if Luna's output is acceptable progress toward goal
https://t.co/mB1m0lDa6w
For anyone with endless ideas, this agent age is nirvana as those ideas are met with endless execution, endless exploration. I've never had has much fun working with computers as I do right now. What a time to be alive.
역사도 그렇고, 길은 직선이 아니라고 생각한다. 직선이었으면 좋겠지만. 난 계단식 성장이란 말 이젠 잘 모르겠다. 하강식 성장? 뒷걸음치면서도 얻는 것이 있다. 세상은 그저 종과 횡으로 나뉘지 않는다. 층위의 개념이 있다. 그 층위마저도 다시 떨어지기도 한다. 그러면 질문을 바꿔야 하지 않을까
그리고, 다시 우선 순위에 대해서 생각한다. 모든 일들이 필요한 것들은 아니었다. 이젠 풀어야 할 문제를 푸는 것이 중요하다는 생각이 든다. 하지만, 그러한 생각을 어떠한 병렬적 탐색이 아니고서 다시 느낄 수 있었을까? 돌고 돌아서 다시 원점으로 온다고 하더라도 원점은 같은 원점이 아니다.
지금 사람들은 글의 위력을 과소평가한다. 하라리가 글을 인류의 운영체계OS라고 말한 것은 정확히 옳다. 그 점을 파악한 AI(개발자)가 그걸 자기(AI) 것으로 만들어 가고 있는데도 태연하다. 그런 AI를 잘 쓰기만 하면 된다고 손쉽게 생각한다. ‘잘 쓰기’ 위해서는 자신이 읽기와 쓰기를 통해 말과 글을 제대로 배우고 익히고 발전시켜 가야 한다는 사실을 잘 이해하지 못하거나 알고 있다면서도 진지하게 여기지 않거나 진지하게 여겨도 그에 따른 수고는 기피한다. 이미 평소 시선과 함께 무엇이 중요하지 식별하는 데 필요한 주의를 쉽게 내주는(자신이 제어하지 못하도록 하는) 상황에 빠져 있기 때문이다. 우리는 잘못을 적극적으로 (그것도 지속적으로) 선택하는 경우는 드물다. 점점 타협하거나 미끄러진 끝에 어딘가에 이를 뿐이다.
16년 한강 작가님의 맨부커상 수상(한국인 최초)이후 열린 기자간담회에서 인상깊던 구절
"우리 삶에 다른 사람들의 삶도,현재와 과거도 다 들어와 있기 때문에 고통을 피하는 것은 불가능하다고 생각해요.할 수 있는 것은 그것을 응시하는 것이고 우리 삶의 일과로서 같이 가야하는 거라 생각합니다"