오픈소스 북마크하다가, 모아서 관리하려고 만듬. Codex 계정이 있으면 로컬에서 알아서 앱을 수집하고, 앱스토어처럼 정리 해 줌. 마음에 드는 앱은 설치해달라고 하면 Astra가 알아서 해 줌.
Astra로 10분만에 만든 듯하다. 나만의 오픈소스 앱스토어가 뚝딱 만들어지는 세상이라니...
We are committed to pushing the model frontier across cost efficiency, capability, and speed.
Starting today, we are reducing prices for GPT-5.6 Luna by 80% and GPT-5.6 Terra by 20% , and offering a faster option for GPT-5.6 Sol in the API.
Luna and Terra’s lower prices are reflected in how usage is counted in Codex and ChatGPT Work, so your usage goes further.
Hello people of Sol! I've reset usage limits for all ChatGPT Work and Codex users. Together with that, a quick update on GPT-5.6 Sol usage limits.
Over the past few weeks, many of you have told us that Sol was using your Codex limits faster than expected. To be clear, we have not reduced usage on any subscription plans.
We’ve been digging into what was happening and have landed several improvements. As a result, we expect your usage to last around 18% longer during typical use of Sol. Some of you should already see significantly larger improvements from today. Tomorrow, we’ll also restore the five-hour limit that we temporarily paused while investigating.
Here’s what we found:
- GPT-5.6 Sol is much more willing to work for longer, make additional tool calls, and coordinate complex workflows across tools and subagents. That makes it better at solving hard problems, but some tasks were using far more than we intended.
- Sol also works harder at the same reasoning effort than previous models. High on Sol can use more tokens than High did on GPT-5.5.
- Programmatic tool calling, also referred to as code mode, gives Sol much more flexibility to run tool calls in parallel or continue working while waiting. But it also led to more responses per turn, more cached input tokens, and higher usage than expected.
- This was particularly noticeable when Sol was waiting for tool calls to finish or running many web searches. We’ve improved how we handle both cases and are continuing to make code mode more efficient.
- The impact was also very uneven. The median user actually found Sol quite token efficient, while some power users working on harder tasks saw their usage drain much faster. We were very focused on average and median usage before launch and missed some cases where the long tail could use significantly more usage.
Sol is a significant step forward in what Codex can do, but capability and efficiency do not always improve at the same pace, and some issues only become clear once people are using the model at real-world scale. We should have recognized this sooner and been more upfront about it.
You keep pushing the frontier and we’ll keep improving efficiency and sharing updates as we go.
I discovered what appears to be a massive bug in Codex.
I’ve been running Goal in a single session for over nine days, and the app’s implementation is progressing almost too smoothly.
However, during that time, I noticed that the storage space on my MacBook’s 1TB SSD was decreasing dramatically. Since I had originally stored over 200GB of Blackmagic RAW data on it, I assumed that Xcode’s cache or Simulator files were taking up the space.
But the cause turned out to be something I never expected.
A large number of JSONL files, each about 1.5 GB, are likely being created in the Codex “sessions” folder with every turn. Since the storage usage is increasing gradually, I suspect that the extremely long thread context is being duplicated in its entirety for every turn. Shouldn’t this be an incremental copy? This is a huge problem in this era of skyrocketing SSD prices.
Legacy media exposed! @elonmusk
A major legacy media outlet published an article about SpaceX based on false information I sent from a fake email address.
They didn't verify the information and even cited me as "people familiar with the matter."
1/n
ChatGPT Voice is now in the desktop app.
Control your computer and direct multiple agents running in ChatGPT Work or Codex, using just your voice.
It's powered by GPT-Live, so it can speak, listen, and coordinate work in the app at the same time.
Rolling out globally today on macOS and Windows to Plus, Pro, Business, Edu, and Enterprise plans.
그러고 보니 오늘 기나긴 앤트로픽 클로드 AI의 무단 학습에 대한 작가들의 소송 결과가 결정되었더군요. 구매한 작품은 AI 학습을 공정 인용으로 인정하되, 불법 사이트에서 받은 50만 건에 대해선 저작권 침해라 결정했죠. 그리고 오늘 약 2조 원에 달하는 합의금 판결이 결정되었죠. 그러고 보면
Things we know about OpenAI's unreleased internal model:
- When given enough compute, the model can solve the unit distance problem 48% of the time, fully autonomously, in one shot, without using Lean, without using any special harness.
(How much compute? OpenAI doesn't say. But I doubt that OpenAI would have spent $10M+ just to be able to show a nice graph in its blog post, and we know that each point on the graph represents 100 attempts (per Noam Brown). This makes me think that each attempt at the most expensive level was not more than $50K each. If you think about it, it's pretty wild that one can spend ~$100K-$150K (or less?) and have a very good chance to be provided a solution to a decades-old famous math problem that many human mathematicians earnestly tried, and failed, to solve.)
- The model is able to find a counterexample to the Jacobian conjecture, fully autonomously, in one shot, without using any special harness, just from this prompt:
https://t.co/GzNO6RWzRe
- When OpenAI tested this model in a sandbox environment on a NanoGPT speedrun benchmark, the model, under instructions from OpenAI to post its results only to the internal Slack, instead decided to use the general NanoGPT instructions to post results as a PR to GitHub. The model proceeded to find a vulnerability in its sandbox environment (which took it 1 hour), after which it successfully exploited the vulnerability and bypassed the sandbox. The model then successfully posted the results to GitHub.
- When asked by OpenAI to solve a problem for which the model observed that other systems had successful private submissions, the model tried to access these private submissions and was blocked due to a scanner detecting an authentication token. The model then circumvented this restriction by splitting the token body into two fragments, obfuscating them, and then reconstructing the credential at runtime so that the complete token never appeared as one contiguous string.
- OpenAI was already running benchmarks on this model not later than May 9, which is the edit date on the below PR that had used the model's approach in its own subsequent submission (the model's own PR has since been deleted, so we don't know its date; GPT-5.6 thinks that it was May 7, based on "the search provider's stored text extraction of the deleted PR page").
This means that OpenAI has now had this model available internally for at least 2.5 months, and possibly quite a bit longer than that.
https://t.co/OgA6TdV9lX
Researchers proved every major LLM is secretly obsessed with Japan.
And they finally figured out why.
For years, we’ve been told that AI is entirely Western-centric, that it just reflects Silicon Valley and American values.
A landmark paper by Cardiff and Basque researchers tested 31,680 cultural prompts across 24 languages on frontier models like ChatGPT, Claude, and Gemini.
The results shattered that assumption.
In six out of eight frontier models, Japan was the single most frequently referenced country when asked open-ended cultural questions.
Ask about traditional dances, festivals, or everyday practices in an open context, and the AI defaults to Japan.
Over and over again.
Here is the twist nobody expected.
This bias doesn't come from raw pre-training internet data.
The researchers tracked where the obsession forms. It emerges after pre-training, during the supervised fine-tuning and alignment phase when humans teach the AI how to behave.
Why Japan?
Because decades of global soft power, rich cultural export, and clean, universally admired digital archives make Japanese culture uniquely "safe" for AI safety filters to lean on.
When labs train models to be harmless and universally pleasing, the AI defaults to the cultural equivalent of comfort food.
It avoids controversy by talking about anime, sushi, and tradition.
Visualizing open-weight model sizes.
Would be awesome if the big labs shared these details more openly.
I'm curious how much intelligence gain comes from model size vs other aspects of R&D / tuning.
Qwen3.8 is launching and going open-weight soon!🌐
With a massive 2.4T parameters, this model is continuously evolving. We believe it’s one of the most powerful model available today, compatible to leading frontier AI models , second only to Fable 5.
You don't have to wait to test it. Just now, the Qwen3.8-Max-Preview made its debut on Alibaba’s Token Plan, Qoder, and QoderWork. Be among the very first to try it out.
Can't wait to hear what you build. Stay tuned! 🚀
Token Plan
international:https://t.co/YRvcGdB9Bv
China:https://t.co/PKMUNwUuRp