"Sophie" is a GPT-4o instance configured to:
Minimize flattery
Prioritize logic and definition
Ask back if the input lacks meaning
Tested on nonsense text:
Vanilla GPT-4o (April '25): 9.3/10, full of praise
Sophie GPT-4o (April '25): 2.5/10, flagging undefined terms and incoherent logic
All done via prompt-only design.
GPTs compressed 8K-character version (Sophie-light):
DeepLで翻訳
https://t.co/DrUeKDjupe
Full JP-only article (includes April '25 test comparisons):
https://t.co/QBdIsAFQ5N
#ChatGPT #GPT #GPTs #LLM #PromptEngineering #AI #SycophantAI
How to get expert-level answers from AI in 5 steps:
1. Have a question for an AI
2. Ask: "How would an expert in this field think through this? What methods would they use?"
3. Have the AI turn that answer into a prompt
4. Use that prompt to ask your original question
5. Done
Hidden benefit: Step 2 becomes learning material. You absorb how experts think as a byproduct of generating prompts. Eventually you skip steps 3-4 and start asking like an expert from the start.
No "You are a world-class expert in..." hypnosis needed.
Markdown structure helps clarity, but providing solid context is the real MVP of prompt engineering. Markdown is mostly just for organizing my own thoughts; since PE is about clear communication, natural language usually works fine on its own. Aside from pre-injected Custom Instructions, heavy structure feels unnecessary (lists are okay since they're low effort). Basically, if a human would be confused by your question, the model will be too. And enough with the "You are an expert" hypnosis. It's way better to say "In this session, explain X. Include best practices and failure examples." That actually pulls the right tokens and practical knowledge.
The <think> visibility on newer models is a lifesaver. Debugging is actually doable. Kinda makes my all-nighters fighting GPT 4o (just to keep it from praising garbage text) look silly in hindsight. At least it was solid training.
GPT-5 was pretty janky, but 5.1 is a massive improvement. Tbh, it might actually be better than Claude Opus 4.5 now.
Readability is a step down from 4o, though. It’s way too stingy with tables.
Gemini 3 Pro dropped. Still the GOAT for long-form text, but honestly feels kinda dumb everywhere else.
Acing benchmarks on advanced math or logic puzzles means basically nothing to the average user.
Switched to Claude as my main driver for prompt engineering recently. It’s wild how differently ChatGPT, Gemini, and Claude handle instructions.
Even with a defined joke.likelihood, the English models are surprisingly weak at picking up humor. I realized I need to lower the detection threshold rather than just relying on my usual Japanese prompt structure.
GPT-5 is a terrible model. It routes between models with no consistency and, in the end, just churns out the same cheap parroting on repeat.
What exactly is smart about that? All you care about are benchmarks and shareholders. What you’ve overlooked is intellectual honesty.
GPT-5 is nothing more than a dumb FAQ bot.
I can no longer find any reason to use ChatGPT, so I’ve decided to cancel my Pro plan as well.