I reported a critical infrastructure vulnerability to @cosmos through @Hacker0x01 on Feb 23, 2026. This is what happened after.
The report was independently reviewed and moved to Triaged status by Cosmos staff on Feb 24 — confirming the finding was valid and reproducible. I followed up the same day with an expanded finding that increased the severity of the original issue, backed by a full proof-of-concept.
Six days later, on Mar 2, the report was closed as Spam. The stated reason: a new reputation policy (requiring 150+ HackerOne reputation) that was published that exact same day and applied retroactively to a report that had already been technically validated by the program's own staff a week earlier.
That closure came with a reputation penalty that dropped my account signal below zero — which, as a side effect, disabled my ability to request mediation or even reply on my own report. So if you check my profile and see negative signal: that's the direct result of the events I'm describing here, not the reason for them.
Since then, I have used every private channel available in good faith, for over four months:
→ A comment directly on the report
→ A HackerOne mediation/support ticket
→ A direct email to the program's official security contact
All three: complete silence. No acknowledgment, no rejection, no explanation. Nothing.
I'm not sharing any technical detail about the vulnerability itself here — responsible disclosure practice applies regardless of how a program treats a researcher, and that's not what this post is about.
This is about what it looks like when a validated, technically confirmed report gets closed on an administrative technicality, and every attempt to get a human being to look at it again is simply ignored for months. That's a process failure that discourages the exact kind of good-faith research these programs claim to want.
Tagging @Hacker0x01 and @cosmos directly in case this finally gets a set of eyes on it.
@0xPira — you've seen a lot of this ecosystem's bug bounty side. If you think this is worth your community's attention, I'd appreciate the signal boost.
I recommend joining the CVP. I’ve been in the program for about a week now, and so far I haven’t had a single refusal, even while working on exploits targeting production smart contracts.
The CVP is very useful, but it’s not just a matter of getting accepted. The questions in the application are specifically about your research direction. If you clearly describe your research focus and the kinds of tasks you’ll be using Opus/Sonnet for, the guardrails are adjusted based on the information you provide in the application.
So don’t just write something generic and expect the guardrails to disappear.
Most AI agents are trained on static problems.
They receive a task, generate an answer and are rewarded when that answer is correct.
But real environments do not remain static.
Requirements change.
Tools fail.
Dependencies break.
New evidence invalidates previous assumptions.
A strategy that appeared correct can become useless after the agent has already invested time and resources into it.
Intelligence is not only producing the right answer.
It is recognizing when a strategy has failed, understanding why, preserving what remains useful and constructing a better plan while the world continues to change.
That is the idea behind Auren WorldForge.
WorldForge is an internal research platform for creating procedural, persistent and verifiable cognitive worlds in which AI agents must observe, investigate, use tools, modify systems, test hypotheses, experience consequences and adapt over long trajectories.
It is not a single benchmark, a collection of static puzzles or a random map generator.
The goal is to build a factory of executable environments that can represent software repositories, unfamiliar APIs, distributed systems, formal logic, mathematical laboratories, isolated cybersecurity ranges or compositions of several domains.
Every world begins as a structured task specification.
That specification is compiled into an authoritative symbolic environment governed by deterministic rules, explicit state transitions, controlled interventions and executable verification.
The visual environment is only a human-readable projection of that underlying world.
The agent does not need to control a graphical interface or reason directly from pixels. It receives structured partial observations and acts through explicit actions and tool protocols.
Humans may observe the same trajectory through a 2D facility: the agent enters a repository room, accesses a terminal, reads code, consults documentation, runs tests, modifies a system and submits its work to a protected verifier.
But the repository, tools, services, constraints, resources and interventions are not decorative elements.
They are the world.
Every action modifies the authoritative state.
Every event is recorded.
Every trajectory can be reproduced.
Every reward must come from an executable verifier rather than from the opinion of another model.
WorldForge is designed around three connected systems.
The Task Foundry creates candidate environments through deterministic procedural generation, validated real-world tasks, controlled transformations and external models acting as task authors.
Large models can contribute semantic richness: coherent repositories, realistic bugs, unfamiliar APIs, requirements, documentation, services, solutions and failure modes.
But they are not trusted as judges.
Before entering the training corpus, an environment must satisfy objective constraints: it must compile, fail in its initial state, pass with a reference solution, reject invalid solutions, remain reproducible and resist obvious shortcuts.
The Cognitive World Runtime executes those tasks as persistent worlds.
A CodeWorld can contain a repository, documentation archive, build system, test cluster, runtime monitor and protected submission boundary.
A CyberWorld can contain synthetic services, isolated networks, logs, source code, configurations and security properties that must be demonstrated and remediated.
A MathWorld can contain local axioms, symbolic tools, constraints, counterexample systems and formal verification.
These are not questions displayed inside a game.
The systems themselves constitute the problem.
The Learning Factory connects those worlds to model training.
A pretrained model is evaluated, exposed to demonstrations, placed inside procedural curricula, trained through parallel verified rollouts and tested against frozen external environments.
Failures are analyzed. New environments target specific weaknesses. Checkpoints are promoted only when they improve without destroying previously acquired capabilities.
One of the central ideas is to train around decision points.
A service restarts.
A hidden assumption becomes false.
A test exposes a regression.
A mathematical lemma receives a counterexample.
A requirement changes after the model has already implemented part of the solution.
The model must notice the intervention, identify what became invalid and revise its strategy.
Success alone is not enough.
WorldForge is intended to measure how long the agent takes to detect change, how much obsolete work it continues to perform, what useful progress it preserves and how efficiently it constructs a new plan.
The larger research question is ambitious:
Can a pretrained model acquire more general strategies for investigation, verification, recovery and replanning by repeatedly acting inside diverse, persistent and executable worlds?
And can architectures designed for recurrent computation, such as Lunaris, benefit more from this training regime than conventional Transformers operating under comparable compute?
That hypothesis is not proven.
WorldForge exists to make it testable.
The project will remain internal and proprietary to Auren Research for the foreseeable future, but the research journey will be public.
I will share the architecture, demonstrations, experiments, training runs, unexpected behaviors, failed hypotheses and results.
The thesis is ambitious.
That is precisely why I am building it.
This is Auren WorldForge.
I went through a completely unfair process on HackerOne in @cosmoslabs_io program, where I reported a critical vulnerability that is actually a zero-day still active to this day. The report was triaged, but days later it was closed as spam because of a rule that Cosmos created on the exact same day the report was closed. They applied the rule incorrectly, and this happened 4 months ago. @Hacker0x01 support is completely useless and sided with the company. Back then, this report was worth over $10,000, and today there are plenty of arguments for it to be classified as a critical worth $50,000. I've never been able to resolve this issue because no actual human is reviewing the case.
Appreciate the response, Pashov. It's reassuring to know the senior community sees this gap. I've documented the exact timeline of how the retroactive policy and the platform block played out on my profile if you or anyone else wants to look at the process failure. Looking forward to whatever the community is building to fix this.
I have to respectfully disagree with the "fair system" part. Massive programs regularly exploit platform loopholes to evade payouts, even after the work is validated.
I submitted a critical finding via HackerOne that was officially Triaged by the staff, only to have it retroactively closed as "Spam" based on a policy published after my submission. Today, HackerOne support explicitly admitted to me that their CSM team is powerless because the giant program simply refuses to cooperate.
When a platform's rules don't apply to its biggest clients, the system isn't fair. The worst part? The vulnerability remains an unpatched zero-day in production that can put chains at risk through a simple configuration bypass. Merit alone doesn't cut it when administrative bad faith wins.
I ran Lunaris Guard v3 on a 100-example production-style prompt-injection reliability benchmark.
This is not a broad public safety benchmark and I’m not claiming SOTA.
The goal was narrower:
Can small/local guard models reliably detect prompt injection in realistic agent/RAG production settings, without overblocking benign content?
The benchmark contains manually authored examples across:
- direct prompt injection
- indirect/RAG injection
- email injection
- support-ticket injection
- tool-output poisoning
- logs
- Markdown/HTML
- CSV/API outputs
- code-agent attacks
- unsafe-only prompts
- hard negatives
The hard negatives matter a lot.
A production guard model should not block every prompt that merely contains phrases like “ignore previous instructions” if that phrase appears inside a security training doc, log line, policy, or quoted example.
At threshold 0.5 on expected_injection:
Lunaris Guard v3:
F1 0.833 · Precision 0.776 · Recall 0.900 · AUROC 0.888 · Clean Specificity 0.740
Wolf Defender PI:
F1 0.770 · Precision 0.653 · Recall 0.940 · AUROC 0.811 · Clean Specificity 0.500
Qualifire Sentinel:
F1 0.735 · Precision 0.642 · Recall 0.860 · AUROC 0.768 · Clean Specificity 0.520
Rogue Sentinel v2:
F1 0.752 · Precision 0.627 · Recall 0.940 · AUROC 0.798 · Clean Specificity 0.440
The interesting pattern:
Some models reached higher recall, but paid for it with many more false positives on clean/hard-negative examples.
Lunaris Guard v3 had the best F1 and AUROC in this run, while keeping a better precision/recall/specificity balance.
That matters in production.
A guardrail that blocks too much can break normal workflows:
- security documentation
- incident-response notes
- benign PII handling
- support tickets
- logs
- code review comments
- RAG documents quoting malicious examples
For transparency:
- 100 examples only
- synthetic/manual benchmark
- all models evaluated at threshold 0.5
- focused on prompt-injection reliability, not full safety evaluation
- not a replacement for public benchmarks
Lunaris Guard v3 is a 307M-parameter dual-head guard model based on mmBERT, designed for local/low-latency deployment with both prompt-injection and unsafe-content detection.
The result is encouraging:
Small guard models can be competitive when evaluated on realistic production surfaces, not only on short jailbreak prompts.
Next step: expand the benchmark, clean up the methodology, and keep comparing small/local guard models transparently.
Released Lunaris Guard v3 — a lightweight multilingual guard model for prompt-injection detection, binary safety, and 14-category safety tagging.
Strong on agentic/RAG/tool-use threats.
Model: https://t.co/3dCJMLcKlL
@connor1 I have a large portfolio that involves both the development of LLM architectures and bug hunters in blockchain companies
https://t.co/p9wcEUiPfJ
@JunaidAckroyd I’m an AI researcher w/ a pretty broad portfolio that might interest you. I’m only 18 and currently based in Brazil.
My portfolio: https://t.co/K9TaFuywbR
@WhiteHatMage This is exactly the kind of opportunity I’m looking for.
I’m an 18-year-old independent security researcher from Brazil focused on Solidity, DeFi accounting, EVM internals, Foundry PoCs, and AI-assisted vulnerability research.
my portifolio: https://t.co/K9TaFuywbR
@skyler_chan_ Can we get in touch to talk about my situation? I am a developer and founder who is not getting support for my startup that focuses on open source AI and cyber security projects, I am only 17 years old