Today marks the first-ever release of Cross-Layer Transcoders for Qwen3.
BluelightAI has trained CLTs for Qwen3-0.6B and 1.7B, creating an explorable set of interpretable features that capture how Qwen3 represents concepts and transforms information across its layers.
The Qwen3 Explorer allows you to examine these features directly, identify structure in the model’s representations, and use this understanding to analyze behavior, diagnose failures, and guide adaptations of Qwen3-based systems.
Black Box to Glass Box: Audit-Ready AI
Modern AI models, including large language models, can perform a wide range of tasks with minimal adaptation. However, this flexibility can come at the cost of transparency and interpretability: why did the model give the answer it did? At BluelightAI we are working to make AI systems more transparent so that models can be deployed to critical roles while maintaining trust and observability. This is especially critical in regulated industries like banking and insurance, where frameworks such as SR 11-7 require documented proof of what drove each model decision before deployment can proceed.
In this blog post we’ll outline one approach we are developing for interpretable classification and present a case study using this approach for intent recognition.
https://t.co/S4pPlfP7UI
Mapping Concept Evolution in Qwen3
We often describe Large Language Models (LLMs) as "black boxes." We observe the input and the output, but the internal machinery – the billions of calculations occurring in between – remains largely opaque. We observe that the model understands concepts, but we rarely discern how it constructs them. It is vital that we understand the “how”, because it will give us better information about how to control the LLMs and AI, and also diagnose possible malfunctions, like the introduction of unacceptable biases or the production of undesirable language or modes of communication. This kind of control will also make the adaptation of LLM technology to specific application domains, such as financial or legal documents, simpler and more direct. https://t.co/TZJ7YIF52T
Mapping Concept Evolution in Qwen3
We often describe Large Language Models (LLMs) as "black boxes." We observe the input and the output, but the internal machinery – the billions of calculations occurring in between – remains largely opaque. We observe that the model understands concepts, but we rarely discern how it constructs them. It is vital that we understand the “how”, because it will give us better information about how to control the LLMs and AI, and also diagnose possible malfunctions, like the introduction of unacceptable biases or the production of undesirable language or modes of communication. This kind of control will also make the adaptation of LLM technology to specific application domains, such as financial or legal documents, simpler and more direct. https://t.co/TZJ7YIF52T
I'm pretty proud of this: We trained cross-layer transcoders for Qwen3 and built a dashboard for exploring the features using TDA-based graph visualizations.
Today marks the first-ever release of Cross-Layer Transcoders for Qwen3.
BluelightAI has trained CLTs for Qwen3-0.6B and 1.7B, creating an explorable set of interpretable features that capture how Qwen3 represents concepts and transforms information across its layers.
The Qwen3 Explorer allows you to examine these features directly, identify structure in the model’s representations, and use this understanding to analyze behavior, diagnose failures, and guide adaptations of Qwen3-based systems.
Today marks the first-ever release of Cross-Layer Transcoders for Qwen3.
BluelightAI has trained CLTs for Qwen3-0.6B and 1.7B, creating an explorable set of interpretable features that capture how Qwen3 represents concepts and transforms information across its layers.
The Qwen3 Explorer allows you to examine these features directly, identify structure in the model’s representations, and use this understanding to analyze behavior, diagnose failures, and guide adaptations of Qwen3-based systems.