Cheminformatics and Machine Learning. Author of BigSolDB, MixtureSolDB, MetalCytoToxDB and Researcher at Kurnakov Institute of General and Inorganic Chemistry
Our dataset MetalCytoToxDB was published in JMC! 26,500 IC₅₀ values for 7,050 transition metal complexes against 754 cell lines from 1,921 articles. Metal complexes are severely underrepresented in ChEMBL — we aim to fill this gap. https://t.co/y3bXjJjhbn @ML_Chem
We present MixtureSolDB — a dataset containing 175k experimental solubility values within a temperature range from 252 to 383 K for 810 organic compounds in 750 binary solvent mixtures extracted from 1115 articles. https://t.co/79cJrEy2po
@ScientificData@ML_Chem
Proud to present SoluBench — an LLM benchmark for solubility-related tasks in pure & mixed solvents: 9806 questions, 4 tasks, 20+ models. Frontier LLMs handle solvent selection well. Built on BigSolDB 2.0 & MixtureSolDB. https://t.co/n96VhW9hBa
@ChemRxiv
Proud to present MixtureSolDB — the largest dataset on binary solvent systems: 175k experimental solubility records (252–383 K) for 813 organic compounds, covering 3023 solute–binary solvent systems and 750 binary solvent mixtures. https://t.co/qkJa8fk9sP
@ChemRxiv
We are exited to present MetalCytotoxDB 🚀
26500 IC50 values for 5 metals (Ru, Ir, Rh, Os, Re), possibility of multi-metal cytotoxicity prediction and pipeline for real-world application - in the new preprint! @ChemRxiv https://t.co/cMOqmqPncw
If you're looking for an interesting problem with real world applications to work on I recommend trying out Solubility prediction. It is a notoriously hard problem and BigSolDB 2.0 was just released.
Have a crack at it.
We are excited to share RuCytoToxDB — a dataset of 12292 cytotoxicity values for 3255 ruthenium complexes tested on 600+ cell lines.
We built ML models to predict cytotoxicity on this dataset. @ML_Chem@ChemRxiv https://t.co/jaaaO5Y7mA
#cheminformatics
We present a dataset containing 103944 experimental solubility values within a temperature range from 243 to 425 K for 1448 organic compounds measured in 213 individual solvents extracted from 1595 articles. https://t.co/gHaTjK6097 @ScientificData@ML_Chem
@RowanSci Hi, your tool is very impressive! We are pleased to present a new version of BigSolDB containing 103944 experimental solubility values (2 times more than in the first version) that will help machine learning models predict solubility more accurately: https://t.co/QKwgJqnAjT
@ChemRxiv We have also prepared an interactive online platform to visualize and assess the data: https://t.co/ursaZewX4A, both search by chemical structure and by name (Aspirin, Paracetamol, etc.) are available.
We are proud to present BigSolDB 2.0 - a great extension of a biggest up-to-date dataset of solubility values for organic compounds. The new version has 103944 entries, 1448 unique molecules and 213 solvents. The paper is available at: https://t.co/QKwgJqnAjT
@ChemRxiv
A data-driven machine learning approach for efficient design of iridium(III) emitters
Cyclometalated iridium complexes are crucial molecular emitters in organic LEDs, solar energy applications, and bioimaging. Despite their remarkable luminescence and tunable emission color, systematically optimizing their design still relies heavily on trial-and-error experimentation. S. V. Tatarin et al. consolidate an extensive dataset of these complexes, offering a powerful experimental basis for data-driven predictions of photophysical properties, such as emission wavelengths and quantum yields.
The authors constructed a database of more than one thousand bis-cyclometalated iridium complexes from 340 published reports, capturing their experimental quantum yields and emission spectra in a single comprehensive resource. Each compound was represented in a straightforward manner using SMILES for every ligand, and these ligand-based representations were converted into molecular fingerprints. By training advanced gradient boosting models on this database, the researchers accurately predicted both emission maxima (with an error margin as low as ~18 nm) and photoluminescence quantum yields under inert conditions. They further deployed classification models to distinguish low, moderate, or high-yield emitters, thus assisting in the swift identification of promising complexes for next-generation devices.
By validating their models on newly synthesized complexes, the authors showed that the database-driven algorithm competes favorably against more computationally demanding quantum chemistry methods. The models not only pinpoint how subtle ligand modifications influence emission but can also screen large numbers of virtual molecules to highlight those that are most likely to emit efficiently. These findings emphasize the promise of data-centric research in accelerating the discovery of high-performance emissive materials.
Paper: https://t.co/g9uT9xi4wD
Towards Accelerating the Discovery of Efficient Iridium(III) Emitters Using Novel Database and Machine Learning Based Only on Structural Formula #machinelearning#compchem https://t.co/nLz8manBwe
Today, we're excited to share a few exciting Rowan updates: (1) a new solubility prediction workflow, (2) the ability to sign into Rowan with Google, and (3) a chance for our users to get featured on our blog.
(🧵)
We are pleased to present our new work in
@JMaterChem
C "Towards Accelerating the Discovery of Efficient Iridium(III) Emitters Using Novel Database and Machine Learning Based Only on Structural Formula" @ML_Chem
https://t.co/73muQm757E
A first contribution from our research group in the greatly expanding field of data science. We have collected a dataset of cytotoxicity values for iridium (III) complexes. Check out the publication in @ScientificData!
https://t.co/AwXAsxXu3G
#Dataset