Search pages, resources, and actions
Models, datasets, and services for materials science research.
Text-guided crystal generation — natural-language prompt to structure; the obvious LLM-agent-facing generative tool for PRISM.
Autoregressive transformer over Wyckoff positions — generate BY space group (symmetry-conditioned de novo + CSP in one checkpoint); tiny, CPU-tolerant; RL-finetuning support.
Latent diffusion trained JOINTLY on molecules (QM9) + crystals (MP20) — one deploy covers periodic and non-periodic generation.
Newest verified-weights property predictor; also does anisotropic displacement parameters (thermal ellipsoids) — a task nobody else serves.
Only verified source for pretrained bulk/shear modulus + Fermi energy + Poisson ratio checkpoints — covers property slots nothing newer ships weights for.
Property head of M3GNet (the registry previously carried only the potential); 3-body graph features beat MEGNet on formation energy.
Multi-fidelity band gap (PBE/GLLB-SC/HSE/SCAN in one model) — users pick fidelity per credit tier.
Formation-energy workhorse with the cleanest deploy path in its category. NOTE: matgl >=1.3 moved all pretrained models to the HF materialyze org — the old GitHub path 404s.
MatPES foundation potential; unique r2SCAN-fidelity variant available (everything else here is PBE); CPU-friendly size.
CORRECTION (verified 2026-07-06): there is NO microsoft/mattersim repo on Hugging Face (404) — weights come via pip/GitHub. 1M and 5M variants, MIT.
The current recommended MACE default (MPtrj+sAlex) — materially better than mace-mp-0; same MIT release channel. Deployable alongside our mace-mp-0 serving container lineage.
3.27M params (smallest serious UIP), OpenLAM 163M structures; LGPL fine for hosted use; DPA-4.0.1-Pro-MPtrj (#14, CC-BY-4.0) is the newer sibling.
One model, five task heads (materials/molecules/catalysis/MOFs/ODAC) via MoLE routing — the most versatile deployable potential; avoid archived uma-s-1 (known bug); use uma-s-1p2.pt.
Deploy the conservative-inf-mpa variant — the direct variants break MD energy conservation.
Best F1 on the board under plain MIT for BOTH code and weights — the top pick when license simplicity matters.
NON-COMMERCIAL (ASL) — research/evaluation tier only; commercial hosting requires negotiating with ICAMS. Graph-ACE, extremely fast in LAMMPS.
Canonical refs confirmed 2026-07-06: HF microsoft/mattergen (MIT, verified) with base + property-conditioned (chemistry, symmetry, magnetic density, bulk modulus) checkpoints; pip package exists.
Multi-task 'Omni' trained on 15 datasets (243M structures) — best cross-domain transfer of the compliant set; native LAMMPS parallel MD.
Only 10.4M params for near-top accuracy AND predicts magmoms — best small permissive model.
GPT-2-style LLM that writes CIFs from composition prompts; small model runs CPU-only — the cheapest generative SKU possible.
44.9M-param E(3)-equivariant attention-free transformer; near-top everything; cuEquivariance dependency makes it effectively GPU-only.
Foundation NequIP (32M); strict equivariance + ZBL core = robust MD; smaller OAM-L and LAMMPS-scalable Allegro-OAM-L from the same registry.
Smooth-energy network; checkpoint esen_30m_oam.pt; GATED download — the deploy pipeline needs a baked-in HF token; same repo holds the eqV2 OMat24 checkpoints.
730M-param unconstrained transformer (Ceriotti lab, Nat. Commun. 2025); biggest and heaviest UIP here — GPU-tier product. Checkpoint models/pet-oam-xl-v1.0.0.ckpt.
Current Matbench Discovery leader; Cartesian/spherical tensor ACE; permissive license — best accuracy-per-license in the catalog. Checkpoint file TECE-OAM-RRA-1.0.pt (TACE-OAM-L and RRA-Preview in the same repo).
~100 languages, fully open training recipe/data, MRL 768-to-256; good open-stack default.
The scientific-paper specialist: citation-informed paper-level embeddings for related-work/similar-paper linking in the KG; English only.
Unique triple-mode: dense + sparse + ColBERT multi-vector from one checkpoint — hybrid retrieval without extra models; 8k context fits full paper sections.
Decoupled two-stage VLM; top-tier formula CDM ~97 and table TEDS ~93; zh/en focus; ships with the full MinerU PDF pipeline. NOTE: base (non-Pro) MinerU2.5 is AGPL-3.0 — this Pro checkpoint is the Apache-2.0 one.
CogViT + GLM-0.5B decoder; SOTA key-information extraction; strong formulas; zh/en/fr/es/ru/de/ja/ko.
Robustness release (scan/skew/warp/photo distortions); seal recognition; 109 languages; most battle-tested of the top tier.
One-shot long-horizon parsing: 40+ PDF pages in a single pass (32k ctx); built on DeepSeek-OCR DeepEncoder + 3B MoE (~500M active); newest entrant.
Two-stage (PP-DocLayoutV3 + 0.9B ERNIE VLM); best-in-class tables (TEDS ~93) and formulas (CDM ~94+); 109 languages; cross-page table merge.
End-to-end single-model image-to-Markdown (no layout stage); strong tables/formulas; easiest deploy of the top tier (standard arch, GGUF ports exist).
Visual-causal-flow rearchitecture of optical context compression (~80% fewer vision tokens); best blank-page handling in hands-on tests; grounding support.
Single VLM does layout JSON + content in one pass; ~100 languages incl. low-resource — top pick for multilingual books.
The optical-context-compression original (10-20x token compression); superseded on accuracy but proven and very widely deployed (29M downloads).
End-to-end, no pipeline; strong arXiv-math and tables; European languages + zh/ja; EU lab with region:eu weights — a natural fit for EU tenants.
Qwen2.5-VL-7B + RL training; fully open (data+code+weights); equations/tables/handwriting; English-focused; best-documented batch pipeline.
Top-tier on messy real-world docs, handwriting, forms, 40+ languages. LICENSE GATE: free only below $2M revenue/funding — commercial use above needs a Datalab agreement.
NON-COMMERCIAL weights — research/evaluation tier only. Semantic tagging output (LaTeX equations, HTML tables, image descriptions, signatures/watermarks), handwriting, VQA.
NON-COMMERCIAL weights — research/evaluation tier only; commercial license requires contacting the HUST authors. Structure-Recognition-Relation paradigm; strong formulas.
Emits DocTags (lossless structured doc markup); equations, code, charts, tables; English (experimental zh/ja/ar); the cheap-tier/edge entry.
Battle-tested high-throughput pipeline, 90+ languages, optional LLM-boost mode. LICENSE GATE: weights free only below $2M revenue AND funding.
Most permissive full pipeline (MIT); PDF/DOCX/PPTX/HTML in; can swap in granite-docling VLM; the safe CPU-only enterprise entry.
NON-COMMERCIAL weights. Historical scientific-PDF-to-math-markdown pioneer; listed for legacy compatibility only.
EU provider; markdown + HTML tables + bbox + structured-annotation JSON. API-only: deploys as a metered pass-through (bring-your-own key or platform key), never as hosted weights.
The default premium pick: Apache-2.0, 119 languages, instruction-aware, MRL dims 32-4096.
Best consistency across 250+ languages; training data released; Llama 3.1 terms (fine for <700M-MAU commercial use).
The default small pick: multilingual + instruction-aware + long context at 0.6B; pairs with Qwen3-Reranker.
Compute: sign in to see available GPUs, your nodes, and running deployments.