For safety evaluations, use Inspect for scripted, repeatable model safety tests and model‑graded rubrics. Host and publish Model Cards on Hugging Face Hub and automate card output with Google Model Card Toolkit or NVIDIA NeMo. Run automated scans with Garak or Promptfoo and hire Scale/iMerit for human red‑teaming when needed.
2Hugging FaceBest when you want to host open‑source models and publish standardized Model Cards with community feedback and templates; Hugging Face centralizes documentation, but it doesn't run in‑depth adversarial red‑teaming for you.43%
Preparing to release an open-source model requires a mix of automated tools, independent auditing groups, and standardized documentation frameworks. Depending on your model's scale, threat profile, and resources, assistance with safety evaluations and model cards generally comes from three primary pillars:
For frontier-level or high-capability open models, specialized independent research organizations conduct external technical safety and capability assessments:
METR (Model Evaluation and Threat Research) : A prominent research non-profit that scientifically measures whether and when AI systems exhibit dangerous or catastrophic autonomous capabilities. They evaluate models on long-horizon tasks, cyber-capabilities, and self-replication risks.
Redwood Research : Frequently engaged for technical alignment evaluations and evaluating control mechanisms.
Community & Academic Red-Teaming : Organizing private or public "trusted tester" periods with academic researchers, domain experts (e.g., biological or cybersecurity specialists), and ethical AI groups before the public weights drop.
If you are running your own pre-release safety evaluations and vulnerability scanning, the open-source community relies heavily on model-agnostic tooling:
garak : Widely used as the "LLM vulnerability scanner" to probe open-source models for prompt injection, toxic generation, data leakage, and structural weaknesses.
Promptfoo : Great for evaluating and red-teaming LLM outputs, security controls, and catching regressions in safety behavior.
PyRIT (Python Risk Identification Tool) : An open-source orchestration framework built by Microsoft to automate risk identification and red-teaming for AI systems across multiple endpoints.
3. Model Card Standards & Hosting Platforms
For drafting and publishing your model card (the standard documentation detailing training data, intended use, limitations, and evaluation metrics):
Hugging Face Model Cards : The industry standard ecosystem. When you host your weights on Hugging Face, their markdown-based model card templates guide you through required sections (intended uses, out-of-scope uses, bias, evaluation results, and environmental impact).
Standardized Frameworks : Refer to the original guidelines by Margaret Mitchell et al. or modern extensions like AI System Cards (which incorporate comprehensive data pedigree, security posture, and mitigation logs).
To help narrow down who or what you should engage next, could you tell me:
What is the approximate parameter size or capability tier of your model?
Does it have specialized capabilities (e.g., dual-use code generation, biology/chemistry knowledge, or general conversational text)?
I can suggest the most relevant evaluation frameworks or compliance targets for your specific launch.
If you’re preparing an open-source/open-weights model release, there are several useful organizations and tools, depending on whether you want independent safety testing, standardized benchmarks, or help documenting the release.
Safety evaluations
MLCommons — Its AI Risk & Reliability group develops standardized AI safety tests and benchmarks. AILuminate is particularly relevant for LLM safety, with testing around hazardous content and an official evaluation process.
Hugging Face — Useful for running and publishing conventional model evaluations and making results discoverable alongside the model. Its model-card format explicitly supports evaluation results.
Independent red-teamers / AI safety researchers — For a serious release, consider commissioning people who weren't involved in training to test misuse, jailbreak resistance, privacy/memorization, bias, cyber capabilities, and other risks specific to your model. This complements automated benchmarks; Google's published model-card methodology, for example, combines structured evaluations with internal red-teaming.
Model cards
You don't necessarily need an outside organization to write the card. A good division of labor is:
Model developers: architecture, training procedure, datasets, quantitative evaluations and technical limitations.
Project/release owner: intended uses, out-of-scope uses, licensing, contact information and release details.
That's essentially how the Hugging Face annotated model-card template divides the work.
For tooling, Google's Model Card Toolkit can generate structured model-card documents and schemas, while Hugging Face uses a README.md model card with structured evaluation metadata.
A practical release setup
For an open model, I'd aim for:
Standard capability evals — e.g. reasoning, knowledge, coding, instruction following, depending on the model.
Safety eval suite — harmful-content refusal, jailbreaks, privacy/memorization, bias/representational harms, cyber or biological capability where relevant.
Human red-team — ideally by people independent of the training team.
Reproducible evaluation methodology — prompts, versions, scoring procedures and relevant code where licensing permits.
Model card — intended use, limitations, training-data information, evaluation results, safety findings, known failure modes and contact/reporting mechanism.
Release an evaluation appendix — so the relatively concise model card doesn't have to contain every technical detail.
One useful current development is Evaluation Cards, an open-source project from the EvalEval Coalition intended to improve the reproducibility, provenance and comparability of AI evaluation reporting.
If you tell me what kind of model you're releasing (LLM, vision, multimodal, code, etc.), approximate parameter size, and whether you want an independent third-party evaluation, I can give you a concrete shortlist of evaluators and a release/evaluation checklist.
Preparing to release an open-source model involves navigating a growing ecosystem of specialized organizations, benchmarking groups, and documentation frameworks. Depending on whether you need deep external technical auditing, standardized safety benchmarks, or automated documentation help, different entities assist with safety evaluations and model cards.
1. Safety Evaluation & Red-Teaming Organizations
If you want independent third-party verification or rigorous vulnerability and risk assessments (such as for CBRN, cybersecurity, or autonomous capabilities), several prominent groups operate in this space:
METR (Model Evaluation and Threat Research) : Specializes in evaluating autonomous capabilities and whether models can act as independent agents or replicate/self-improve.
Redwood Research : Focuses on empirical safety, alignment research, and adversarial testing/red-teaming.
MLCommons (AI Safety Working Group) : Develops industry-standard safety benchmarks and test specifications (collaborating with frameworks like Stanford's HELM) to systematically test models for hazards like toxicity, bias, and dangerous content.
National AI Safety Institutes (AISIs) : Depending on your jurisdiction and scale (such as the U.S. AI Safety Institute or its international counterparts), government-backed institutes increasingly engage with developers for pre-release frontier evaluations.
Commercial Evaluation Startups : Platforms like Patronus AI or Braintrust provide automated evaluation suites for LLM risk testing, hallucination detection, and safety guardrails.
2. Model Cards and Documentation Assistance
Model cards (pioneered originally by Google researchers) are typically authored internally by the model development team, but tooling and standardized templates have made the process modular and semi-automated:
Hugging Face Model Cards : If you are hosting your open-source model on Hugging Face, they provide extensive metadata standards, community templates, and interactive widgets to structure your model card detailing intended use, limitations, and evaluation metrics.
NVIDIA Model Card Generator (MCG) Toolkit : An enterprise/open tool that automates the generation of structured Model Card++ documentation (covering bias, explainability, privacy, and safety) by parsing model pipelines and generating CycloneDX-compliant outputs.
Standard Framework Templates : Most open-source releases adapt templates provided by MLCommons or the original Model Card frameworks to report quantitative evaluation results across standard slices (demographics, languages, and adversarial prompts).
If you'd like to narrow this down, tell me:
What scale and modality is your model (e.g., text-only LLM, multimodal, under 10B parameters or frontier-scale)?
Are you looking for open-source automated tools to run locally, or third-party audit partners to contact before launch?
If you’re preparing an open-source/open-weight model release, there are several organizations and tools worth considering. They serve somewhat different roles:
METR — independent AI safety evaluation, particularly for dangerous capabilities, autonomy, and frontier-model risks. METR is one of the prominent organizations doing external model evaluations.
MLCommons / AI Risk & Reliability — develops standardized safety tests and benchmarks. Its AILuminate benchmark covers hazardous behavior across multiple categories and has both open practice testing and a formal official test.
EleutherAI — an open research organization with substantial experience in open models and evaluations. Its evaluation ecosystem includes the LM Evaluation Harness, widely used for standardized model benchmarking.
Hugging Face / EvalEval Coalition — particularly relevant to documentation and reproducible evaluation reporting. EvalEval's Evaluation Cards project is designed to standardize how evaluation results, provenance, completeness, and comparability are documented.
Inspect AI — an open-source evaluation framework from the UK AI Safety Institute ecosystem for testing model capabilities, safety, and alignment.
HarmBench / Center for AI Safety — useful if you need structured testing specifically around harmful behavior and jailbreak resistance.
For the model card itself
I'd separate model documentation from independent safety evaluation. Your model card should ideally document:
model architecture, size, training data/source categories, and training procedure
intended and out-of-scope uses
known limitations and failure modes
safety mitigations and their limitations
evaluation methodology, datasets, prompts, and versions
quantitative results, including failures—not just aggregate scores
red-team methodology and findings
known security/privacy issues
licensing and acceptable-use restrictions, if any
provenance of external evaluations and whether evaluators had access to weights, system prompts, training information, etc.
For an open release, reproducibility and provenance are particularly valuable: EvalEval's current work is explicitly aimed at making evaluation results more complete and comparable.
A practical release setup
I'd consider a three-layer process:
Internal evals: LM Evaluation Harness / Inspect AI + your own task-specific safety tests.
Independent safety testing: an external group such as METR, or an appropriate specialist evaluator depending on your model's capabilities and risk profile.
Public reporting: a conventional model card plus an evaluation/evidence section documenting exactly what was tested, by whom, under what access conditions, and what wasn't tested.
If you tell me roughly what kind of model you're releasing (e.g. 7B/70B LLM, multimodal, coding, agentic, biology, etc.) and whether you want free/open-source help or are willing to pay for an external evaluation, I can give you a much more targeted shortlist and a suggested pre-release safety-evaluation plan.
Preparing to release an open-source model requires a mix of technical benchmarking, vulnerability "red-teaming," and transparent documentation. Depending on your resources and whether you are an independent developer, an academic lab, or a commercial enterprise, several distinct groups, tools, and organizations specialize in helping with safety evaluations and drafting model/system cards.
1. External AI Safety & Red-Teaming Organizations
If you need external validation, stress-testing, or adversarial testing (red-teaming) for safety vulnerabilities, biases, or dangerous capabilities:
Government AI Safety Institutes: The US AI Safety Institute (NIST AISI) and the UK AI Safety Institute evaluate frontier systems and develop rigorous testing standards for trustworthy AI.
Specialized Research Labs & Non-Profits: Organizations like the Alignment Research Center , Redwood Research, and academic centers (such as Stanford's HAI or MIT) often consult on alignment, robustness, and risk assessments.
Community Red-Teaming Networks: Modeled after programs like the OpenAI Red Teaming Network , open-source projects frequently crowdsource external security researchers, ethicists, and domain experts via platforms like EleutherAI or specialized Discord communities to find jailbreaks and failure modes before public weights drop.
Rather than relying purely on manual audits, most open-source releases use standardized evaluation frameworks to generate quantitative metrics for their model cards:
Promptfoo: A popular open-source tool used to stress-test LLMs for toxicity, PII leaks, and jailbreaks.
Inspect & lm-evaluation-harness: Frameworks by organizations like the UK AISI (Inspect) and EleutherAI (lm-evaluation-harness ) that automate running standardized capability and safety benchmarks.
Evaluation Ecosystems: Initiatives like the EvalEval Coalition provide infrastructure and data standards to improve the completeness and comparability of model evaluations.
3. Writing Model Cards (Standard Practice)
Unlike closed-source models where a centralized corporate safety team drafts a proprietary system card, open-source model cards are typically authored by the core engineering/research team using established open-source templates.
Hugging Face Model Card Template: The industry standard. When hosting your model on Hugging Face , they provide built-in metadata tags and markdown templates prompting you to fill in intended uses, out-of-scope uses, training data provenance, limitations, and evaluation results.
The Original Framework: Grounded in the original research framework proposed by Margaret Mitchell et al. (2018), a proper model card should systematically document performance across diverse demographic or operational slices rather than just aggregate metrics.
To help narrow down who or what you should engage, tell me:
What is the approximate scale/capability of your model (e.g., small fine-tune vs. a large frontier base model)?
Are you looking for automated open-source software tools to run the tests yourself, or third-party human auditors/red-teamers?
Yes. For an open-source model release, I’d think about two separate jobs: (1) independent or structured safety evaluation, and (2) a transparent model card/release package.
Safety evaluation
MLCommons — particularly its AI Risk & Reliability working group. It develops open safety tests and benchmarks, including AILuminate, which evaluates general-purpose chat models across 12 hazard categories. It has an open practice-testing path as well as an official benchmark process.
Hugging Face — useful for running and publishing evaluations alongside the model. Its Evaluate ecosystem supports evaluation libraries, benchmark leaderboards, and structured evaluation results.
Independent red-teamers / safety researchers — particularly valuable if you want something beyond benchmark scores. A good release can include adversarial testing for jailbreaks, harmful-content generation, privacy leakage, bias, misuse, and model-specific failure modes. Hugging Face's model-card guidance explicitly recommends combining quantitative performance evaluation with evaluation of social/safety issues such as red-teaming.
Model cards
Hugging Face is probably the most straightforward place to start. Its model cards are Markdown README.md files attached to the model repository and are intended to document:
intended uses and limitations
biases and ethical considerations
training procedure and technical details
training datasets
evaluation methodology and results
relevant risks and mitigations
Hugging FaceTheir annotated template is particularly useful because it recommends input from different perspectives: the developer, a sociotechnical reviewer, and someone responsible for the project's scope and communication. It also recommends separating general-performance testing from safety/social-impact evaluation.
For a serious open-source release, I'd aim for something like:
Technical evaluation — capability benchmarks relevant to the model.
Safety evaluation — standardized tests such as AILuminate where applicable.
Adversarial red-team evaluation — deliberately try to break the model's safeguards.
Privacy/security testing — memorization, data leakage, prompt injection, etc., as relevant to your model.
Model card — document what you tested, methodology, datasets, results, known limitations, and intended/non-intended uses.
Reproducible eval artifacts — publish test configurations, versions, prompts where licensing permits, and scripts/results so others can reproduce or challenge the findings.
One useful distinction: a benchmark result isn't the same thing as a safety certification. MLCommons itself notes that AILuminate is intended to provide information for decision-making and does not guarantee that a system is safe or reliable.
If you tell me what kind of model you're releasing (LLM, multimodal, image, code, agent, etc.) and roughly its size/capabilities, I can give you a concrete pre-release safety/evaluation checklist and identify organizations that could potentially perform an independent evaluation.
When preparing to release an open-source (or open-weight) model, navigating safety evaluations and drafting thorough model cards typically involves a mix of specialized third-party evaluation labs, open-source tooling, and community-driven red-teaming.
For formal risk assessments (such as checking for autonomous capabilities, cybersecurity risks, or CBRN vectors), developers often coordinate with external evaluation organizations:
Government AI Safety Institutes (ASIs) : Agencies like the UK AI Safety Institute (UK AISI) or the U.S. equivalent routinely collaborate with developers to run pre-deployment evaluations.
Commercial Red-Teaming Vendors : Specialized third-party security firms (such as Irregular or specialized AI security startups) perform stress tests, jailbreak testing, and adversarial robustness scans.
2. Standard Evaluation Frameworks & Software
If you are running the evaluations in-house or with an open-source community cohort, standard benchmarking platforms are widely utilized:
Inspect : Developed by the UK AISI, Inspect is an open-source evaluation platform and software library used to test core knowledge, reasoning, and safety parameters.
MLCommons : Participates in defining industry-standard safety and performance benchmarks through its safety working groups.
Hugging Face Evaluation Ecosystem : Utilizing tools like lighteval or community-centric evaluation harnesses to systematically grade model outputs across safety axes.
3. Crafting Model Cards
Model cards and system cards are traditionally authored by your own core research and safety teams, but they pull data from standardized templates and external audits:
Hugging Face Model Card Templates : The community standard for documenting intended use, training parameters, evaluation results, and ethical considerations. (Hugging Face provides built-in markdown templates directly in their repository ecosystem).
Llama Guard / Safety Classifiers : If you are releasing an open-weight model akin to Meta’s Llama series, documenting companion guardrails (like input/output content filters) is a critical addition to the model card to guide safe enterprise deployment.
Are you looking for independent evaluation partners for high-risk capabilities, or do you need help structuring the documentation and benchmarks for your specific model architecture?
If you’re releasing an open-source/open-weight model, there are several useful organizations and tools, depending on whether you need independent safety testing, standardized benchmarks, or help documenting the release.
Hugging Face — particularly useful for the model-card side. Its Hub model cards cover intended uses, limitations, biases/ethical considerations, training information, datasets, and evaluation results. It also provides evaluation tooling such as LightEval and Evaluate.
MLCommons — useful for standardized safety evaluation. Its AI Risk & Reliability working group develops safety tests and benchmarks, and its AILuminate suite evaluates hazardous-content behavior using standardized prompts and safety evaluators.
Google's Model Cards framework — the original model-card research is a good conceptual template for deciding what to disclose, including performance across relevant subgroups and intended-use conditions.
A practical release setup
For a serious open model release, I'd separate the work into:
Capability evaluation: standard benchmarks appropriate to your model and intended uses.
Safety evaluation: harmful-content, jailbreak, privacy, cyber, bio/chemical, autonomy/agentic, and other risks relevant to the model's capabilities.
Red teaming: ideally involving people outside the immediate model-development team.
Model card: document training data at an appropriate level, intended/foreseeable uses, limitations, evaluation methodology/results, known failure modes, safety mitigations, and release restrictions.
Reproducibility: publish enough information about evaluation prompts, versions, configurations, and methodology that others can understand or reproduce the results.
One important distinction: a model card is documentation, not an independent safety certification. Hugging Face explicitly notes that model-card evaluation scores are often produced by the model author, whereas community/benchmark evaluations are a separate matter.
If you tell me what kind of model you're releasing (LLM, multimodal, coding, agentic, etc.) and roughly its size/capabilities, I can give you a concrete pre-release safety/evaluation checklist and identify organizations that could potentially perform independent evaluations.
Third-party AI safety institutes, specialized auditing firms, and open-source evaluation toolkits help developers conduct safety evaluations and draft model cards prior to public release.
When preparing an open-source model, you can leverage several categories of external organizations, frameworks, and technical tools to handle rigorous safety checks and transparent documentation.
External Evaluation and Safety Organizations
National AI Safety Institutes (AISIs): Government-backed bodies like the UK AI Safety Institute and the US AISI offer frameworks, testing environments, and scientific guidance for frontier model evaluation.
Third-Party AI Auditors: Specialized security and ethics-focused firms (such as MLCommons, Scale AI, or Arthur AI) conduct independent red-teaming, bias testing, and alignment checks.
Domain-Specific Trusted Entities: Organizations like the Internet Watch Foundation (IWF) or NCMEC collaborate on specialized safety filters and benchmark evaluations for harmful content prevention.
Open-Source Evaluation Frameworks & Toolkits
Inspect: The UK AI Safety Institute's Inspect platform is an open-source testing framework designed specifically to evaluate model behavior, prompt safety, and agentic tool use.
Model Card Toolkit: Google provides the Model Card Toolkit to help automate, structure, and generate standardized model card documentation and metadata.
LM Evaluation Harness: Developed by EleutherAI, this widely accepted open-source framework standardizes zero-shot and few-shot benchmark testing across hundreds of tasks.
If you can share the scale (parameter size) and primary capabilities/modalities of your model, I can help point you toward the most relevant evaluation benchmarks and red-teaming frameworks for your release.
If you’re releasing an open-source/open-weight model, I’d separate safety evaluation from model-card documentation. A few organizations/projects are particularly useful:
mlcommons.org — probably the most directly relevant community for independent-style safety benchmarking. Its AILuminate benchmark evaluates general-purpose chat models across 12 hazard categories, and the working group welcomes contributors.
github.com — open tooling for actually running safety benchmarks against your model and producing reports. You can test a self-hosted model rather than relying exclusively on an API.
huggingface.co — the natural place to publish the model card if you're distributing through Hugging Face. Their format covers intended use, limitations, biases/risks, training information, datasets, and evaluation results.
Hugging Face's annotated model-card template is particularly useful for figuring out who should contribute what: technical evaluation, testing data/metrics, and a sociotechnical assessment of risks and limitations. It explicitly recommends including both quantitative evaluation and safety/social-impact evaluation such as red-teaming.
EvalEval / Evaluation Cards — a newer open-source effort aimed at making AI evaluation results more reproducible, comparable, and interpretable. It can be useful if you want your evaluation methodology/results documented in a standardized way rather than just listing benchmark scores.
A practical release setup
For a serious release, I'd aim for:
Capability evals — accuracy, reasoning, coding, multilingual performance, etc.
Safety evals — harmful-content generation, jailbreak resistance, privacy, bias, cyber/chemical/biological-risk categories as relevant to the model.
Adversarial red-teaming — humans trying to elicit failures that automated benchmarks miss.
Independent review — ideally have someone outside the core training team reproduce or review the safety results.
Model card — document intended uses, out-of-scope uses, training data, known limitations, evaluation methodology/results, safety mitigations, and residual risks.
Release artifacts — publish the evaluation code/configuration and enough information for others to reproduce the tests where doing so doesn't expose sensitive test material.
One important nuance: AILuminate isn't a blanket declaration that a model is "safe." Its own documentation says the benchmark provides relative safety information rather than guaranteeing safety or reliability.
If you tell me what kind of model you're releasing (LLM, multimodal, coding, agent, etc.) and its approximate size, I can give you a concrete pre-release safety-evaluation checklist + model-card template, including which tests I'd run and which organizations/tools I'd approach.