Hermes 3 is a generalist language model with many improvements over Hermes 2, including advanced agentic capabilities, better roleplaying, reasoning, multi-turn conversation, and long context coherence. This 70B model is a competitive finetune of Llama-3.1-70B focused on aligning LLMs to the user with powerful steering capabilities.
Added Jan 7, 2026
Model weightsContext Window
65.5K
Max Output
8.2K
Avg output tokens (7d)
419 tokens
Input Price (Auto)
$0.43/1M
Output Price (Auto)
$0.43/1M
Cache Read (Auto)
$0.21/1M
Benchmarks
Benchmarks
Performance metrics and benchmarks
Sourced from Artificial Analysis.
Intelligence Index
5.1
Reasoning
GPQA Diamond
Graduate-level scientific reasoning
40.1%
Better than 17% of models compared
HLE
Humanity's Last Exam
4.1%
Better than 15% of models compared
IFBench
Instruction-following benchmark
29.0%
Better than 14% of models compared
T²-Bench Telecom
Conversational AI agents in dual-control scenarios
21.6%
Better than 23% of models compared
AA-LCR
Long context reasoning evaluation
3.0%
Better than 16% of models compared
CritPt
Research-level physics reasoning
0.0%
Coding
SciCode
Python programming for scientific computing
23.1%
Better than 28% of models compared
Terminal-Bench Hard
Agentic coding and terminal use
0.0%
Better than 6% of models compared
LiveCodeBench
Contamination-free coding benchmark
18.8%
Better than 20% of models compared
Math
AIME
American Invitational Mathematics Examination
2.3%
Better than 13% of models compared
Math-500
Diverse mathematical problem solving benchmark
53.8%
Better than 17% of models compared
Knowledge
MMLU-Pro
Professional and academic subject knowledge
57.1%
Better than 19% of models compared
AA-Omniscience Accuracy
Proportion of correctly answered questions
18.7%
AA-Omniscience Hallucination Rate
Rate of incorrect answers among non-correct responses
82.7%
Last updated Jun 28, 2026
Artificial AnalysisProviders
Choose explicit providers for this model. Auto routing remains available as the default option.
Loading provider options…
Related text models
Compare Hermes 3 70B with similar models from the same provider or model family.
Hermes 4 (Thinking)
NousResearch/Hermes-4-70B:thinkingHermes 4 70B with thinking enabled. Emits explicit reasoning content before final answer when streamed.
Hermes 4 Medium
nousresearch/hermes-4-70bEfficient reasoning model based on Llama-3.1-70B. Offers hybrid thinking capabilities with strong performance in math, code, and logical reasoning tasks. Supports structured outputs and JSON mode with enhanced steerability.
Hermes 4 Large
nousresearch/hermes-4-405bAdvanced reasoning model built on Llama-3.1-405B with hybrid thinking modes. Features internal deliberation capabilities, excels at math, code, STEM, and logical reasoning while supporting structured outputs with improved steerability and neutral alignment.
Hermes 4 Large (Thinking)
nousresearch/hermes-4-405b:thinkingHermes 4 Large with thinking enabled. Streams visible reasoning before the final answer and supports structured outputs.
EVA-LLaMA-3.33-70B-v0.1
EVA-UNIT-01/EVA-LLaMA-3.33-70B-v0.1A RP/storywriting specialist model, full-parameter finetune of Llama-3.3-70B-Instruct on mixture of synthetic and natural data. It uses Celeste 70B 0.1 data mixture, greatly expanding it to improve versatility, creativity and flavor of the resulting model.
Llama 3.3 70B Instruct abliterated
huihui-ai/Llama-3.3-70B-Instruct-abliteratedAn abliterated (removed restrictions and censorship) version of Llama 3.3 70b.