Model family by Anthropic · used by 2 applications
Claude Sonnet 4 is a hybrid reasoning large language model developed by Anthropic, released in May 2025. It is optimized for speed, cost-efficiency, and agentic tasks, such as coding, computer use, and tool-based workflows. The model supports extended thinking mode for complex tasks and is designed for scalable, high-volume applications. Claude Sonnet 4 was trained on a proprietary mix of publicly available internet data (up to March 2025), private datasets, opted-in user data, and synthetic data. It underwent RLHF (Reinforcement Learning from Human Feedback) and Constitutional AI fine-tuning to align with principles like helpfulness, honesty, and harmlessness. The model is text-only and multilingual. Safety evaluations show strong alignment, with a 98.99% harmless response rate on violative requests and low over-refusal rates (0.23%) on benign prompts. It was deployed under AI Safety Level 2 (ASL-2) protections.
Six dimensions, evidence-linked
Work at Anthropic? Claim this listing to correct or complete the data.