Granite 3.3 models feature enhanced reasoning capabilities and support for Fill-in-the-Middle (FIM) code completion. They are built on a foundation of open-source instruction datasets with permissive licenses, alongside internally curated synthetic datasets tailored for long-context problem-solving. These models preserve the key strengths of previous Granite versions, including support for a 128K context length, strong performance in retrieval-augmented generation (RAG) and function calling, and controls for response length and originality. Granite 3.3 also delivers competitive results across general, enterprise, and safety benchmarks. Released as open source, the models are available under the Apache 2.0 license.
Benchmarks
评测成绩
| 评测基准 | 类别 | 分数 | 来源 |
|---|---|---|---|
| HumanEval | reasoning code | 89.7 | 来源 |
| AttaQ | safety | 88.5 | 来源 |
| HumanEval+ | reasoning | 86.1 | 来源 |
| AIME 2024 | math reasoning | 81.2 | 来源 |
| GSM8k | math reasoning | 80.9 | 来源 |
| IFEval | general | 74.8 | 来源 |
| BIG-Bench Hard | reasoning math language | 69.1 | 来源 |
| MATH-500 | math reasoning | 69.0 | 来源 |
| TruthfulQA | general reasoning legal healthcare finance | 66.9 | 来源 |
| MMLU | general reasoning language math | 65.5 | 来源 |
| AlpacaEval 2.0 | general creativity reasoning | 62.7 | 来源 |
| DROP | reasoning math | 59.4 | 来源 |
| Arena Hard | general reasoning creativity | 57.6 | 来源 |
| PopQA | general reasoning | 26.2 | 来源 |
Pricing
API 价格对比
暂无 API 价格。