PDF理解功能使模型能夠解析並理解PDF文檔,提取文檔中的文字與圖片內容進行分析。您可以通過 OpenAI 相容的 Chat Completions 介面或 DashScope 介面,以 URL 或 Base 64 編碼方式傳入 PDF 檔案。
當前 PDF 理解功能僅支援華北2(北京)地區調用,且暫不支援通過 Responses API 呼叫。使用 Responses API傳入 PDF 時,請求會返回 HTTP 200,但檔案不會傳遞給模型。
支援的模型
qwen3.8-max、qwen3.8-flash、qwen3.8-27b
快速開始
運行以下代碼,向模型傳入PDF檔案。
OpenAI 相容
Python
範例程式碼
from openai import OpenAI
import os
client = OpenAI(
# 若沒有配置環境變數,請用百鍊API Key將下行替換為:api_key="sk-xxx"(不建議),
api_key=os.getenv("DASHSCOPE_API_KEY"),
# 以下為新加坡地區的URL,調用時請將 {WorkspaceId} 替換為真實的業務空間ID,各地區的URL不同。
base_url="https://{WorkspaceId}.ap-southeast-1.maas.aliyuncs.com/compatible-mode/v1",
)
completion = client.chat.completions.create(
model="qwen3.8-max",
messages=[
{
"role": "user",
"content": [
{
"type": "file",
"file": {
"file_url": "https://help-static-aliyun-doc.aliyuncs.com/file-manage-files/zh-CN/20260616/qmycjl/1506.02640v5.pdf"
}
},
{
"type": "text",
"text": "總結一下這個PDF文檔的內容"
}
]
}
],
stream=True,
stream_options={"include_usage": True}
)
for chunk in completion:
if not chunk.choices:
print(f"\nUsage: {chunk.usage}")
continue
delta = chunk.choices[0].delta
if hasattr(delta, "content") and delta.content:
print(delta.content, end="", flush=True)
Node.js
範例程式碼
import OpenAI from "openai";
import process from 'process';
const openai = new OpenAI({
// 若沒有配置環境變數,請用百鍊API Key將下行替換為:apiKey: "sk-xxx",
apiKey: process.env.DASHSCOPE_API_KEY,
// 以下為新加坡地區的URL,調用時請將 {WorkspaceId} 替換為真實的業務空間ID,各地區的URL不同。
baseURL: 'https://{WorkspaceId}.ap-southeast-1.maas.aliyuncs.com/compatible-mode/v1'
});
async function main() {
const stream = await openai.chat.completions.create({
model: 'qwen3.8-max',
messages: [
{
role: 'user',
content: [
{
type: 'file',
file: {
file_url: 'https://help-static-aliyun-doc.aliyuncs.com/file-manage-files/zh-CN/20260616/qmycjl/1506.02640v5.pdf'
}
},
{
type: 'text',
text: '總結一下這個PDF文檔的內容'
}
]
}
],
stream: true,
stream_options: { include_usage: true }
});
for await (const chunk of stream) {
if (!chunk.choices?.length) {
console.log('\nUsage:', chunk.usage);
continue;
}
const delta = chunk.choices[0].delta;
if (delta.content) {
process.stdout.write(delta.content);
}
}
}
main();
HTTP
範例程式碼
# 以下為新加坡地區的URL,調用時請將 {WorkspaceId} 替換為真實的業務空間ID,各地區的URL不同。
curl -X POST https://{WorkspaceId}.ap-southeast-1.maas.aliyuncs.com/compatible-mode/v1/chat/completions \
-H "Authorization: Bearer $DASHSCOPE_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "qwen3.8-max",
"messages": [
{
"role": "user",
"content": [
{
"type": "file",
"file": {
"file_url": "https://help-static-aliyun-doc.aliyuncs.com/file-manage-files/zh-CN/20260616/qmycjl/1506.02640v5.pdf"
}
},
{
"type": "text",
"text": "總結一下這個PDF文檔的內容"
}
]
}
],
"stream": true,
"stream_options": {
"include_usage": true
}
}'
DashScope
Python
範例程式碼
import os
from dashscope import MultiModalConversation
import dashscope
# 以下為新加坡地區的URL,調用時請將 {WorkspaceId} 替換為真實的業務空間ID,各地區的URL不同。
dashscope.base_http_api_url = "https://{WorkspaceId}.ap-southeast-1.maas.aliyuncs.com/api/v1"
messages = [
{
"role": "user",
"content": [
{
"file_url": "https://help-static-aliyun-doc.aliyuncs.com/file-manage-files/zh-CN/20260616/qmycjl/1506.02640v5.pdf"
},
{
"text": "總結一下這個PDF文檔的內容"
}
]
}
]
completion = MultiModalConversation.call(
# 若沒有配置環境變數,請用阿里雲百鍊API Key將下行替換為:api_key = "sk-xxx",
api_key=os.getenv("DASHSCOPE_API_KEY"),
model="qwen3.8-max",
messages=messages,
stream=True,
incremental_output=True
)
for chunk in completion:
message = chunk.output.choices[0].message
if message.content:
print(message.content[0]["text"], end="", flush=True)
HTTP
範例程式碼
# 以下為新加坡地區的URL,調用時請將 {WorkspaceId} 替換為真實的業務空間ID,各地區的URL不同。
curl -X POST "https://{WorkspaceId}.ap-southeast-1.maas.aliyuncs.com/api/v1/services/aigc/multimodal-generation/generation" \
-H "Authorization: Bearer $DASHSCOPE_API_KEY" \
-H "Content-Type: application/json" \
-H "X-DashScope-SSE: enable" \
-d '{
"model": "qwen3.8-max",
"input": {
"messages": [
{
"role": "user",
"content": [
{
"file_url": "https://help-static-aliyun-doc.aliyuncs.com/file-manage-files/zh-CN/20260616/qmycjl/1506.02640v5.pdf"
},
{
"text": "總結一下這個PDF文檔的內容"
}
]
}
]
},
"parameters": {
"incremental_output": true,
"result_format": "message"
}
}'
使用Base64輸入
如果無法提供檔案的URL地址,也可以將PDF檔案以Base64編碼字串的形式傳入。使用 file_data 時,filename 欄位為必填項。
說明Base 64 編碼會使資料體積增大約 1/3,例如 150MB 的檔案編碼後約 200MB,會超出請求體大小上限。大檔案請改用 URL 方式傳入。
import base64
from openai import OpenAI
import os
client = OpenAI(
api_key=os.getenv("DASHSCOPE_API_KEY"),
# 以下為新加坡地區的URL,調用時請將 {WorkspaceId} 替換為真實的業務空間ID,各地區的URL不同。
base_url="https://{WorkspaceId}.ap-southeast-1.maas.aliyuncs.com/compatible-mode/v1",
)
# 讀取並編碼PDF檔案
with open("report.pdf", "rb") as f:
pdf_base64 = base64.b64encode(f.read()).decode("utf-8")
completion = client.chat.completions.create(
model="qwen3.8-max",
messages=[
{
"role": "user",
"content": [
{
"type": "file",
"file": {
"file_data": f"data:application/pdf;base64,{pdf_base64}",
"filename": "report.pdf"
}
},
{
"type": "text",
"text": "這份報告的核心結論是什嗎?"
}
]
}
]
)
print(completion.choices[0].message.content)
import base64
import os
from dashscope import MultiModalConversation
import dashscope
# 以下為新加坡地區的URL,調用時請將 {WorkspaceId} 替換為真實的業務空間ID,各地區的URL不同。
dashscope.base_http_api_url = "https://{WorkspaceId}.ap-southeast-1.maas.aliyuncs.com/api/v1"
# 讀取並編碼PDF檔案
with open("report.pdf", "rb") as f:
pdf_base64 = base64.b64encode(f.read()).decode("utf-8")
messages = [
{
"role": "user",
"content": [
{
"file_data": f"data:application/pdf;base64,{pdf_base64}",
"filename": "report.pdf"
},
{
"text": "這份報告的核心結論是什嗎?"
}
]
}
]
response = MultiModalConversation.call(
api_key=os.getenv("DASHSCOPE_API_KEY"),
model="qwen3.8-max",
messages=messages
)
print(response.output.choices[0].message.content[0]["text"])
請求參數
檔案輸入通過content數組中的元素指定,OpenAI相容協議使用 type: "file" 類型,DashScope協議使用包含 file_url/file_data 的元素。
說明協議中 URL 部分僅支援字串輸入,不支援 list(數組)形式。
參數 | 類型 | 是否必填 | 說明 |
|---|---|---|---|
file_url | string | 二選一必填 | 指定PDF檔案的下載地址,和 |
file_data | string | Base64格式的PDF檔案輸入,格式為 | |
filename | string | 條件必填 | 檔案名稱,使用 |
file_format | string | 否 | 檔案格式,選擇性參數,當前僅支援 |
限制說明
專案 | 限制 |
|---|---|
單檔案大小限制 | 150MB |
單文檔頁數限制 | 256頁 |
說明PDF解析可能比普通文本請求耗時更長,首包逾時時間最長為300秒,建議使用流式輸出方式即時擷取結果,避免長時間等待。
計費說明
計費涉及以下方面:
- 模型調用費用:PDF檔案解析出的文字與圖片會計入模型的輸入Token,按照模型的標準輸入價格計費。
- 文檔解析費用:按PDF文檔解析的頁數計費(地區支援範圍見頁首說明),$0.00275/頁。
各模型的輸入輸出單價請參見模型調用計費。