PDF の理解機能により、モデルは PDF ドキュメントを解析・理解し、テキストとイメージを抽出して分析できます。PDF ファイルは、URL または Base64 エンコーディングで渡すことができます。
PDF の理解機能は現在、中国 (北京) リージョンでのみ利用可能です。また、現時点では Responses API を介した呼び出しはサポートされていません。Responses API を使用して PDF を渡した場合、呼び出しは HTTP 200 を返しますが、ファイルはモデルに配信されません。
サポートされるモデル
qwen3.8-max, qwen3.8-flash, qwen3.8-27b
クイックスタート
以下の例は、PDF ファイルをモデルに送信する方法を示しています。
API キーの取得およびAPI キーを環境変数として設定を完了している必要があります。
OpenAI 互換
Python
サンプルコード
from openai import OpenAI
import os
client = OpenAI(
# 環境変数が設定されていない場合は、ご自身の Model Studio API キーに置き換えます: api_key="sk-xxx"
api_key=os.getenv("DASHSCOPE_API_KEY"),
# 以下はシンガポールリージョンの URL です。API を呼び出す際は、{WorkspaceId} を実際のワークスペース ID に置き換えてください。URL はリージョンによって異なります。
base_url="https://{WorkspaceId}.ap-southeast-1.maas.aliyuncs.com/compatible-mode/v1",
)
completion = client.chat.completions.create(
model="qwen3.8-max",
messages=[
{
"role": "user",
"content": [
{
"type": "file",
"file": {
"file_url": "https://help-static-aliyun-doc.aliyuncs.com/file-manage-files/zh-CN/20260616/qmycjl/1506.02640v5.pdf"
}
},
{
"type": "text",
"text": "Summarize this PDF document"
}
]
}
],
stream=True,
stream_options={"include_usage": True}
)
for chunk in completion:
if not chunk.choices:
print(f"\nUsage: {chunk.usage}")
continue
delta = chunk.choices[0].delta
if hasattr(delta, "content") and delta.content:
print(delta.content, end="", flush=True)
Node.js
サンプルコード
import OpenAI from "openai";
import process from 'process';
const openai = new OpenAI({
// 環境変数が設定されていない場合は、ご自身の Model Studio API キーに置き換えます: apiKey: "sk-xxx"
apiKey: process.env.DASHSCOPE_API_KEY,
// 以下はシンガポールリージョンの URL です。API を呼び出す際は、{WorkspaceId} を実際のワークスペース ID に置き換えてください。URL はリージョンによって異なります。
baseURL: 'https://{WorkspaceId}.ap-southeast-1.maas.aliyuncs.com/compatible-mode/v1'
});
async function main() {
const stream = await openai.chat.completions.create({
model: 'qwen3.8-max',
messages: [
{
role: 'user',
content: [
{
type: 'file',
file: {
file_url: 'https://help-static-aliyun-doc.aliyuncs.com/file-manage-files/zh-CN/20260616/qmycjl/1506.02640v5.pdf'
}
},
{
type: 'text',
text: 'Summarize this PDF document'
}
]
}
],
stream: true,
stream_options: { include_usage: true }
});
for await (const chunk of stream) {
if (!chunk.choices?.length) {
console.log('\nUsage:', chunk.usage);
continue;
}
const delta = chunk.choices[0].delta;
if (delta.content) {
process.stdout.write(delta.content);
}
}
}
main();
HTTP
サンプルコード
# 以下はシンガポールリージョンの URL です。API を呼び出す際は、{WorkspaceId} を実際のワークスペース ID に置き換えてください。URL はリージョンによって異なります。
curl -X POST https://{WorkspaceId}.ap-southeast-1.maas.aliyuncs.com/compatible-mode/v1/chat/completions \
-H "Authorization: Bearer $DASHSCOPE_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "qwen3.8-max",
"messages": [
{
"role": "user",
"content": [
{
"type": "file",
"file": {
"file_url": "https://help-static-aliyun-doc.aliyuncs.com/file-manage-files/zh-CN/20260616/qmycjl/1506.02640v5.pdf"
}
},
{
"type": "text",
"text": "Summarize this PDF document"
}
]
}
],
"stream": true,
"stream_options": {
"include_usage": true
}
}'
DashScope
Python
サンプルコード
import os
from dashscope import MultiModalConversation
import dashscope
# 以下はシンガポールリージョンの URL です。API を呼び出す際は、{WorkspaceId} を実際のワークスペース ID に置き換えてください。URL はリージョンによって異なります。
dashscope.base_http_api_url = "https://{WorkspaceId}.ap-southeast-1.maas.aliyuncs.com/api/v1"
messages = [
{
"role": "user",
"content": [
{
"file_url": "https://help-static-aliyun-doc.aliyuncs.com/file-manage-files/zh-CN/20260616/qmycjl/1506.02640v5.pdf"
},
{
"text": "Summarize this PDF document"
}
]
}
]
completion = MultiModalConversation.call(
# 環境変数が設定されていない場合は、ご自身の Alibaba Cloud Model Studio API キーに置き換えます: api_key="sk-xxx"
api_key=os.getenv("DASHSCOPE_API_KEY"),
model="qwen3.8-max",
messages=messages,
stream=True,
incremental_output=True
)
for chunk in completion:
message = chunk.output.choices[0].message
if message.content:
print(message.content[0]["text"], end="", flush=True)
HTTP
サンプルコード
# 以下はシンガポールリージョンの URL です。API を呼び出す際は、{WorkspaceId} を実際のワークスペース ID に置き換えてください。URL はリージョンによって異なります。
curl -X POST "https://{WorkspaceId}.ap-southeast-1.maas.aliyuncs.com/api/v1/services/aigc/multimodal-generation/generation" \
-H "Authorization: Bearer $DASHSCOPE_API_KEY" \
-H "Content-Type: application/json" \
-H "X-DashScope-SSE: enable" \
-d '{
"model": "qwen3.8-max",
"input": {
"messages": [
{
"role": "user",
"content": [
{
"file_url": "https://help-static-aliyun-doc.aliyuncs.com/file-manage-files/zh-CN/20260616/qmycjl/1506.02640v5.pdf"
},
{
"text": "Summarize this PDF document"
}
]
}
]
},
"parameters": {
"incremental_output": true,
"result_format": "message"
}
}'
Base64 入力の使用
URL を指定できない場合は、PDF ファイルを Base64 エンコードされた文字列として渡すことができます。file_data を使用する場合、filename フィールドは必須です。
注記Base64 エンコーディングにより、データサイズが約 3 分の 1 増加します。例えば、150 MB のファイルはエンコード後に約 200 MB となり、リクエストボディのサイズ制限を超えてしまいます。サイズの大きいファイルの場合は、代わりに URL メソッドを使用してください。
import base64
from openai import OpenAI
import os
client = OpenAI(
api_key=os.getenv("DASHSCOPE_API_KEY"),
# 以下はシンガポールリージョンの URL です。API を呼び出す際に、{WorkspaceId} を実際のワークスペース ID に置き換えてください。URL はリージョンによって異なります。
base_url="https://{WorkspaceId}.ap-southeast-1.maas.aliyuncs.com/compatible-mode/v1",
)
# PDF ファイルを読み取り、エンコードします
with open("report.pdf", "rb") as f:
pdf_base64 = base64.b64encode(f.read()).decode("utf-8")
completion = client.chat.completions.create(
model="qwen3.8-max",
messages=[
{
"role": "user",
"content": [
{
"type": "file",
"file": {
"file_data": f"data:application/pdf;base64,{pdf_base64}",
"filename": "report.pdf"
}
},
{
"type": "text",
"text": "このレポートの主要な調査結果は何ですか?"
}
]
}
]
)
print(completion.choices[0].message.content)
import base64
import os
from dashscope import MultiModalConversation
import dashscope
# 以下はシンガポールリージョンの URL です。API を呼び出す際に、{WorkspaceId} を実際のワークスペース ID に置き換えてください。URL はリージョンによって異なります。
dashscope.base_http_api_url = "https://{WorkspaceId}.ap-southeast-1.maas.aliyuncs.com/api/v1"
# PDF ファイルを読み取り、エンコードします
with open("report.pdf", "rb") as f:
pdf_base64 = base64.b64encode(f.read()).decode("utf-8")
messages = [
{
"role": "user",
"content": [
{
"file_data": f"data:application/pdf;base64,{pdf_base64}",
"filename": "report.pdf"
},
{
"text": "このレポートの主要な調査結果は何ですか?"
}
]
}
]
response = MultiModalConversation.call(
api_key=os.getenv("DASHSCOPE_API_KEY"),
model="qwen3.8-max",
messages=messages
)
print(response.output.choices[0].message.content[0]["text"])
リクエストパラメータ
ファイル入力は、type: "file" を持つコンテンツアイテム (OpenAI 互換) として、または file_url/file_data を含むコンテンツアイテム (DashScope) として指定します。
注記URL フィールドは文字列のみを受け付けます。URL のリスト (配列) はサポートされていません。
パラメータ | 型 | 必須 | 説明 |
|---|---|---|---|
file_url | 文字列 | 条件付き | ダウンロードする PDF ファイルの URL です。 |
file_data | 文字列 | 条件付き |
|
filename | 文字列 | 条件付き | ファイル名です。 |
file_format | 文字列 | いいえ | ファイル形式です。現在、 |
制限事項
項目 | 制限 |
|---|---|
最大ファイルサイズ | 150 MB |
最大ページ数 | 256 ページ |
注記PDF 解析は、通常のテキストリクエストよりも時間がかかる場合があります。最初のトークンのタイムアウトは最大 300 秒です。長時間の待機を避けるため、ストリーミング出力のご利用を推奨します。
課金
課金には、以下が含まれます:
- モデルの入力トークン:PDF から抽出されたテキストおよびページから解析されたイメージは、入力トークンとしてカウントされ、モデルの標準入力トークンレートで課金されます。
- ドキュメント解析手数料:解析された PDF ドキュメントのページごとに課金され、測定項目 document_parsing (pdf) に対応します。中国 (北京):1 ページあたり 0.00275 USD。
モデルごとの入出力トークン価格については、「モデル推論の料金」をご参照ください。