文档 · 模型能力
视觉理解
带「视觉理解」能力的模型可以同时理解图片和文字:识别票据、读图表、看截图、描述照片等。
import base64
# Option 1: a publicly reachable image URL
url_part = {"type": "image_url", "image_url": {"url": "https://example.com/receipt.jpg"}}
# Option 2: a local file as a base64 data URL
b64 = base64.b64encode(open("receipt.jpg", "rb").read()).decode()
data_part = {"type": "image_url", "image_url": {"url": f"data:image/jpeg;base64,{b64}"}}
resp = client.chat.completions.create(
model="gpt-4.1-mini",
messages=[{
"role": "user",
"content": [
{"type": "text", "text": "What is the total amount on this receipt?"},
data_part,
],
}],
)
print(resp.choices[0].message.content)图片输入方式
- 公网 URL:图片必须能被公网访问,内网地址和需要登录的链接无法读取。
- base64 data URL:推荐用于本地文件或内网图片。
- 一条消息中可以包含多张图片,与文字混排。
Anthropic 格式
msg = client.messages.create(
model="claude-sonnet-4-5",
max_tokens=1024,
messages=[{
"role": "user",
"content": [
{"type": "image", "source": {"type": "base64", "media_type": "image/jpeg", "data": b64}},
# or {"type": "image", "source": {"type": "url", "url": "https://example.com/receipt.jpg"}}
{"type": "text", "text": "What is the total amount on this receipt?"},
],
}],
)计费与限制
- 图片会折算成输入 tokens(部分模型单独列出图片输入价格),分辨率越高消耗越多。
- 上传前把图片缩放到够用的分辨率,可以明显降低成本和延迟。
- 常见支持格式为 JPEG、PNG、WebP、GIF;单张图片过大会返回 400。