.. DO NOT EDIT. .. THIS FILE WAS AUTOMATICALLY GENERATED BY SPHINX-GALLERY. .. TO MAKE CHANGES, EDIT THE SOURCE PYTHON FILE: .. "tutorial/task_model.py" .. LINE NUMBERS ARE GIVEN BELOW. .. only:: html .. note:: :class: sphx-glr-download-link-note :ref:`Go to the end ` to download the full example code. .. rst-class:: sphx-glr-example-title .. _sphx_glr_tutorial_task_model.py: .. _model: Model ==================== In this tutorial, we introduce the model APIs integrated in AgentScope, how to use them and how to integrate new model APIs. The supported model APIs and providers include: .. list-table:: :header-rows: 1 * - API - Class - Compatible - Streaming - Tools - Vision - Reasoning * - OpenAI - ``OpenAIChatModel`` - vLLM, DeepSeek - ✅ - ✅ - ✅ - ✅ * - DashScope - ``DashScopeChatModel`` - - ✅ - ✅ - ✅ - ✅ * - Anthropic - ``AnthropicChatModel`` - - ✅ - ✅ - ✅ - ✅ * - Gemini - ``GeminiChatModel`` - - ✅ - ✅ - ✅ - ✅ * - Ollama - ``OllamaChatModel`` - - ✅ - ✅ - ✅ - ✅ .. note:: When using vLLM, you need to configure the appropriate tool calling parameters for different models during deployment, such as ``--enable-auto-tool-choice``, ``--tool-call-parser``, etc. For more details, refer to the `official vLLM documentation `_. .. note:: For OpenAI-compatible models (e.g. vLLM, Deepseek), developers can use the ``OpenAIChatModel`` class, and specify the API endpoint by the ``client_kwargs`` parameter: ``client_kwargs={"base_url": "http://your-api-endpoint"}``. For example: .. code-block:: python OpenAIChatModel(client_kwargs={"base_url": "http://localhost:8000/v1"}) .. note:: Model behavior parameters (such as temperature, maximum length, etc.) can be preset in the constructor function via the ``generate_kwargs`` parameter. For example: .. code-block:: python OpenAIChatModel(generate_kwargs={"temperature": 0.3, "max_tokens": 1000}) To provide unified model interfaces, the above model classes has the following common methods: - The first three arguments of the ``__call__`` method are ``messages`` , ``tools`` and ``tool_choice``, representing the input messages, JSON schema of tool functions, and tool selection mode, respectively. - The return type are either a ``ChatResponse`` instance or an async generator of ``ChatResponse`` in streaming mode. .. note:: Different model APIs differ in the input message format, refer to :ref:`prompt` for more details. The ``ChatResponse`` instance contains the generated thinking/text/tool use content, identity, created time and usage information. .. GENERATED FROM PYTHON SOURCE LINES 80-105 .. code-block:: Python import asyncio import json import os from agentscope.message import TextBlock, ToolUseBlock, ThinkingBlock, Msg from agentscope.model import ChatResponse, DashScopeChatModel response = ChatResponse( content=[ ThinkingBlock( type="thinking", thinking="I should search for AgentScope on Google.", ), TextBlock(type="text", text="I'll search for AgentScope on Google."), ToolUseBlock( type="tool_use", id="642n298gjna", name="google_search", input={"query": "AgentScope?"}, ), ], ) print(response) .. rst-class:: sphx-glr-script-out .. code-block:: none ChatResponse(content=[{'type': 'thinking', 'thinking': 'I should search for AgentScope on Google.'}, {'type': 'text', 'text': "I'll search for AgentScope on Google."}, {'type': 'tool_use', 'id': '642n298gjna', 'name': 'google_search', 'input': {'query': 'AgentScope?'}}], id='2026-05-25 08:41:44.019_024cae', created_at='2026-05-25 08:41:44.019', type='chat', usage=None, metadata=None) .. GENERATED FROM PYTHON SOURCE LINES 106-107 Taking ``DashScopeChatModel`` as an example, we can use it to create a chat model instance and call it with messages and tools: .. GENERATED FROM PYTHON SOURCE LINES 107-132 .. code-block:: Python async def example_model_call() -> None: """An example of using the DashScopeChatModel.""" model = DashScopeChatModel( model_name="qwen-max", api_key=os.environ["DASHSCOPE_API_KEY"], stream=False, ) res = await model( messages=[ {"role": "user", "content": "Hi!"}, ], ) # You can directly create a ``Msg`` object with the response content msg_res = Msg("Friday", res.content, "assistant") print("The response:", res) print("The response as Msg:", msg_res) asyncio.run(example_model_call()) .. rst-class:: sphx-glr-script-out .. code-block:: none The response: ChatResponse(content=[{'type': 'text', 'text': 'Hello! How can I assist you today?'}], id='f42a49fa-dacf-98be-83fb-5c1d6be69c18', created_at='2026-05-25 08:41:45.449', type='chat', usage=ChatUsage(input_tokens=10, output_tokens=9, time=1.429052, type='chat', metadata=GenerationUsage(input_tokens=10, output_tokens=9)), metadata=None) The response as Msg: Msg(id='GvmSwwf9wGz8K3ReR3fuvz', name='Friday', content=[{'type': 'text', 'text': 'Hello! How can I assist you today?'}], role='assistant', metadata={}, timestamp='2026-05-25 08:41:45.449', invocation_id='None') .. GENERATED FROM PYTHON SOURCE LINES 133-140 Streaming ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ To enable streaming model, set the ``stream`` parameter in the model constructor to ``True``. When streaming is enabled, the ``__call__`` method will return an **async generator** that yields ``ChatResponse`` instances as they are generated by the model. .. note:: The streaming mode in AgentScope is designed to be **cumulative**, meaning the content in each chunk contains all the previous content plus the newly generated content. .. GENERATED FROM PYTHON SOURCE LINES 140-170 .. code-block:: Python async def example_streaming() -> None: """An example of using the streaming model.""" model = DashScopeChatModel( model_name="qwen-max", api_key=os.environ["DASHSCOPE_API_KEY"], stream=True, ) generator = await model( messages=[ { "role": "user", "content": "Count from 1 to 20, and just report the number without any other information.", }, ], ) print("The type of the response:", type(generator)) i = 0 async for chunk in generator: print(f"Chunk {i}") print(f"\ttype: {type(chunk.content)}") print(f"\t{chunk}\n") i += 1 asyncio.run(example_streaming()) .. rst-class:: sphx-glr-script-out .. code-block:: none The type of the response: Chunk 0 type: ChatResponse(content=[{'type': 'text', 'text': '1'}], id='0ba2980e-4354-9f06-918e-5210621495db', created_at='2026-05-25 08:41:46.446', type='chat', usage=ChatUsage(input_tokens=27, output_tokens=1, time=0.995459, type='chat', metadata=GenerationUsage(input_tokens=27, output_tokens=1)), metadata=None) Chunk 1 type: ChatResponse(content=[{'type': 'text', 'text': '1\n2\n'}], id='0ba2980e-4354-9f06-918e-5210621495db', created_at='2026-05-25 08:41:46.515', type='chat', usage=ChatUsage(input_tokens=27, output_tokens=4, time=1.064213, type='chat', metadata=GenerationUsage(input_tokens=27, output_tokens=4)), metadata=None) Chunk 2 type: ChatResponse(content=[{'type': 'text', 'text': '1\n2\n3\n4'}], id='0ba2980e-4354-9f06-918e-5210621495db', created_at='2026-05-25 08:41:46.586', type='chat', usage=ChatUsage(input_tokens=27, output_tokens=7, time=1.13498, type='chat', metadata=GenerationUsage(input_tokens=27, output_tokens=7)), metadata=None) Chunk 3 type: ChatResponse(content=[{'type': 'text', 'text': '1\n2\n3\n4\n5\n'}], id='0ba2980e-4354-9f06-918e-5210621495db', created_at='2026-05-25 08:41:47.436', type='chat', usage=ChatUsage(input_tokens=27, output_tokens=10, time=1.985567, type='chat', metadata=GenerationUsage(input_tokens=27, output_tokens=10)), metadata=None) Chunk 4 type: ChatResponse(content=[{'type': 'text', 'text': '1\n2\n3\n4\n5\n6\n7\n8\n'}], id='0ba2980e-4354-9f06-918e-5210621495db', created_at='2026-05-25 08:41:47.721', type='chat', usage=ChatUsage(input_tokens=27, output_tokens=16, time=2.270519, type='chat', metadata=GenerationUsage(input_tokens=27, output_tokens=16)), metadata=None) Chunk 5 type: ChatResponse(content=[{'type': 'text', 'text': '1\n2\n3\n4\n5\n6\n7\n8\n9\n10\n1'}], id='0ba2980e-4354-9f06-918e-5210621495db', created_at='2026-05-25 08:41:47.845', type='chat', usage=ChatUsage(input_tokens=27, output_tokens=22, time=2.394715, type='chat', metadata=GenerationUsage(input_tokens=27, output_tokens=22)), metadata=None) Chunk 6 type: ChatResponse(content=[{'type': 'text', 'text': '1\n2\n3\n4\n5\n6\n7\n8\n9\n10\n11\n12\n1'}], id='0ba2980e-4354-9f06-918e-5210621495db', created_at='2026-05-25 08:41:48.313', type='chat', usage=ChatUsage(input_tokens=27, output_tokens=28, time=2.862106, type='chat', metadata=GenerationUsage(input_tokens=27, output_tokens=28)), metadata=None) Chunk 7 type: ChatResponse(content=[{'type': 'text', 'text': '1\n2\n3\n4\n5\n6\n7\n8\n9\n10\n11\n12\n13\n14\n1'}], id='0ba2980e-4354-9f06-918e-5210621495db', created_at='2026-05-25 08:41:48.399', type='chat', usage=ChatUsage(input_tokens=27, output_tokens=34, time=2.948456, type='chat', metadata=GenerationUsage(input_tokens=27, output_tokens=34)), metadata=None) Chunk 8 type: ChatResponse(content=[{'type': 'text', 'text': '1\n2\n3\n4\n5\n6\n7\n8\n9\n10\n11\n12\n13\n14\n15\n16\n1'}], id='0ba2980e-4354-9f06-918e-5210621495db', created_at='2026-05-25 08:41:49.527', type='chat', usage=ChatUsage(input_tokens=27, output_tokens=40, time=4.076106, type='chat', metadata=GenerationUsage(input_tokens=27, output_tokens=40)), metadata=None) Chunk 9 type: ChatResponse(content=[{'type': 'text', 'text': '1\n2\n3\n4\n5\n6\n7\n8\n9\n10\n11\n12\n13\n14\n15\n16\n17\n18\n1'}], id='0ba2980e-4354-9f06-918e-5210621495db', created_at='2026-05-25 08:41:49.681', type='chat', usage=ChatUsage(input_tokens=27, output_tokens=46, time=4.230569, type='chat', metadata=GenerationUsage(input_tokens=27, output_tokens=46)), metadata=None) Chunk 10 type: ChatResponse(content=[{'type': 'text', 'text': '1\n2\n3\n4\n5\n6\n7\n8\n9\n10\n11\n12\n13\n14\n15\n16\n17\n18\n19\n20'}], id='0ba2980e-4354-9f06-918e-5210621495db', created_at='2026-05-25 08:41:49.850', type='chat', usage=ChatUsage(input_tokens=27, output_tokens=50, time=4.399119, type='chat', metadata=GenerationUsage(input_tokens=27, output_tokens=50)), metadata=None) Chunk 11 type: ChatResponse(content=[{'type': 'text', 'text': '1\n2\n3\n4\n5\n6\n7\n8\n9\n10\n11\n12\n13\n14\n15\n16\n17\n18\n19\n20'}], id='0ba2980e-4354-9f06-918e-5210621495db', created_at='2026-05-25 08:41:49.872', type='chat', usage=ChatUsage(input_tokens=27, output_tokens=50, time=4.420958, type='chat', metadata=GenerationUsage(input_tokens=27, output_tokens=50)), metadata=None) .. GENERATED FROM PYTHON SOURCE LINES 171-175 Reasoning ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ AgentScope supports reasoning models by providing the ``ThinkingBlock``. .. GENERATED FROM PYTHON SOURCE LINES 175-200 .. code-block:: Python async def example_reasoning() -> None: """An example of using the reasoning model.""" model = DashScopeChatModel( model_name="qwen-turbo", api_key=os.environ["DASHSCOPE_API_KEY"], enable_thinking=True, ) res = await model( messages=[ {"role": "user", "content": "Who am I?"}, ], ) last_chunk = None async for chunk in res: last_chunk = chunk print("The final response:") print(last_chunk) asyncio.run(example_reasoning()) .. rst-class:: sphx-glr-script-out .. code-block:: none The final response: ChatResponse(content=[{'type': 'thinking', 'thinking': 'Okay, the user asked, "Who am I?" I need to figure out how to respond. First, I should consider that they might be asking about their identity in a philosophical sense, or maybe they\'re looking for a personal connection. But since I\'m an AI, I don\'t have a personal identity. I should clarify that.\n\nI should start by acknowledging that I can\'t know their personal identity. Then, maybe offer to help them explore their own identity. I can ask them to share more about themselves, like their interests or experiences. That way, I can provide a more tailored response. I need to keep the tone friendly and open-ended. Also, make sure not to make assumptions. Let them guide the conversation. Alright, that makes sense.'}, {'type': 'text', 'text': "You are a unique individual with your own thoughts, experiences, and perspectives. While I can't know your specific identity, I can help you explore questions about who you are, what you value, or how you see yourself. If you'd like, you can share more about your interests, goals, or experiences, and I can help you reflect on them. What would you like to discuss? 😊"}], id='0f5ff933-be53-953c-941e-525416a00598', created_at='2026-05-25 08:41:53.829', type='chat', usage=ChatUsage(input_tokens=12, output_tokens=238, time=3.954212, type='chat', metadata=GenerationUsage(input_tokens=12, output_tokens=238)), metadata=None) .. GENERATED FROM PYTHON SOURCE LINES 201-209 Tools API ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ Different model providers differ in their tools APIs, e.g. the tools JSON schema, the tool call/response format. To provide a unified interface, AgentScope solves the problem by: - Providing unified tool call block :ref:`ToolUseBlock ` and tool response block :ref:`ToolResultBlock `, respectively. - Providing a unified tools interface in the ``__call__`` method of the model classes, that accepts a list of tools JSON schemas as follows: .. GENERATED FROM PYTHON SOURCE LINES 209-230 .. code-block:: Python json_schemas = [ { "type": "function", "function": { "name": "google_search", "description": "Search for a query on Google.", "parameters": { "type": "object", "properties": { "query": { "type": "string", "description": "The search query.", }, }, "required": ["query"], }, }, }, ] .. GENERATED FROM PYTHON SOURCE LINES 231-236 Further Reading ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ - :ref:`message` - :ref:`prompt` .. rst-class:: sphx-glr-timing **Total running time of the script:** (0 minutes 9.815 seconds) .. _sphx_glr_download_tutorial_task_model.py: .. only:: html .. container:: sphx-glr-footer sphx-glr-footer-example .. container:: sphx-glr-download sphx-glr-download-jupyter :download:`Download Jupyter notebook: task_model.ipynb ` .. container:: sphx-glr-download sphx-glr-download-python :download:`Download Python source code: task_model.py ` .. container:: sphx-glr-download sphx-glr-download-zip :download:`Download zipped: task_model.zip ` .. only:: html .. rst-class:: sphx-glr-signature `Gallery generated by Sphinx-Gallery `_