Environment Preparations
Set up the environment by referring to the , and set parameters as required by referring to the .
Procedure
The Server supports service applications compatible with third-party framework APIs such as Triton, OpenAI, TGI, and vLLM. You are advised to enable HTTPS communication and configure the required service certificates and private keys by referring to the .
The default IP address and port number for Server startup are
[object Object]. You can modify[object Object]and[object Object]in the[object Object]file to configure the startup IP address and port number.The Server can implement functions such as service status query, model information query, and text/streaming inference.
[object Object]
Start the service in either of the following ways.
The startup command must be executed in the {MindIE installation directory} directory. Run the following command to view the installation path:
[object Object][object Object]
Method 1 (recommended): Start the service using a background process. After the service is started in background process mode, the process is retained when the window is closed.
[object Object]If the following information is printed in the file captured by the standard output stream, the startup is successful:
[object Object]Method 2: Directly start the service.
[object Object]If the following information is displayed, the service is started successfully.
[object Object]
[object Object]
You can use an HTTPS client (such as Linux cURL or Postman) to send HTTPS requests. The following uses Linux cURL commands as an example.
Open a new window and run the following command to send a request: for example, list the current model list.
[object Object][object Object]
This section uses the v1/chat streaming inference API and v1/completions streaming inference API as examples to describe how to call APIs. For details about how to call other APIs, see .
1. v1/chat streaming inference API
[object Object]data: {"id":"endpoint_common_10","object":"chat.completion.chunk","created":1744038509,"model":"llama","choices":[{"index":0,"delta":{"role":"assistant","content":" are"},"logprobs":null,"finish_reason":null}]}
data: {"id":"endpoint_common_10","object":"chat.completion.chunk","created":1744038509,"model":"llama","choices":[{"index":1,"delta":{"role":"assistant","content":" are"},"logprobs":null,"finish_reason":null}]}
data: {"id":"endpoint_common_10","object":"chat.completion.chunk","created":1744038509,"model":"llama","choices":[{"index":0,"delta":{"role":"assistant","content":" a"},"logprobs":null,"finish_reason":null}]}
data: {"id":"endpoint_common_10","object":"chat.completion.chunk","created":1744038509,"model":"llama","choices":[{"index":1,"delta":{"role":"assistant","content":" a"},"logprobs":null,"finish_reason":null}]}
data: {"id":"endpoint_common_10","object":"chat.completion.chunk","created":1744038509,"model":"llama","choices":[{"index":0,"delta":{"role":"assistant","content":" helpful"},"logprobs":null,"finish_reason":null}]}
data: {"id":"endpoint_common_10","object":"chat.completion.chunk","created":1744038509,"model":"llama","choices":[{"index":1,"delta":{"role":"assistant","content":" helpful"},"logprobs":null,"finish_reason":null}]}
data: {"id":"endpoint_common_10","object":"chat.completion.chunk","created":1744038509,"model":"llama","usage":{"prompt_tokens":24,"prompt_tokens_details": {"cached_tokens": 0},"completion_tokens":5,"total_tokens":29,"batch_size":[1,1,1,1,1],"queue_wait_time":[5318,117,82,72,196]},"choices":[{"index":0,"delta":{"role":"assistant","content":" assistant"},"logprobs":null,"finish_reason":"length"}]}
data: {"id":"endpoint_common_10","object":"chat.completion.chunk","created":1744038509,"model":"llama","usage":{"prompt_tokens":24,"prompt_tokens_details": {"cached_tokens": 0},"completion_tokens":5,"total_tokens":29,"batch_size":[1,1,1,1,1],"queue_wait_time":[5318,117,82,72,196]},"choices":[{"index":1,"delta":{"role":"assistant","content":" assistant"},"logprobs":null,"finish_reason":"length"}]}
data: [DONE]
[object Object]2. v1/completions streaming inference API
[object Object]data: [DONE]
[object Object]This section uses the text inference API and streaming inference API as examples to describe how to call APIs. For details about how to call other APIs, see .
1. Text inference API
[object Object]2. Streaming inference API
[object Object][object Object]
[object Object]data: {"prefill_time":null,"decode_time":128.32,"token":{"id":[263],"text":" a"}}
data: {"prefill_time":null,"decode_time":18.17,"token":{"id":[5176],"text":" French"}}
data: {"prefill_time":null,"decode_time":16.80,"token":{"id":[17739],"text":" photograph"}}
data: {"prefill_time":null,"decode_time":16.80,"token":{"id":[261],"text":"er"}}
data: {"prefill_time":null,"decode_time":16.80,"token":{"id":[2729],"text":" based"}}
data: {"prefill_time":null,"decode_time":16.80,"token":{"id":[297],"text":" in"}}
data: {"prefill_time":null,"decode_time":16.80,"token":{"id":[3681],"text":" Paris"}}
data: {"prefill_time":null,"decode_time":16.80,"token":{"id":[29889],"text":"."}}
data: {"prefill_time":null,"decode_time":16.80,"token":{"id":[13],"text":"\n"}}
data: {"prefill_time":null,"decode_time":16.80,"token":{"id":[29902],"text":"I"}}
data: {"prefill_time":null,"decode_time":16.80,"token":{"id":[505],"text":" have"}}
data: {"prefill_time":null,"decode_time":16.80,"token":{"id":[1063],"text":" been"}}
data: {"prefill_time":null,"decode_time":16.80,"token":{"id":[27904],"text":" shooting"}}
data: {"prefill_time":null,"decode_time":16.80,"token":{"id":[1951],"text":" since"}}
data: {"prefill_time":null,"decode_time":16.80,"token":{"id":[306],"text":" I"}}
data: {"prefill_time":null,"decode_time":16.80,"token":{"id":[471],"text":" was"}}
data: {"prefill_time":null,"decode_time":16.80,"token":{"id":[29871],"text":" "}}
data: {"prefill_time":null,"decode_time":16.80,"token":{"id":[29896],"text":"1"}}
data: {"prefill_time":null,"decode_time":16.80,"generated_text":"am a French photographer based in Paris.\nI have been shooting since I was 15","details":{"finish_reason":"length","generated_tokens":20,"seed":846930886},"token":{"id":[29945],"text":null}}
[object Object]MindIE currently supports precision testing with the AISBench tool. An example is shown below. For detailed usage, see .
Procedure
Download and install AISBench.
[object Object][object Object]
Prepare a dataset.
Using gsm8k as an example: download the dataset by clicking , then extract the archive and place the
[object Object]folder under[object Object]in the tool root directory.Configure the
[object Object]file. The following is an example:[object Object]Run the following command to start the serving accuracy test:
[object Object]The command is executed successfully if the command output is as follows:
[object Object]
MindIE supports performance testing with the AISBench tool, as shown in the example below. For detailed usage, see .
Procedure
Download and install AISBench.
[object Object][object Object]
Prepare a dataset.
Using gsm8k as an example: download the dataset by clicking , then extract the archive and place the
[object Object]folder under[object Object]in the tool root directory.Configure the
[object Object]file. The following is an example:[object Object]Run the following command to start the serving performance test:
[object Object]The command is executed successfully if the command output is as follows:
[object Object]In the performance test result, pay attention to the output parameters
[object Object],[object Object],[object Object], and[object Object]. For details about the parameters, see .[object Object]
Log in to the installation node as the installation user and stop the Server service in either of the following ways:
Method 1 (recommended): When the service is started using a background process, you can stop the service in either of the following ways:
Run the
[object Object]command to stop the process.[object Object][object Object]
Alternatively, run the
[object Object]command to stop the process.[object Object]
Method 2: If the service is started by directly starting the process, press
[object Object]to stop the service.