[object Object][object Object]

Environment Preparations

Set up the environment by referring to the , and set parameters as required by referring to the .

Procedure

  • The Server supports service applications compatible with third-party framework APIs such as Triton, OpenAI, TGI, and vLLM. You are advised to enable HTTPS communication and configure the required service certificates and private keys by referring to the .

  • The default IP address and port number for Server startup are [object Object]. You can modify [object Object] and [object Object] in the [object Object] file to configure the startup IP address and port number.

  • The Server can implement functions such as service status query, model information query, and text/streaming inference.

[object Object]
  1. Start the service in either of the following ways.

    The startup command must be executed in the {MindIE installation directory} directory. Run the following command to view the installation path:

    [object Object]
    [object Object]
    • Method 1 (recommended): Start the service using a background process. After the service is started in background process mode, the process is retained when the window is closed.

      [object Object]

      If the following information is printed in the file captured by the standard output stream, the startup is successful:

      [object Object]
    • Method 2: Directly start the service.

      [object Object]

      If the following information is displayed, the service is started successfully.

      [object Object]
    [object Object]
  2. You can use an HTTPS client (such as Linux cURL or Postman) to send HTTPS requests. The following uses Linux cURL commands as an example.

    Open a new window and run the following command to send a request: for example, list the current model list.

    [object Object]
    [object Object]
[object Object][object Object]

This section uses the v1/chat streaming inference API and v1/completions streaming inference API as examples to describe how to call APIs. For details about how to call other APIs, see .

1. v1/chat streaming inference API

[object Object]

data: {"id":"endpoint_common_10","object":"chat.completion.chunk","created":1744038509,"model":"llama","choices":[{"index":0,"delta":{"role":"assistant","content":" are"},"logprobs":null,"finish_reason":null}]}

data: {"id":"endpoint_common_10","object":"chat.completion.chunk","created":1744038509,"model":"llama","choices":[{"index":1,"delta":{"role":"assistant","content":" are"},"logprobs":null,"finish_reason":null}]}

data: {"id":"endpoint_common_10","object":"chat.completion.chunk","created":1744038509,"model":"llama","choices":[{"index":0,"delta":{"role":"assistant","content":" a"},"logprobs":null,"finish_reason":null}]}

data: {"id":"endpoint_common_10","object":"chat.completion.chunk","created":1744038509,"model":"llama","choices":[{"index":1,"delta":{"role":"assistant","content":" a"},"logprobs":null,"finish_reason":null}]}

data: {"id":"endpoint_common_10","object":"chat.completion.chunk","created":1744038509,"model":"llama","choices":[{"index":0,"delta":{"role":"assistant","content":" helpful"},"logprobs":null,"finish_reason":null}]}

data: {"id":"endpoint_common_10","object":"chat.completion.chunk","created":1744038509,"model":"llama","choices":[{"index":1,"delta":{"role":"assistant","content":" helpful"},"logprobs":null,"finish_reason":null}]}

data: {"id":"endpoint_common_10","object":"chat.completion.chunk","created":1744038509,"model":"llama","usage":{"prompt_tokens":24,"prompt_tokens_details": {"cached_tokens": 0},"completion_tokens":5,"total_tokens":29,"batch_size":[1,1,1,1,1],"queue_wait_time":[5318,117,82,72,196]},"choices":[{"index":0,"delta":{"role":"assistant","content":" assistant"},"logprobs":null,"finish_reason":"length"}]}

data: {"id":"endpoint_common_10","object":"chat.completion.chunk","created":1744038509,"model":"llama","usage":{"prompt_tokens":24,"prompt_tokens_details": {"cached_tokens": 0},"completion_tokens":5,"total_tokens":29,"batch_size":[1,1,1,1,1],"queue_wait_time":[5318,117,82,72,196]},"choices":[{"index":1,"delta":{"role":"assistant","content":" assistant"},"logprobs":null,"finish_reason":"length"}]}

data: [DONE]

[object Object]

2. v1/completions streaming inference API

[object Object]

data: [DONE]

[object Object][object Object]

This section uses the text inference API and streaming inference API as examples to describe how to call APIs. For details about how to call other APIs, see .

1. Text inference API

[object Object]

2. Streaming inference API

[object Object][object Object]

[object Object]

data: {"prefill_time":null,"decode_time":128.32,"token":{"id":[263],"text":" a"}}

data: {"prefill_time":null,"decode_time":18.17,"token":{"id":[5176],"text":" French"}}

data: {"prefill_time":null,"decode_time":16.80,"token":{"id":[17739],"text":" photograph"}}

data: {"prefill_time":null,"decode_time":16.80,"token":{"id":[261],"text":"er"}}

data: {"prefill_time":null,"decode_time":16.80,"token":{"id":[2729],"text":" based"}}

data: {"prefill_time":null,"decode_time":16.80,"token":{"id":[297],"text":" in"}}

data: {"prefill_time":null,"decode_time":16.80,"token":{"id":[3681],"text":" Paris"}}

data: {"prefill_time":null,"decode_time":16.80,"token":{"id":[29889],"text":"."}}

data: {"prefill_time":null,"decode_time":16.80,"token":{"id":[13],"text":"\n"}}

data: {"prefill_time":null,"decode_time":16.80,"token":{"id":[29902],"text":"I"}}

data: {"prefill_time":null,"decode_time":16.80,"token":{"id":[505],"text":" have"}}

data: {"prefill_time":null,"decode_time":16.80,"token":{"id":[1063],"text":" been"}}

data: {"prefill_time":null,"decode_time":16.80,"token":{"id":[27904],"text":" shooting"}}

data: {"prefill_time":null,"decode_time":16.80,"token":{"id":[1951],"text":" since"}}

data: {"prefill_time":null,"decode_time":16.80,"token":{"id":[306],"text":" I"}}

data: {"prefill_time":null,"decode_time":16.80,"token":{"id":[471],"text":" was"}}

data: {"prefill_time":null,"decode_time":16.80,"token":{"id":[29871],"text":" "}}

data: {"prefill_time":null,"decode_time":16.80,"token":{"id":[29896],"text":"1"}}

data: {"prefill_time":null,"decode_time":16.80,"generated_text":"am a French photographer based in Paris.\nI have been shooting since I was 15","details":{"finish_reason":"length","generated_tokens":20,"seed":846930886},"token":{"id":[29945],"text":null}}

[object Object][object Object]

MindIE currently supports precision testing with the AISBench tool. An example is shown below. For detailed usage, see .

Procedure

  1. Download and install AISBench.

    [object Object]
    [object Object]
  2. Prepare a dataset.

    Using gsm8k as an example: download the dataset by clicking , then extract the archive and place the [object Object] folder under [object Object] in the tool root directory.

  3. Configure the [object Object] file. The following is an example:

    [object Object]
  4. Run the following command to start the serving accuracy test:

    [object Object]

    The command is executed successfully if the command output is as follows:

    [object Object]
[object Object]

MindIE supports performance testing with the AISBench tool, as shown in the example below. For detailed usage, see .

Procedure

  1. Download and install AISBench.

    [object Object]
    [object Object]
  2. Prepare a dataset.

    Using gsm8k as an example: download the dataset by clicking , then extract the archive and place the [object Object] folder under [object Object] in the tool root directory.

  3. Configure the [object Object] file. The following is an example:

    [object Object]
  4. Run the following command to start the serving performance test:

    [object Object]

    The command is executed successfully if the command output is as follows:

    [object Object]

    In the performance test result, pay attention to the output parameters [object Object], [object Object], [object Object], and [object Object]. For details about the parameters, see .

    [object Object]
[object Object]

Log in to the installation node as the installation user and stop the Server service in either of the following ways:

  • Method 1 (recommended): When the service is started using a background process, you can stop the service in either of the following ways:

    • Run the [object Object] command to stop the process.

      [object Object]
      [object Object]
    • Alternatively, run the [object Object] command to stop the process.

      [object Object]
  • Method 2: If the service is started by directly starting the process, press [object Object] to stop the service.