XuanShu API
How do I troubleshoot request latency?
Separate connection time, time to first Token, and total response time before comparing model, context, and network.
Last verified:
How do I troubleshoot request latency?
Separate connection time, time to first Token, and total response time before comparing model, context, and network.
Steps
- Record DNS/TLS, first byte, first Token, and total duration.
- Build a baseline with a short prompt and non-streaming request.
- Compare models, regions, concurrency, and observation periods.
Limits and exceptions
Generation length, tool calls, upstream queues, and network variance affect latency; one sample is not an SLA.