HKUST Generative AI API Service - Supported AI Models
Below is a list of the Generative AI models currently accessible through our API service. Please note that models are deployed and retired on a regular basis. To obtain the current list of supported models, please retrieve them from following endpoint: https://hkust.azure-api.net/hkust-genai/v1/models.
Pricing for these AI models is available in the pricing page,
| Provider | Model name | Suggested Use Case |
|---|---|---|
| DeepSeek | deepseek/deepseek-v4-flash | (preview) Ultra-fast, cost-efficient everyday tasks, high-speed agent routing, and light coding |
| DeepSeek | deepseek/deepseek-v4-flash-0731 | (final) Ultra-fast, cost-efficient everyday tasks, high-speed agent routing, and light coding |
| DeepSeek |
deepseek/deepseek-v4-pro |
(preview) Advanced multilingual reasoning, complex programming, and deep-dive analysis |
| DeepSeek |
deepseek/deepseek-v4-pro-0813 |
(final) Advanced multilingual reasoning, complex programming, and deep-dive analysis |
| gemini-2.5-flash | Fast, low-cost multimodal tasks and chatbots | |
| gemini-2.5-flash-image | Image generation | |
| gemini-2.5-flash-lite | Ultra-cheap, fast classification and routing | |
| gemini-3-flash-preivew | Next-generation high-speed multimodal features | |
| gemini-3.1-pro-preivew | Deep complex analysis and advanced coding | |
| gemini-3.6-flash | Ultra-low latency, real-time conversational agents | |
| gemini-3.7-flash | Multi-step reasoning tasks and lightweight coding assistance | |
| MiniMax | minimax/minimax-m3 | Long-context multimodal reasoning, autonomous agentic workflows |
| Moonshot | moonshotai/kimi-k3 | Extremely long-context document synthesis, deep academic research |
| OpenAI | gpt-4.1-mini | Efficient, low-cost daily text generation |
| OpenAI | gpt-4o | High-stakes complex reasoning |
| OpenAI | gpt-4o-mini | Massive-scale lightweight tasks, summarization, and simple queries |
| OpenAI | gpt-5-mini | Next-gen everyday automation balancing high intelligence and cost |
| OpenAI | gpt-5-nano | On-device processing, mobile apps, and offline micro-tasks |
| OpenAI | gpt-5.6-luna | Low latency and cost efficiency, best for chatbots, and data classification tasks |
| OpenAI | gpt-5.6-sol | Deep academic or legal research, and multi-step agent |
| OpenAI | gpt-5.6-terra | Balanced workflow automation, and moderate coding assistance |
| OpenAI | o3-mini | Budget-friendly deep reasoning for coding and math logic |
| OpenAI | o4-mini | Next-generation advanced STEM solving at a smaller scale |
| OpenAI | text-embedding-3-large | High-precision semantic search and deep RAG applications |
| OpenAI | text-embedding-3-small | Fast, low-cost vector generation for standard search |
| OpenAI | text-embedding-ada-002 | Legacy support for existing, pre-built vector databases |
| Qwen | qwen/qwen3.8-flash | Fast, low-cost agentic workflows, function calling |
| Qwen | qwen/qwen3.8-max | High-complexity reasoning, advanced visual understanding |
| Zhipuai | z-ai/glm-5.3 | Deep-effort software engineering, defensive cybersecurity audits |
| Zhipuai | z-ai/glm-5.3-flash | High-speed developer assistants, and agile agent actions |