ENTERPRISE MODEL & COMPUTE CLOUD LIVE

Call models, deploy services,
and rent GPUs — in one place

O17Pod provides DeepSeek V4 Pro, DeepSeek V4 Flash, Kimi K3 and other leading model services through one API or dedicated inference deployments on Kubernetes and Docker hosts.

›_ Try the API
✓ DeepSeek V4 / Kimi K3✓ OpenAI compatible✓ Per-key billing controls
O17Pod COMMAND CENTERTenant: Production ● Healthy
PRODUCTION WORKSPACEGood evening, BuilderLIVE
Token balance86.4M↑ 12.8%
Instances12/ 50 quota
API calls today1.28M↑ 18.6%
API traffic
RUNNINGDeepSeek-R1-32B4 × A100128 ms
RUNNINGQwen2.5-72B2 × L40S96 ms
›_ Cluster  Healthy GPU 63% ━━━API 99.98%
Request routing12.8K RPS+24.6%
0140+Models
02500+GPUs
0399.95%API availability
04<120msGateway overhead
053Compatible protocols
Demonstration data. Production targets are defined by the service agreement.
01 / PRODUCT CAPABILITIES

From model calls to compute delivery

O17Pod combines an AI gateway, inference deployment, GPU scheduling and commercial billing in one MaaS platform.

01

Unified Model API

One key and one endpoint across providers.

  • OpenAI & Anthropic compatibility
  • Streaming and tool use
  • Routing, fallback and retries
Learn more
02

Dedicated Deployment

Turn open or fine-tuned models into private APIs.

  • vLLM, TGI and SGLang
  • Kubernetes and Docker
  • Autoscaling, canary and rollback
Learn more
03

GPU Compute Rental

Launch training, inference and notebook containers.

  • H100, A100, L40S and 4090
  • Custom images and storage
  • Hourly, daily and monthly billing
Learn more
04

Tokens & Billing

Control prepaid and usage-based spend.

  • Token plans and PAYG
  • Tenant, project and key budgets
  • Orders, invoices and alerts
Learn more
02 / MODEL CATALOG

DeepSeek, Kimi, MiniMax and other leading models

O17Pod provides unified model API and dedicated enterprise deployment services, with model choice based on capability, latency, cost and compliance.

Operational
D

DeepSeek V4 Pro

DeepSeek · Reasoning / Agents
AVAILABLE
Context / Output1MService modeUnified API / Dedicated
O17Pod model serviceIntegrated
D

DeepSeek V4 Flash

DeepSeek · Chat / Reasoning
AVAILABLE
Context / Output1MService modeUnified API / Dedicated
O17Pod model serviceIntegrated
K

Kimi K3

Moonshot AI · Multimodal / Reasoning
AVAILABLE
Context / Output1MService modeUnified API / Dedicated
O17Pod model serviceIntegrated
M

MiniMax M3

MiniMax · Multimodal / Agents
AVAILABLE
Context / Output1MService modeUnified API / Dedicated
O17Pod model serviceIntegrated
H

MiniMax H3

MiniMax · Video / Omni-generation
AVAILABLE
Context / Output15s · 2KService modeUnified API / Dedicated
O17Pod model serviceIntegrated
D

DeepSeek R1

DeepSeek · Reasoning
AVAILABLE
Context / Output128KService modeUnified API / Dedicated
O17Pod model serviceIntegrated
Q

Qwen 2.5 72B

Alibaba · Chat
AVAILABLE
Context / Output128KService modeUnified API / Dedicated
O17Pod model serviceIntegrated
L

Llama 3.3 70B

Meta · Chat
AVAILABLE
Context / Output128KService modeUnified API / Dedicated
O17Pod model serviceIntegrated
V

Qwen2-VL 72B

Alibaba · Multimodal
AVAILABLE
Context / Output32KService modeUnified API / Dedicated
O17Pod model serviceIntegrated
B

BGE-M3

BAAI · Embedding
AVAILABLE
Context / Output8KService modeUnified API / Dedicated
O17Pod model serviceIntegrated
S

Stable Diffusion XL

Stability AI · Image
AVAILABLE
Context / OutputService modeUnified API / Dedicated
O17Pod model serviceIntegrated
03 / COMPUTE RENTAL

Two delivery modes,
one compute experience

Launch cloud-native Kubernetes Pods or rent containers on designated Docker GPU hosts with unified inventory, images, networking, monitoring and billing.

Kubernetes PodElastic scheduling · self-healing · per-second billingRECOMMENDED
Docker on selected hostsFixed node · port mapping · container lifecycle
GPU INVENTORY / SHANGHAI LIVE
01NVIDIA H100 SXM80 GB HBM3Available24Utilization 62%¥32.80/GPU·h
02NVIDIA A100 SXM80 GB HBM2eAvailable48Utilization 71%¥18.60/GPU·h
03NVIDIA L40S48 GB GDDR6Available36Utilization 54%¥11.20/GPU·h
04NVIDIA RTX 409024 GB GDDR6XAvailable17Utilization 83%¥5.80/GPU·h
04 / DEVELOPER EXPERIENCE

Change one URL.
Keep your code.

Keep your OpenAI SDK and application logic. Replace the API key, base URL and model name.

  • ● Chat Completions
  • ● Responses API
  • ● Anthropic Messages
  • ● Embedding & Rerank
Developer support
quickstart.py
from openai import OpenAI

client = OpenAI(
    api_key="YOUR_O17POD_API_KEY",
    base_url="https://api.o17.dev/v1"
)

response = client.chat.completions.create(
    model="deepseek-r1",
    messages=[{"role":"user","content":"Hello O17Pod"}],
    stream=True
)
● 200 OKTTFT 84msstream: true
05 / ENTERPRISE PLATFORM

Built for production

Multi-tenant isolation, RBAC, encryption, audit trails, private networking and full-stack observability.

Web AppAI AgentEnterprise RAGOpenAI SDK
Multi-Protocol API GatewayAUTHRATE LIMITROUTINGMETERING
KubernetesPod · Service · HPA
Docker HostsNode · Container · Port
External APIsProvider · Failover · SLA
06 / PRICING

Scale from prototype to production

Choose public models, token plans, private deployments and GPU compute as your workload grows.

DEVELOPER

Developer

¥0
  • Public model API
  • Starter token plans
  • 20 RPM default
  • Community support
ENTERPRISE

Enterprise

Custom
  • Dedicated deployments
  • Reserved GPU pools
  • Enterprise SLA and support
  • Private and hybrid cloud
07 / OPEN ECOSYSTEM

Built on proven open source

O17Pod adds one product, billing and operations layer over dependable infrastructure.

Kubernetes
Docker
vLLM
SGLang
Hugging Face
Prometheus
Grafana
OpenTelemetry
READY TO BUILD?

From your first API call
to production-scale inference

Manage models, compute, tokens and spend in one place — and focus on your AI product.