AI Localized Deployment
AI Gateway delivers unified access and governance for private LLMs — organization-level control, end-to-end observability, and a public-cloud-grade experience for in-house models, with Token costs, usage boundaries, and permissions fully controlled.
After private deployment, the real challenges begin
Four typical challenges constrain the scale-up of private AI
Scattered model resources and inconsistent access standards
Multiple private models differ in protocol, endpoints, and authentication
Scattered credentials and usage with security and cost risks
Real API keys distributed directly, with no constraint on usage and cost
Lack of organization-level control and cost attribution
Usage and cost cannot be attributed by department, permissions not recycled
Lack of usage observability and timely anomaly response
Call chains are opaque; failures and abnormal usage are hard to trace
Unified management and intelligent routing for private model resources
AI Gateway sits between enterprise applications and private model resources — unified access upward, unified management downward.
Employees / Internal Applications
Access via secondary credentials or SSO identity, all calls go through AI Gateway
AI Gateway
The hub of unified access, routing, and governance, hosting all control capabilities
Private LLMs / Public Cloud Models
OpenAI-compatible / Anthropic-compatible / custom endpoints, centrally managed
Three Access Modes
OpenAI Compatible
Internal / dedicated endpoints using the OpenAI protocol
Self-hosted inference for mainstream open-source models
Anthropic Compatible
Internal endpoints using the Anthropic protocol
Claude-series private / dedicated deployment
Custom Type
Custom endpoints with request/response Schema adaptation
Appliances, dedicated gateways, special protocols
Provider-level Key Capabilities
Endpoint configuration
intranet addresses, calls never leave the network
Credential security
AES-256 encrypted storage, masked display only
Model whitelist
authorized models only, others rejected
Rate limiting
TPM/RPM limits shield inference from traffic spikes
Routing priority
smart routing by priority with active/standby switching
Model sync
automatic sync of provider-side model lists
Cost estimation
billing type and peak periods for cost accounting
Five core capabilities covering the full lifecycle of private AI
Unified access, governance, quota control, and observability — a solid operational foundation for private AI
Unified private model integration (provider management)
Centrally manage private model resources, shield protocol differences, and make everything available after one integration.
- OpenAI / Anthropic compatible and custom endpoints
- Model whitelist and priority routing
- AES-256 encrypted credential storage
Organizational governance: multi-tenant and org structure
Bring AI usage under organization-level governance with usage aggregation and cost allocation.
- Strict tenant-level data isolation
- Hierarchical departments with independent quotas
- Bulk import/export of org structure
User access and self-service: credential lifecycle management
A secondary credential system keeps real keys inside the network, with self-service usage dashboards for employees.
- Secondary credentials replace real keys dynamically
- Credentials bound to employees/apps and permissions
- Personal usage dashboard and report export
sk-***-a1b2Masked onlyUsage quota and budget control
A multi-level, stackable, real-time quota system that keeps cost risk within budget.
- Four-level quotas: credential / employee / department / group
- Real-time check on the proxy path, HTTP 429 on overage
- 80% alert threshold with policy orchestration
429End-to-end observability and usage audit
Full logs, multi-dimensional analytics, and cost accounting — failures and abnormal usage can be traced to a single call.
- Full call-chain logs (async Kafka delivery)
- Employee / department / app multi-dimensional analytics
- Cost accounting and compliance audit reports
Need a tailored solution for your private environment?
Covering intranet, private cloud, and appliance scenarios with dedicated deployment support