Solution

AI Localized Deployment

AI Gateway delivers unified access and governance for private LLMs — organization-level control, end-to-end observability, and a public-cloud-grade experience for in-house models, with Token costs, usage boundaries, and permissions fully controlled.

Unified access, one-time integrationCredentials never leave the networkOrganization-level governance and cost attributionFull observability and audit
Key Challenges

After private deployment, the real challenges begin

Four typical challenges constrain the scale-up of private AI

Scattered model resources and inconsistent access standards

Multiple private models differ in protocol, endpoints, and authentication

Inconsistent protocolsScattered endpointsRepeated adaptation

Scattered credentials and usage with security and cost risks

Real API keys distributed directly, with no constraint on usage and cost

Key leak riskNo cost accountingNo quota limits

Lack of organization-level control and cost attribution

Usage and cost cannot be attributed by department, permissions not recycled

No usage aggregationHard cost allocationNo permission revoke

Lack of usage observability and timely anomaly response

Call chains are opaque; failures and abnormal usage are hard to trace

No call logsNo dashboardsNo audit
Architecture

Unified management and intelligent routing for private model resources

AI Gateway sits between enterprise applications and private model resources — unified access upward, unified management downward.

Employees / Internal Applications

Access via secondary credentials or SSO identity, all calls go through AI Gateway

Secondary credentials or SSO access
OpenAI-compatible API calls
Personal usage self-service

AI Gateway

The hub of unified access, routing, and governance, hosting all control capabilities

Credential Management
Routing & Rate Limiting
Quota Control
Call Logs
Usage Statistics
Compliance Audit
OpenAI Compatible / Anthropic Compatible / Custom Type

Private LLMs / Public Cloud Models

OpenAI-compatible / Anthropic-compatible / custom endpoints, centrally managed

Private LLM A (OpenAI-compatible)
Private LLM B (Anthropic-compatible)
Public cloud models (optional)
AES-256 · Internal link

Three Access Modes

OpenAI Compatible

Internal / dedicated endpoints using the OpenAI protocol

Self-hosted inference for mainstream open-source models

Anthropic Compatible

Internal endpoints using the Anthropic protocol

Claude-series private / dedicated deployment

Custom Type

Custom endpoints with request/response Schema adaptation

Appliances, dedicated gateways, special protocols

Provider-level Key Capabilities

Endpoint configuration

intranet addresses, calls never leave the network

Credential security

AES-256 encrypted storage, masked display only

Model whitelist

authorized models only, others rejected

Rate limiting

TPM/RPM limits shield inference from traffic spikes

Routing priority

smart routing by priority with active/standby switching

Model sync

automatic sync of provider-side model lists

Cost estimation

billing type and peak periods for cost accounting

Core Capabilities

Five core capabilities covering the full lifecycle of private AI

Unified access, governance, quota control, and observability — a solid operational foundation for private AI

Unified Access

Unified private model integration (provider management)

Centrally manage private model resources, shield protocol differences, and make everything available after one integration.

  • OpenAI / Anthropic compatible and custom endpoints
  • Model whitelist and priority routing
  • AES-256 encrypted credential storage
OpenAI-compatible API
AI GatewayAES-256
OpenAI Compatible
Anthropic Compatible
Custom Type
OpenAI / AnthropicTPM / RPM
Governance

Organizational governance: multi-tenant and org structure

Bring AI usage under organization-level governance with usage aggregation and cost allocation.

  • Strict tenant-level data isolation
  • Hierarchical departments with independent quotas
  • Bulk import/export of org structure
Tenant A
Dept A30M
A-1A-2
Dept B12M
Tenant B
Dept A30M
A-1A-2
Dept B12M
Strictly isolatedA ↔ B
User Access

User access and self-service: credential lifecycle management

A secondary credential system keeps real keys inside the network, with self-service usage dashboards for employees.

  • Secondary credentials replace real keys dynamically
  • Credentials bound to employees/apps and permissions
  • Personal usage dashboard and report export
sk-***-a1b2Masked only
AES-256
Create
Bind
Use
Revoke
Self-service · usage dashboard / reports
Quota Control

Usage quota and budget control

A multi-level, stackable, real-time quota system that keeps cost risk within budget.

  • Four-level quotas: credential / employee / department / group
  • Real-time check on the proxy path, HTTP 429 on overage
  • 80% alert threshold with policy orchestration
Group quota
80%
Dept quota
58%
Employee quota
80%
Credential quota
66%
HTTP 429 on overage429
Observability

End-to-end observability and usage audit

Full logs, multi-dimensional analytics, and cost accounting — failures and abnormal usage can be traced to a single call.

  • Full call-chain logs (async Kafka delivery)
  • Employee / department / app multi-dimensional analytics
  • Cost accounting and compliance audit reports
12.4M
Calls
99.97%
Success rate
¥8,320
Cost
Log streamAsync Kafka delivery
12:00:03sk-***-a1deepseek-chat2001.2s
12:00:05sk-***-b2qwen-max2002.8s
12:00:09sk-***-c3gpt-4o4290.1s
12:00:11sk-***-d4deepseek-r12004.6s

Need a tailored solution for your private environment?

Covering intranet, private cloud, and appliance scenarios with dedicated deployment support