LLM Gateway 아키텍처¶
개요¶
LLM Gateway는 애플리케이션과 LLM 제공자 사이에 위치하는 프록시/미들웨어 레이어다. API Key 관리, 비용 제어, 라우팅, 캐싱, 보안 필터링 등을 중앙에서 처리한다.
핵심 가치:
| 영역 | 문제 | Gateway 해결책 |
|---|---|---|
| 비용 | 모델별 과금 추적 불가 | 요청별 토큰/비용 집계 |
| 안정성 | 단일 제공자 장애시 서비스 중단 | Fallback + Load Balancing |
| 보안 | API Key 노출, PII 유출 | Key Vault 연동, PII 마스킹 |
| 관측성 | 지연/품질 모니터링 부재 | 중앙 로깅, 메트릭 수집 |
| 거버넌스 | 팀별 사용량 제어 불가 | Rate Limiting, 쿼터 관리 |
아키텍처 다이어그램¶
전체 구조¶
+-------------+ +-------------+ +-------------+
| Frontend | | Backend | | Batch Job |
| (Chat UI) | | (API) | | (Pipeline) |
+------+------+ +------+------+ +------+------+
| | |
+------------------+------------------+
|
v
+-------------------+
| LLM Gateway |
| |
| +---------------+ |
| | Auth / ACL | |
| +---------------+ |
| | Rate Limiter | |
| +---------------+ |
| | PII Filter | |
| +---------------+ |
| | Cache Layer | |
| +---------------+ |
| | Router | |
| +---------------+ |
| | Logger/Meter | |
| +---------------+ |
+--------+----------+
|
+--------------+--------------+
| | |
v v v
+------+------+ +----+-----+ +------+------+
| OpenAI | | Anthropic| | Self-hosted|
| API | | API | | (vLLM) |
+-------------+ +----------+ +-------------+
요청 처리 흐름¶
Client Request
|
v
[1] Authentication -----> 401 Unauthorized
|
v
[2] Rate Limit Check ---> 429 Too Many Requests
|
v
[3] Input Filter -------> 400 Blocked Content
|
v
[4] Cache Lookup -------> Cache Hit -> Return Cached Response
|
v (Cache Miss)
[5] Route Selection
|
+---> Primary Provider
| |
| [Success] ---> [6] Output Filter -> [7] Cache Store -> Response
| |
| [Failure] ---> Fallback Provider
| |
| [Success] ---> Response
| |
| [Failure] ---> 503 Service Unavailable
v
[8] Logging & Metrics
핵심 설계 패턴¶
1. Rate Limiting¶
요청 빈도와 토큰 소비를 모두 제어해야 한다.
다중 차원 Rate Limiting:
| 차원 | 단위 | 예시 |
|---|---|---|
| 요청 횟수 (RPM) | requests/min | 60 RPM per API key |
| 토큰 처리량 (TPM) | tokens/min | 100K TPM per team |
| 동시 요청 (Concurrency) | active requests | 10 concurrent per user |
| 일일 예산 (Budget) | USD/day | $50/day per project |
구현 전략:
import time
import asyncio
from dataclasses import dataclass, field
from collections import defaultdict
@dataclass
class RateLimitConfig:
rpm: int = 60 # requests per minute
tpm: int = 100_000 # tokens per minute
daily_budget_usd: float = 50.0
max_concurrent: int = 10
class TokenBucketLimiter:
"""토큰 버킷 기반 Rate Limiter"""
def __init__(self, capacity: int, refill_rate: float):
self.capacity = capacity
self.tokens = capacity
self.refill_rate = refill_rate # tokens per second
self.last_refill = time.monotonic()
self._lock = asyncio.Lock()
async def acquire(self, cost: int = 1) -> bool:
async with self._lock:
now = time.monotonic()
elapsed = now - self.last_refill
self.tokens = min(
self.capacity,
self.tokens + elapsed * self.refill_rate
)
self.last_refill = now
if self.tokens >= cost:
self.tokens -= cost
return True
return False
@property
def wait_time(self) -> float:
if self.tokens > 0:
return 0.0
return (1 - self.tokens) / self.refill_rate
class MultiDimensionLimiter:
"""RPM + TPM + Budget 복합 제한"""
def __init__(self, config: RateLimitConfig):
self.rpm_limiter = TokenBucketLimiter(
capacity=config.rpm,
refill_rate=config.rpm / 60.0
)
self.tpm_limiter = TokenBucketLimiter(
capacity=config.tpm,
refill_rate=config.tpm / 60.0
)
self.daily_spend: dict[str, float] = defaultdict(float)
self.budget_limit = config.daily_budget_usd
async def check(self, api_key: str, estimated_tokens: int) -> dict:
rpm_ok = await self.rpm_limiter.acquire(1)
if not rpm_ok:
return {"allowed": False, "reason": "RPM limit exceeded"}
tpm_ok = await self.tpm_limiter.acquire(estimated_tokens)
if not tpm_ok:
return {"allowed": False, "reason": "TPM limit exceeded"}
today = time.strftime("%Y-%m-%d")
key = f"{api_key}:{today}"
if self.daily_spend[key] >= self.budget_limit:
return {"allowed": False, "reason": "Daily budget exceeded"}
return {"allowed": True}
def record_spend(self, api_key: str, cost_usd: float):
today = time.strftime("%Y-%m-%d")
self.daily_spend[f"{api_key}:{today}"] += cost_usd
2. Load Balancing & Routing¶
모델 선택과 부하 분산을 위한 라우팅 전략.
라우팅 전략 비교:
| 전략 | 설명 | 적합한 상황 |
|---|---|---|
| Round Robin | 순차 분배 | 동일 모델 다중 엔드포인트 |
| Weighted | 가중치 기반 분배 | 성능/비용 차이가 있는 제공자 |
| Latency-based | 지연 시간 기반 | 실시간 응답 중요 |
| Cost-based | 비용 최적화 라우팅 | 예산 제한 환경 |
| Content-based | 요청 내용 분석 | 복잡도별 모델 분기 |
from abc import ABC, abstractmethod
from dataclasses import dataclass
import random
import time
@dataclass
class Provider:
name: str
model: str
endpoint: str
weight: float = 1.0
cost_per_1k_tokens: float = 0.0
avg_latency_ms: float = 0.0
is_healthy: bool = True
error_count: int = 0
last_error_at: float = 0.0
class RouterStrategy(ABC):
@abstractmethod
def select(self, providers: list[Provider], request: dict) -> Provider:
...
class WeightedRouter(RouterStrategy):
def select(self, providers: list[Provider], request: dict) -> Provider:
healthy = [p for p in providers if p.is_healthy]
if not healthy:
raise Exception("No healthy providers available")
total = sum(p.weight for p in healthy)
r = random.uniform(0, total)
cumulative = 0.0
for p in healthy:
cumulative += p.weight
if r <= cumulative:
return p
return healthy[-1]
class CostBasedRouter(RouterStrategy):
"""비용 최적화: 저렴한 제공자 우선, 품질 임계값 적용"""
def __init__(self, quality_threshold: float = 0.8):
self.quality_threshold = quality_threshold
self.quality_scores: dict[str, float] = {}
def select(self, providers: list[Provider], request: dict) -> Provider:
eligible = [
p for p in providers
if p.is_healthy
and self.quality_scores.get(p.name, 1.0) >= self.quality_threshold
]
if not eligible:
eligible = [p for p in providers if p.is_healthy]
return min(eligible, key=lambda p: p.cost_per_1k_tokens)
class ContentBasedRouter(RouterStrategy):
"""요청 복잡도에 따라 모델 티어 선택"""
COMPLEXITY_THRESHOLDS = {
"simple": 100, # 토큰 100 이하 -> 경량 모델
"medium": 1000, # 1000 이하 -> 중간 모델
"complex": float("inf") # 나머지 -> 고성능 모델
}
MODEL_TIERS = {
"simple": ["gpt-4o-mini", "claude-3-haiku"],
"medium": ["gpt-4o", "claude-3.5-sonnet"],
"complex": ["gpt-4o", "claude-3-opus"],
}
def select(self, providers: list[Provider], request: dict) -> Provider:
token_count = self._estimate_tokens(request)
tier = self._classify(token_count)
preferred_models = self.MODEL_TIERS.get(tier, [])
for model in preferred_models:
match = [p for p in providers if p.model == model and p.is_healthy]
if match:
return match[0]
healthy = [p for p in providers if p.is_healthy]
return healthy[0] if healthy else providers[0]
def _estimate_tokens(self, request: dict) -> int:
messages = request.get("messages", [])
return sum(len(m.get("content", "")) // 4 for m in messages)
def _classify(self, tokens: int) -> str:
for tier, threshold in self.COMPLEXITY_THRESHOLDS.items():
if tokens <= threshold:
return tier
return "complex"
3. Fallback & Circuit Breaker¶
장애 전파를 차단하고 자동 복구하는 패턴.
Provider A (Primary)
|
[Request] ---> [Success] ---> Return
|
[Failure] ---> Circuit Breaker Check
|
[CLOSED] ---> Retry (max 2)
|
[OPEN] ---> Skip, go to Fallback
|
[HALF_OPEN] --> Test 1 request
|
Provider B (Fallback)
|
[Request] ---> [Success] ---> Return
|
[Failure] ---> Provider C (Last Resort)
import asyncio
import time
from enum import Enum
class CircuitState(Enum):
CLOSED = "closed" # 정상 - 요청 통과
OPEN = "open" # 차단 - 요청 거부
HALF_OPEN = "half_open" # 테스트 - 제한적 통과
class CircuitBreaker:
def __init__(
self,
failure_threshold: int = 5,
recovery_timeout: float = 30.0,
half_open_max: int = 1
):
self.failure_threshold = failure_threshold
self.recovery_timeout = recovery_timeout
self.half_open_max = half_open_max
self.state = CircuitState.CLOSED
self.failure_count = 0
self.last_failure_at = 0.0
self.half_open_count = 0
def can_execute(self) -> bool:
if self.state == CircuitState.CLOSED:
return True
if self.state == CircuitState.OPEN:
if time.monotonic() - self.last_failure_at > self.recovery_timeout:
self.state = CircuitState.HALF_OPEN
self.half_open_count = 0
return True
return False
if self.state == CircuitState.HALF_OPEN:
return self.half_open_count < self.half_open_max
return False
def record_success(self):
if self.state == CircuitState.HALF_OPEN:
self.state = CircuitState.CLOSED
self.failure_count = 0
def record_failure(self):
self.failure_count += 1
self.last_failure_at = time.monotonic()
if self.state == CircuitState.HALF_OPEN:
self.state = CircuitState.OPEN
elif self.failure_count >= self.failure_threshold:
self.state = CircuitState.OPEN
class FallbackChain:
"""순차적 Fallback 실행"""
def __init__(self, providers: list[Provider]):
self.providers = providers
self.breakers = {p.name: CircuitBreaker() for p in providers}
async def execute(self, request: dict) -> dict:
errors = []
for provider in self.providers:
breaker = self.breakers[provider.name]
if not breaker.can_execute():
continue
try:
result = await self._call_provider(provider, request)
breaker.record_success()
return result
except Exception as e:
breaker.record_failure()
errors.append((provider.name, str(e)))
raise Exception(f"All providers failed: {errors}")
async def _call_provider(self, provider: Provider, request: dict) -> dict:
# 실제 HTTP 호출 구현
...
4. Semantic Cache¶
동일하거나 유사한 요청에 대해 캐시된 응답을 반환하여 비용과 지연을 줄인다.
캐시 전략 비교:
| 전략 | Hit Rate | 구현 복잡도 | 적합한 상황 |
|---|---|---|---|
| Exact Match | 낮음 | 낮음 | 정형화된 쿼리 |
| Semantic (Embedding) | 중간 | 중간 | 자연어 질의 |
| Prompt Template | 높음 | 높음 | 템플릿 기반 호출 |
import hashlib
import json
import numpy as np
from datetime import datetime, timedelta
class SemanticCache:
def __init__(
self,
embedding_fn,
similarity_threshold: float = 0.95,
ttl_seconds: int = 3600,
max_entries: int = 10000
):
self.embedding_fn = embedding_fn
self.threshold = similarity_threshold
self.ttl = timedelta(seconds=ttl_seconds)
self.max_entries = max_entries
self.cache: list[dict] = []
def _make_key(self, messages: list[dict], model: str) -> str:
content = json.dumps({"messages": messages, "model": model}, sort_keys=True)
return hashlib.sha256(content.encode()).hexdigest()
async def get(self, messages: list[dict], model: str) -> dict | None:
# 1) Exact match
key = self._make_key(messages, model)
for entry in self.cache:
if entry["key"] == key and not self._is_expired(entry):
entry["hits"] += 1
return entry["response"]
# 2) Semantic match
query_text = messages[-1].get("content", "")
query_embedding = await self.embedding_fn(query_text)
best_score = 0.0
best_entry = None
for entry in self.cache:
if entry["model"] != model or self._is_expired(entry):
continue
score = self._cosine_similarity(query_embedding, entry["embedding"])
if score > best_score:
best_score = score
best_entry = entry
if best_score >= self.threshold and best_entry:
best_entry["hits"] += 1
return best_entry["response"]
return None
async def put(self, messages: list[dict], model: str, response: dict):
if len(self.cache) >= self.max_entries:
self._evict()
query_text = messages[-1].get("content", "")
embedding = await self.embedding_fn(query_text)
self.cache.append({
"key": self._make_key(messages, model),
"model": model,
"embedding": embedding,
"response": response,
"created_at": datetime.utcnow(),
"hits": 0,
})
def _is_expired(self, entry: dict) -> bool:
return datetime.utcnow() - entry["created_at"] > self.ttl
def _cosine_similarity(self, a: np.ndarray, b: np.ndarray) -> float:
return float(np.dot(a, b) / (np.linalg.norm(a) * np.linalg.norm(b)))
def _evict(self):
# LRU + expired 우선 제거
self.cache = [e for e in self.cache if not self._is_expired(e)]
if len(self.cache) >= self.max_entries:
self.cache.sort(key=lambda e: e["hits"])
self.cache = self.cache[len(self.cache) // 4:]
보안¶
API Key 관리¶
+------------------+ +------------------+
| Developer | | LLM Gateway |
| (Virtual Key) +-------->| |
| vk-abc123... | | Virtual Key Map |
+------------------+ | vk-abc -> sk-... |
| |
| Key Rotation |
| Audit Log |
+--------+---------+
|
+-------+-------+
| |
+----+----+ +-----+-----+
| Vault | | Env Secret |
| (prod) | | (dev) |
+---------+ +-----------+
Virtual Key 패턴: 개발자에게는 내부 가상 키를 발급하고, 실제 제공자 API Key는 Gateway 내부에서만 관리한다.
# Virtual Key -> Provider Key 매핑
VIRTUAL_KEY_MAP = {
"vk-team-alpha-001": {
"provider": "openai",
"actual_key_ref": "vault://secrets/openai/team-alpha",
"rate_limit": RateLimitConfig(rpm=30, tpm=50_000, daily_budget_usd=20),
"allowed_models": ["gpt-4o-mini", "gpt-4o"],
},
"vk-team-beta-001": {
"provider": "anthropic",
"actual_key_ref": "vault://secrets/anthropic/team-beta",
"rate_limit": RateLimitConfig(rpm=60, tpm=100_000, daily_budget_usd=100),
"allowed_models": ["claude-3.5-sonnet", "claude-3-opus"],
},
}
PII 마스킹¶
입력에서 개인정보를 탐지하고 마스킹한 후 LLM에 전달, 응답에서 복원한다.
import re
from dataclasses import dataclass
@dataclass
class PIIMask:
original: str
masked: str
category: str
class PIIFilter:
PATTERNS = {
"phone_kr": re.compile(r"01[016789]-?\d{3,4}-?\d{4}"),
"email": re.compile(r"[a-zA-Z0-9._%+-]+@[a-zA-Z0-9.-]+\.[a-zA-Z]{2,}"),
"rrn_kr": re.compile(r"\d{6}-?[1-4]\d{6}"), # 주민등록번호
"card_number": re.compile(r"\d{4}[-\s]?\d{4}[-\s]?\d{4}[-\s]?\d{4}"),
"ip_address": re.compile(r"\d{1,3}\.\d{1,3}\.\d{1,3}\.\d{1,3}"),
}
def mask(self, text: str) -> tuple[str, list[PIIMask]]:
masks = []
masked_text = text
for category, pattern in self.PATTERNS.items():
for match in pattern.finditer(masked_text):
original = match.group()
placeholder = f"[{category.upper()}_{len(masks)}]"
masks.append(PIIMask(original, placeholder, category))
masked_text = masked_text.replace(original, placeholder, 1)
return masked_text, masks
def unmask(self, text: str, masks: list[PIIMask]) -> str:
result = text
for m in masks:
result = result.replace(m.masked, m.original)
return result
# 사용 예시
pii = PIIFilter()
text = "고객 홍길동 전화번호 010-1234-5678 이메일 [email protected]"
masked, masks = pii.mask(text)
# masked: "고객 홍길동 전화번호 [PHONE_KR_0] 이메일 [EMAIL_1]"
입출력 필터링 (Guardrails)¶
class ContentFilter:
"""입출력 안전성 필터"""
BLOCKED_CATEGORIES = [
"prompt_injection",
"jailbreak",
"harmful_content",
"data_exfiltration",
]
def __init__(self, classifier_fn=None):
self.classifier = classifier_fn
async def filter_input(self, messages: list[dict]) -> dict:
last_message = messages[-1].get("content", "")
# 1) 패턴 기반 차단
if self._detect_injection(last_message):
return {"blocked": True, "reason": "Potential prompt injection"}
# 2) 분류기 기반 차단
if self.classifier:
result = await self.classifier(last_message)
if result["category"] in self.BLOCKED_CATEGORIES:
return {"blocked": True, "reason": result["category"]}
return {"blocked": False}
def _detect_injection(self, text: str) -> bool:
injection_patterns = [
r"ignore\s+(all\s+)?previous\s+instructions",
r"system\s*:\s*you\s+are\s+now",
r"<\|im_start\|>system",
r"```\s*system",
]
text_lower = text.lower()
return any(re.search(p, text_lower) for p in injection_patterns)
주요 솔루션 비교¶
| 솔루션 | 유형 | 주요 특징 | 배포 방식 | 가격 |
|---|---|---|---|---|
| LiteLLM | OSS Proxy | 100+ 모델 통합, OpenAI 호환 API | Self-hosted | 무료 (Enterprise 유료) |
| Kong AI Gateway | Commercial | Kong 기반 플러그인, 엔터프라이즈 기능 | Self-hosted / Cloud | 유료 |
| Cloudflare AI Gateway | SaaS | 엣지 캐싱, Analytics, 간편 설정 | Cloud (Cloudflare) | 무료 티어 + 종량제 |
| OpenRouter | SaaS | 다중 모델 라우팅, 통합 API | Cloud | 모델별 종량제 |
| Portkey | SaaS + OSS | Guardrails, 관측성, A/B 테스트 | Cloud / Self-hosted | 무료 티어 + 유료 |
| AWS Bedrock | Cloud | AWS 네이티브, IAM 통합 | Cloud (AWS) | 종량제 |
기능 상세 비교¶
| 기능 | LiteLLM | Kong AI | Cloudflare | OpenRouter | Portkey |
|---|---|---|---|---|---|
| OpenAI 호환 API | O | O | O | O | O |
| Semantic Cache | O | X | O | X | O |
| Rate Limiting | O | O | O | O | O |
| Fallback | O | O | X | O | O |
| PII 마스킹 | X | Plugin | X | X | O |
| 비용 추적 | O | O | O | O | O |
| A/B 테스트 | X | X | X | X | O |
| 자체 호스팅 | O | O | X | X | O |
| Streaming | O | O | O | O | O |
| Guardrails | X | Plugin | O | X | O |
배포 패턴¶
On-Premise 배포¶
+----------------------------------------------------------+
| Kubernetes Cluster |
| |
| +------------------+ +------------------+ |
| | Gateway Pod (x3) | | Redis | |
| | - LiteLLM | | (Rate Limit | |
| | - Sidecar Proxy | | + Cache State) | |
| +--------+---------+ +--------+---------+ |
| | | |
| +--------+-----------------------+---------+ |
| | Internal Network | |
| +--------+-----------------------+---------+ |
| | | |
| +--------+---------+ +--------+---------+ |
| | vLLM Pod (GPU) | | Postgres | |
| | (Self-hosted LLM)| | (Audit Log | |
| +------------------+ | + Key Store) | |
| +------------------+ |
+----------------------------------------------------------+
|
[External API Calls]
|
+------+------+
| OpenAI API |
| Anthropic |
+-------------+
Cloud-Native 배포¶
[CloudFront / ALB]
|
v
[API Gateway (AWS)]
|
v
[Lambda / ECS - Gateway Logic]
|
+-------> [ElastiCache Redis]
|
+-------> [DynamoDB - Key Store]
|
+-------> [CloudWatch - Metrics]
|
v
[External LLM APIs] + [Bedrock - Self-hosted]
하이브리드 패턴¶
민감 데이터는 On-Premise 모델로, 일반 요청은 Cloud API로 분기.
Request
|
v
[Content Classifier]
|
+--------+--------+
| |
[Sensitive] [Non-sensitive]
| |
v v
+----------------+ +----------------+
| On-Prem vLLM | | Cloud API |
| (No data egress)| | (OpenAI, etc.) |
+----------------+ +----------------+
비용 최적화 전략¶
1. 모델 티어링¶
요청 복잡도에 따라 적절한 모델을 선택하여 비용을 줄인다.
| 티어 | 모델 예시 | 입력 비용/1M tok | 용도 |
|---|---|---|---|
| Tier 1 (경량) | GPT-4o-mini, Haiku | $0.15 - $0.25 | 분류, 추출, 간단한 QA |
| Tier 2 (범용) | GPT-4o, Sonnet | $2.50 - $3.00 | 일반 대화, 분석 |
| Tier 3 (고성능) | GPT-4o, Opus | $10.00 - $15.00 | 복잡 추론, 코딩 |
2. Prompt 최적화¶
class PromptOptimizer:
"""토큰 소비를 줄이는 프롬프트 최적화"""
@staticmethod
def compress_system_prompt(prompt: str, max_tokens: int = 500) -> str:
"""시스템 프롬프트 압축"""
# 불필요한 반복/설명 제거
lines = prompt.strip().split("\n")
compressed = []
seen = set()
for line in lines:
normalized = line.strip().lower()
if normalized and normalized not in seen:
seen.add(normalized)
compressed.append(line)
return "\n".join(compressed)
@staticmethod
def truncate_history(
messages: list[dict],
max_tokens: int = 4000,
keep_system: bool = True
) -> list[dict]:
"""대화 이력 잘라내기 - 최근 메시지 우선"""
system_msgs = []
other_msgs = []
for m in messages:
if m["role"] == "system" and keep_system:
system_msgs.append(m)
else:
other_msgs.append(m)
result = list(system_msgs)
token_count = sum(len(m["content"]) // 4 for m in result)
for m in reversed(other_msgs):
msg_tokens = len(m["content"]) // 4
if token_count + msg_tokens > max_tokens:
break
result.insert(len(system_msgs), m)
token_count += msg_tokens
return result
3. 비용 모니터링 대시보드 쿼리¶
-- 일별 팀별 비용 집계
SELECT
DATE(created_at) AS date,
team_id,
model,
COUNT(*) AS request_count,
SUM(input_tokens) AS total_input_tokens,
SUM(output_tokens) AS total_output_tokens,
SUM(cost_usd) AS total_cost_usd,
AVG(latency_ms) AS avg_latency_ms
FROM llm_requests
WHERE created_at >= CURRENT_DATE - INTERVAL '30 days'
GROUP BY DATE(created_at), team_id, model
ORDER BY total_cost_usd DESC;
-- 캐시 히트율
SELECT
DATE(created_at) AS date,
COUNT(CASE WHEN cache_hit THEN 1 END)::float
/ NULLIF(COUNT(*), 0) AS cache_hit_rate,
SUM(CASE WHEN cache_hit THEN estimated_cost_usd ELSE 0 END) AS cost_saved
FROM llm_requests
WHERE created_at >= CURRENT_DATE - INTERVAL '7 days'
GROUP BY DATE(created_at)
ORDER BY date;
실무 구성 예시: LiteLLM¶
# litellm_config.yaml
model_list:
- model_name: gpt-4o
litellm_params:
model: openai/gpt-4o
api_key: os.environ/OPENAI_API_KEY
rpm: 60
tpm: 100000
- model_name: gpt-4o
litellm_params:
model: azure/gpt-4o-deployment
api_key: os.environ/AZURE_API_KEY
api_base: https://my-resource.openai.azure.com
rpm: 120
tpm: 200000
- model_name: claude-sonnet
litellm_params:
model: anthropic/claude-3.5-sonnet-20241022
api_key: os.environ/ANTHROPIC_API_KEY
- model_name: local-llama
litellm_params:
model: openai/meta-llama/Llama-3-8B
api_base: http://vllm-server:8000/v1
api_key: dummy
router_settings:
routing_strategy: least-busy
num_retries: 3
timeout: 60
fallbacks:
- gpt-4o: [claude-sonnet, local-llama]
- claude-sonnet: [gpt-4o]
allowed_fails: 3
cooldown_time: 30
general_settings:
master_key: os.environ/LITELLM_MASTER_KEY
database_url: os.environ/DATABASE_URL
cache: true
cache_params:
type: redis
host: redis-server
port: 6379
ttl: 3600
# LiteLLM 실행
litellm --config litellm_config.yaml --port 4000
# 호출 (OpenAI 호환)
curl -X POST http://localhost:4000/v1/chat/completions \
-H "Authorization: Bearer sk-litellm-master-key" \
-H "Content-Type: application/json" \
-d '{
"model": "gpt-4o",
"messages": [{"role": "user", "content": "Hello"}]
}'
관측성 (Observability)¶
핵심 메트릭¶
| 메트릭 | 설명 | 알림 임계값 |
|---|---|---|
| Request Latency (p50/p95/p99) | 응답 지연 시간 | p95 > 5s |
| Error Rate | 4xx/5xx 비율 | > 5% |
| Token Throughput | 분당 처리 토큰 | TPM 한도 90% |
| Cache Hit Rate | 캐시 적중률 | < 10% (비효율) |
| Cost per Request | 요청당 평균 비용 | 일일 예산 80% |
| Provider Health | 제공자별 성공률 | < 95% |
| TTFT | Time to First Token | > 2s |
OpenTelemetry 연동¶
from opentelemetry import trace, metrics
from opentelemetry.sdk.trace import TracerProvider
from opentelemetry.sdk.metrics import MeterProvider
tracer = trace.get_tracer("llm-gateway")
meter = metrics.get_meter("llm-gateway")
request_counter = meter.create_counter(
"llm.requests.total",
description="Total LLM requests"
)
latency_histogram = meter.create_histogram(
"llm.request.duration",
unit="ms",
description="Request latency distribution"
)
token_counter = meter.create_counter(
"llm.tokens.total",
description="Total tokens processed"
)
async def handle_request(request: dict) -> dict:
with tracer.start_as_current_span("llm.request") as span:
span.set_attribute("llm.model", request.get("model", ""))
span.set_attribute("llm.provider", "openai")
start = time.monotonic()
try:
response = await call_provider(request)
duration = (time.monotonic() - start) * 1000
latency_histogram.record(duration, {"model": request["model"]})
request_counter.add(1, {"model": request["model"], "status": "success"})
token_counter.add(
response["usage"]["total_tokens"],
{"model": request["model"]}
)
span.set_attribute("llm.tokens.input", response["usage"]["prompt_tokens"])
span.set_attribute("llm.tokens.output", response["usage"]["completion_tokens"])
return response
except Exception as e:
request_counter.add(1, {"model": request["model"], "status": "error"})
span.set_status(trace.StatusCode.ERROR, str(e))
raise
참고 자료¶
| 자료 | URL |
|---|---|
| LiteLLM 공식 문서 | https://docs.litellm.ai/ |
| Portkey AI Gateway | https://portkey.ai/docs |
| Cloudflare AI Gateway | https://developers.cloudflare.com/ai-gateway/ |
| Kong AI Gateway | https://docs.konghq.com/gateway/latest/ai-gateway/ |
| OpenTelemetry LLM SIG | https://github.com/open-telemetry/semantic-conventions/tree/main/docs/gen-ai |
| LLM Gateway Design Patterns (블로그) | https://martinfowler.com/ |