콘텐츠로 이동
Data Prep
상세

LLM Gateway 아키텍처


개요

LLM Gateway는 애플리케이션과 LLM 제공자 사이에 위치하는 프록시/미들웨어 레이어다. API Key 관리, 비용 제어, 라우팅, 캐싱, 보안 필터링 등을 중앙에서 처리한다.

핵심 가치:

영역 문제 Gateway 해결책
비용 모델별 과금 추적 불가 요청별 토큰/비용 집계
안정성 단일 제공자 장애시 서비스 중단 Fallback + Load Balancing
보안 API Key 노출, PII 유출 Key Vault 연동, PII 마스킹
관측성 지연/품질 모니터링 부재 중앙 로깅, 메트릭 수집
거버넌스 팀별 사용량 제어 불가 Rate Limiting, 쿼터 관리

아키텍처 다이어그램

전체 구조

+-------------+    +-------------+    +-------------+
|  Frontend   |    |  Backend    |    |  Batch Job  |
|  (Chat UI)  |    |  (API)      |    |  (Pipeline) |
+------+------+    +------+------+    +------+------+
       |                  |                  |
       +------------------+------------------+
                          |
                          v
                +-------------------+
                |   LLM Gateway     |
                |                   |
                | +---------------+ |
                | | Auth / ACL    | |
                | +---------------+ |
                | | Rate Limiter  | |
                | +---------------+ |
                | | PII Filter    | |
                | +---------------+ |
                | | Cache Layer   | |
                | +---------------+ |
                | | Router        | |
                | +---------------+ |
                | | Logger/Meter  | |
                | +---------------+ |
                +--------+----------+
                         |
          +--------------+--------------+
          |              |              |
          v              v              v
   +------+------+ +----+-----+ +------+------+
   |  OpenAI     | |  Anthropic| |  Self-hosted|
   |  API        | |  API      | |  (vLLM)     |
   +-------------+ +----------+ +-------------+

요청 처리 흐름

Client Request
      |
      v
[1] Authentication -----> 401 Unauthorized
      |
      v
[2] Rate Limit Check ---> 429 Too Many Requests
      |
      v
[3] Input Filter -------> 400 Blocked Content
      |
      v
[4] Cache Lookup -------> Cache Hit -> Return Cached Response
      |
      v (Cache Miss)
[5] Route Selection
      |
      +---> Primary Provider
      |         |
      |    [Success] ---> [6] Output Filter -> [7] Cache Store -> Response
      |         |
      |    [Failure] ---> Fallback Provider
      |                       |
      |                  [Success] ---> Response
      |                       |
      |                  [Failure] ---> 503 Service Unavailable
      v
[8] Logging & Metrics

핵심 설계 패턴

1. Rate Limiting

요청 빈도와 토큰 소비를 모두 제어해야 한다.

다중 차원 Rate Limiting:

차원 단위 예시
요청 횟수 (RPM) requests/min 60 RPM per API key
토큰 처리량 (TPM) tokens/min 100K TPM per team
동시 요청 (Concurrency) active requests 10 concurrent per user
일일 예산 (Budget) USD/day $50/day per project

구현 전략:

import time
import asyncio
from dataclasses import dataclass, field
from collections import defaultdict

@dataclass
class RateLimitConfig:
    rpm: int = 60              # requests per minute
    tpm: int = 100_000         # tokens per minute
    daily_budget_usd: float = 50.0
    max_concurrent: int = 10

class TokenBucketLimiter:
    """토큰 버킷 기반 Rate Limiter"""

    def __init__(self, capacity: int, refill_rate: float):
        self.capacity = capacity
        self.tokens = capacity
        self.refill_rate = refill_rate  # tokens per second
        self.last_refill = time.monotonic()
        self._lock = asyncio.Lock()

    async def acquire(self, cost: int = 1) -> bool:
        async with self._lock:
            now = time.monotonic()
            elapsed = now - self.last_refill
            self.tokens = min(
                self.capacity,
                self.tokens + elapsed * self.refill_rate
            )
            self.last_refill = now

            if self.tokens >= cost:
                self.tokens -= cost
                return True
            return False

    @property
    def wait_time(self) -> float:
        if self.tokens > 0:
            return 0.0
        return (1 - self.tokens) / self.refill_rate


class MultiDimensionLimiter:
    """RPM + TPM + Budget 복합 제한"""

    def __init__(self, config: RateLimitConfig):
        self.rpm_limiter = TokenBucketLimiter(
            capacity=config.rpm,
            refill_rate=config.rpm / 60.0
        )
        self.tpm_limiter = TokenBucketLimiter(
            capacity=config.tpm,
            refill_rate=config.tpm / 60.0
        )
        self.daily_spend: dict[str, float] = defaultdict(float)
        self.budget_limit = config.daily_budget_usd

    async def check(self, api_key: str, estimated_tokens: int) -> dict:
        rpm_ok = await self.rpm_limiter.acquire(1)
        if not rpm_ok:
            return {"allowed": False, "reason": "RPM limit exceeded"}

        tpm_ok = await self.tpm_limiter.acquire(estimated_tokens)
        if not tpm_ok:
            return {"allowed": False, "reason": "TPM limit exceeded"}

        today = time.strftime("%Y-%m-%d")
        key = f"{api_key}:{today}"
        if self.daily_spend[key] >= self.budget_limit:
            return {"allowed": False, "reason": "Daily budget exceeded"}

        return {"allowed": True}

    def record_spend(self, api_key: str, cost_usd: float):
        today = time.strftime("%Y-%m-%d")
        self.daily_spend[f"{api_key}:{today}"] += cost_usd

2. Load Balancing & Routing

모델 선택과 부하 분산을 위한 라우팅 전략.

라우팅 전략 비교:

전략 설명 적합한 상황
Round Robin 순차 분배 동일 모델 다중 엔드포인트
Weighted 가중치 기반 분배 성능/비용 차이가 있는 제공자
Latency-based 지연 시간 기반 실시간 응답 중요
Cost-based 비용 최적화 라우팅 예산 제한 환경
Content-based 요청 내용 분석 복잡도별 모델 분기
from abc import ABC, abstractmethod
from dataclasses import dataclass
import random
import time

@dataclass
class Provider:
    name: str
    model: str
    endpoint: str
    weight: float = 1.0
    cost_per_1k_tokens: float = 0.0
    avg_latency_ms: float = 0.0
    is_healthy: bool = True
    error_count: int = 0
    last_error_at: float = 0.0

class RouterStrategy(ABC):
    @abstractmethod
    def select(self, providers: list[Provider], request: dict) -> Provider:
        ...

class WeightedRouter(RouterStrategy):
    def select(self, providers: list[Provider], request: dict) -> Provider:
        healthy = [p for p in providers if p.is_healthy]
        if not healthy:
            raise Exception("No healthy providers available")

        total = sum(p.weight for p in healthy)
        r = random.uniform(0, total)
        cumulative = 0.0
        for p in healthy:
            cumulative += p.weight
            if r <= cumulative:
                return p
        return healthy[-1]

class CostBasedRouter(RouterStrategy):
    """비용 최적화: 저렴한 제공자 우선, 품질 임계값 적용"""

    def __init__(self, quality_threshold: float = 0.8):
        self.quality_threshold = quality_threshold
        self.quality_scores: dict[str, float] = {}

    def select(self, providers: list[Provider], request: dict) -> Provider:
        eligible = [
            p for p in providers
            if p.is_healthy
            and self.quality_scores.get(p.name, 1.0) >= self.quality_threshold
        ]
        if not eligible:
            eligible = [p for p in providers if p.is_healthy]

        return min(eligible, key=lambda p: p.cost_per_1k_tokens)

class ContentBasedRouter(RouterStrategy):
    """요청 복잡도에 따라 모델 티어 선택"""

    COMPLEXITY_THRESHOLDS = {
        "simple": 100,    # 토큰 100 이하 -> 경량 모델
        "medium": 1000,   # 1000 이하 -> 중간 모델
        "complex": float("inf")  # 나머지 -> 고성능 모델
    }

    MODEL_TIERS = {
        "simple": ["gpt-4o-mini", "claude-3-haiku"],
        "medium": ["gpt-4o", "claude-3.5-sonnet"],
        "complex": ["gpt-4o", "claude-3-opus"],
    }

    def select(self, providers: list[Provider], request: dict) -> Provider:
        token_count = self._estimate_tokens(request)
        tier = self._classify(token_count)
        preferred_models = self.MODEL_TIERS.get(tier, [])

        for model in preferred_models:
            match = [p for p in providers if p.model == model and p.is_healthy]
            if match:
                return match[0]

        healthy = [p for p in providers if p.is_healthy]
        return healthy[0] if healthy else providers[0]

    def _estimate_tokens(self, request: dict) -> int:
        messages = request.get("messages", [])
        return sum(len(m.get("content", "")) // 4 for m in messages)

    def _classify(self, tokens: int) -> str:
        for tier, threshold in self.COMPLEXITY_THRESHOLDS.items():
            if tokens <= threshold:
                return tier
        return "complex"

3. Fallback & Circuit Breaker

장애 전파를 차단하고 자동 복구하는 패턴.

Provider A (Primary)
      |
  [Request] ---> [Success] ---> Return
      |
  [Failure] ---> Circuit Breaker Check
                      |
              [CLOSED] ---> Retry (max 2)
                      |
              [OPEN]   ---> Skip, go to Fallback
                      |
              [HALF_OPEN] --> Test 1 request
                      |
Provider B (Fallback)
      |
  [Request] ---> [Success] ---> Return
      |
  [Failure] ---> Provider C (Last Resort)
import asyncio
import time
from enum import Enum

class CircuitState(Enum):
    CLOSED = "closed"        # 정상 - 요청 통과
    OPEN = "open"            # 차단 - 요청 거부
    HALF_OPEN = "half_open"  # 테스트 - 제한적 통과

class CircuitBreaker:
    def __init__(
        self,
        failure_threshold: int = 5,
        recovery_timeout: float = 30.0,
        half_open_max: int = 1
    ):
        self.failure_threshold = failure_threshold
        self.recovery_timeout = recovery_timeout
        self.half_open_max = half_open_max

        self.state = CircuitState.CLOSED
        self.failure_count = 0
        self.last_failure_at = 0.0
        self.half_open_count = 0

    def can_execute(self) -> bool:
        if self.state == CircuitState.CLOSED:
            return True

        if self.state == CircuitState.OPEN:
            if time.monotonic() - self.last_failure_at > self.recovery_timeout:
                self.state = CircuitState.HALF_OPEN
                self.half_open_count = 0
                return True
            return False

        if self.state == CircuitState.HALF_OPEN:
            return self.half_open_count < self.half_open_max

        return False

    def record_success(self):
        if self.state == CircuitState.HALF_OPEN:
            self.state = CircuitState.CLOSED
        self.failure_count = 0

    def record_failure(self):
        self.failure_count += 1
        self.last_failure_at = time.monotonic()

        if self.state == CircuitState.HALF_OPEN:
            self.state = CircuitState.OPEN
        elif self.failure_count >= self.failure_threshold:
            self.state = CircuitState.OPEN


class FallbackChain:
    """순차적 Fallback 실행"""

    def __init__(self, providers: list[Provider]):
        self.providers = providers
        self.breakers = {p.name: CircuitBreaker() for p in providers}

    async def execute(self, request: dict) -> dict:
        errors = []
        for provider in self.providers:
            breaker = self.breakers[provider.name]
            if not breaker.can_execute():
                continue

            try:
                result = await self._call_provider(provider, request)
                breaker.record_success()
                return result
            except Exception as e:
                breaker.record_failure()
                errors.append((provider.name, str(e)))

        raise Exception(f"All providers failed: {errors}")

    async def _call_provider(self, provider: Provider, request: dict) -> dict:
        # 실제 HTTP 호출 구현
        ...

4. Semantic Cache

동일하거나 유사한 요청에 대해 캐시된 응답을 반환하여 비용과 지연을 줄인다.

캐시 전략 비교:

전략 Hit Rate 구현 복잡도 적합한 상황
Exact Match 낮음 낮음 정형화된 쿼리
Semantic (Embedding) 중간 중간 자연어 질의
Prompt Template 높음 높음 템플릿 기반 호출
import hashlib
import json
import numpy as np
from datetime import datetime, timedelta

class SemanticCache:
    def __init__(
        self,
        embedding_fn,
        similarity_threshold: float = 0.95,
        ttl_seconds: int = 3600,
        max_entries: int = 10000
    ):
        self.embedding_fn = embedding_fn
        self.threshold = similarity_threshold
        self.ttl = timedelta(seconds=ttl_seconds)
        self.max_entries = max_entries
        self.cache: list[dict] = []

    def _make_key(self, messages: list[dict], model: str) -> str:
        content = json.dumps({"messages": messages, "model": model}, sort_keys=True)
        return hashlib.sha256(content.encode()).hexdigest()

    async def get(self, messages: list[dict], model: str) -> dict | None:
        # 1) Exact match
        key = self._make_key(messages, model)
        for entry in self.cache:
            if entry["key"] == key and not self._is_expired(entry):
                entry["hits"] += 1
                return entry["response"]

        # 2) Semantic match
        query_text = messages[-1].get("content", "")
        query_embedding = await self.embedding_fn(query_text)

        best_score = 0.0
        best_entry = None
        for entry in self.cache:
            if entry["model"] != model or self._is_expired(entry):
                continue
            score = self._cosine_similarity(query_embedding, entry["embedding"])
            if score > best_score:
                best_score = score
                best_entry = entry

        if best_score >= self.threshold and best_entry:
            best_entry["hits"] += 1
            return best_entry["response"]

        return None

    async def put(self, messages: list[dict], model: str, response: dict):
        if len(self.cache) >= self.max_entries:
            self._evict()

        query_text = messages[-1].get("content", "")
        embedding = await self.embedding_fn(query_text)

        self.cache.append({
            "key": self._make_key(messages, model),
            "model": model,
            "embedding": embedding,
            "response": response,
            "created_at": datetime.utcnow(),
            "hits": 0,
        })

    def _is_expired(self, entry: dict) -> bool:
        return datetime.utcnow() - entry["created_at"] > self.ttl

    def _cosine_similarity(self, a: np.ndarray, b: np.ndarray) -> float:
        return float(np.dot(a, b) / (np.linalg.norm(a) * np.linalg.norm(b)))

    def _evict(self):
        # LRU + expired 우선 제거
        self.cache = [e for e in self.cache if not self._is_expired(e)]
        if len(self.cache) >= self.max_entries:
            self.cache.sort(key=lambda e: e["hits"])
            self.cache = self.cache[len(self.cache) // 4:]

보안

API Key 관리

+------------------+         +------------------+
|  Developer       |         |  LLM Gateway     |
|  (Virtual Key)   +-------->|                  |
|  vk-abc123...    |         | Virtual Key Map  |
+------------------+         | vk-abc -> sk-... |
                             |                  |
                             | Key Rotation     |
                             | Audit Log        |
                             +--------+---------+
                                      |
                              +-------+-------+
                              |               |
                         +----+----+    +-----+-----+
                         | Vault   |    | Env Secret |
                         | (prod)  |    | (dev)      |
                         +---------+    +-----------+

Virtual Key 패턴: 개발자에게는 내부 가상 키를 발급하고, 실제 제공자 API Key는 Gateway 내부에서만 관리한다.

# Virtual Key -> Provider Key 매핑
VIRTUAL_KEY_MAP = {
    "vk-team-alpha-001": {
        "provider": "openai",
        "actual_key_ref": "vault://secrets/openai/team-alpha",
        "rate_limit": RateLimitConfig(rpm=30, tpm=50_000, daily_budget_usd=20),
        "allowed_models": ["gpt-4o-mini", "gpt-4o"],
    },
    "vk-team-beta-001": {
        "provider": "anthropic",
        "actual_key_ref": "vault://secrets/anthropic/team-beta",
        "rate_limit": RateLimitConfig(rpm=60, tpm=100_000, daily_budget_usd=100),
        "allowed_models": ["claude-3.5-sonnet", "claude-3-opus"],
    },
}

PII 마스킹

입력에서 개인정보를 탐지하고 마스킹한 후 LLM에 전달, 응답에서 복원한다.

import re
from dataclasses import dataclass

@dataclass
class PIIMask:
    original: str
    masked: str
    category: str

class PIIFilter:
    PATTERNS = {
        "phone_kr": re.compile(r"01[016789]-?\d{3,4}-?\d{4}"),
        "email": re.compile(r"[a-zA-Z0-9._%+-]+@[a-zA-Z0-9.-]+\.[a-zA-Z]{2,}"),
        "rrn_kr": re.compile(r"\d{6}-?[1-4]\d{6}"),  # 주민등록번호
        "card_number": re.compile(r"\d{4}[-\s]?\d{4}[-\s]?\d{4}[-\s]?\d{4}"),
        "ip_address": re.compile(r"\d{1,3}\.\d{1,3}\.\d{1,3}\.\d{1,3}"),
    }

    def mask(self, text: str) -> tuple[str, list[PIIMask]]:
        masks = []
        masked_text = text

        for category, pattern in self.PATTERNS.items():
            for match in pattern.finditer(masked_text):
                original = match.group()
                placeholder = f"[{category.upper()}_{len(masks)}]"
                masks.append(PIIMask(original, placeholder, category))
                masked_text = masked_text.replace(original, placeholder, 1)

        return masked_text, masks

    def unmask(self, text: str, masks: list[PIIMask]) -> str:
        result = text
        for m in masks:
            result = result.replace(m.masked, m.original)
        return result

# 사용 예시
pii = PIIFilter()
text = "고객 홍길동 전화번호 010-1234-5678 이메일 [email protected]"
masked, masks = pii.mask(text)
# masked: "고객 홍길동 전화번호 [PHONE_KR_0] 이메일 [EMAIL_1]"

입출력 필터링 (Guardrails)

class ContentFilter:
    """입출력 안전성 필터"""

    BLOCKED_CATEGORIES = [
        "prompt_injection",
        "jailbreak",
        "harmful_content",
        "data_exfiltration",
    ]

    def __init__(self, classifier_fn=None):
        self.classifier = classifier_fn

    async def filter_input(self, messages: list[dict]) -> dict:
        last_message = messages[-1].get("content", "")

        # 1) 패턴 기반 차단
        if self._detect_injection(last_message):
            return {"blocked": True, "reason": "Potential prompt injection"}

        # 2) 분류기 기반 차단
        if self.classifier:
            result = await self.classifier(last_message)
            if result["category"] in self.BLOCKED_CATEGORIES:
                return {"blocked": True, "reason": result["category"]}

        return {"blocked": False}

    def _detect_injection(self, text: str) -> bool:
        injection_patterns = [
            r"ignore\s+(all\s+)?previous\s+instructions",
            r"system\s*:\s*you\s+are\s+now",
            r"<\|im_start\|>system",
            r"```\s*system",
        ]
        text_lower = text.lower()
        return any(re.search(p, text_lower) for p in injection_patterns)

주요 솔루션 비교

솔루션 유형 주요 특징 배포 방식 가격
LiteLLM OSS Proxy 100+ 모델 통합, OpenAI 호환 API Self-hosted 무료 (Enterprise 유료)
Kong AI Gateway Commercial Kong 기반 플러그인, 엔터프라이즈 기능 Self-hosted / Cloud 유료
Cloudflare AI Gateway SaaS 엣지 캐싱, Analytics, 간편 설정 Cloud (Cloudflare) 무료 티어 + 종량제
OpenRouter SaaS 다중 모델 라우팅, 통합 API Cloud 모델별 종량제
Portkey SaaS + OSS Guardrails, 관측성, A/B 테스트 Cloud / Self-hosted 무료 티어 + 유료
AWS Bedrock Cloud AWS 네이티브, IAM 통합 Cloud (AWS) 종량제

기능 상세 비교

기능 LiteLLM Kong AI Cloudflare OpenRouter Portkey
OpenAI 호환 API O O O O O
Semantic Cache O X O X O
Rate Limiting O O O O O
Fallback O O X O O
PII 마스킹 X Plugin X X O
비용 추적 O O O O O
A/B 테스트 X X X X O
자체 호스팅 O O X X O
Streaming O O O O O
Guardrails X Plugin O X O

배포 패턴

On-Premise 배포

+----------------------------------------------------------+
|  Kubernetes Cluster                                       |
|                                                           |
|  +------------------+   +------------------+              |
|  | Gateway Pod (x3) |   | Redis            |              |
|  | - LiteLLM        |   | (Rate Limit     |              |
|  | - Sidecar Proxy  |   |  + Cache State)  |              |
|  +--------+---------+   +--------+---------+              |
|           |                       |                        |
|  +--------+-----------------------+---------+              |
|  |              Internal Network            |              |
|  +--------+-----------------------+---------+              |
|           |                       |                        |
|  +--------+---------+   +--------+---------+              |
|  | vLLM Pod (GPU)   |   | Postgres         |              |
|  | (Self-hosted LLM)|   | (Audit Log       |              |
|  +------------------+   |  + Key Store)     |              |
|                         +------------------+              |
+----------------------------------------------------------+
           |
    [External API Calls]
           |
    +------+------+
    | OpenAI API  |
    | Anthropic   |
    +-------------+

Cloud-Native 배포

[CloudFront / ALB]
        |
        v
[API Gateway (AWS)]
        |
        v
[Lambda / ECS - Gateway Logic]
        |
        +-------> [ElastiCache Redis]
        |
        +-------> [DynamoDB - Key Store]
        |
        +-------> [CloudWatch - Metrics]
        |
        v
[External LLM APIs]  +  [Bedrock - Self-hosted]

하이브리드 패턴

민감 데이터는 On-Premise 모델로, 일반 요청은 Cloud API로 분기.

                    Request
                       |
                       v
                [Content Classifier]
                       |
              +--------+--------+
              |                 |
        [Sensitive]       [Non-sensitive]
              |                 |
              v                 v
     +----------------+  +----------------+
     | On-Prem vLLM   |  | Cloud API      |
     | (No data egress)|  | (OpenAI, etc.) |
     +----------------+  +----------------+

비용 최적화 전략

1. 모델 티어링

요청 복잡도에 따라 적절한 모델을 선택하여 비용을 줄인다.

티어 모델 예시 입력 비용/1M tok 용도
Tier 1 (경량) GPT-4o-mini, Haiku $0.15 - $0.25 분류, 추출, 간단한 QA
Tier 2 (범용) GPT-4o, Sonnet $2.50 - $3.00 일반 대화, 분석
Tier 3 (고성능) GPT-4o, Opus $10.00 - $15.00 복잡 추론, 코딩

2. Prompt 최적화

class PromptOptimizer:
    """토큰 소비를 줄이는 프롬프트 최적화"""

    @staticmethod
    def compress_system_prompt(prompt: str, max_tokens: int = 500) -> str:
        """시스템 프롬프트 압축"""
        # 불필요한 반복/설명 제거
        lines = prompt.strip().split("\n")
        compressed = []
        seen = set()
        for line in lines:
            normalized = line.strip().lower()
            if normalized and normalized not in seen:
                seen.add(normalized)
                compressed.append(line)
        return "\n".join(compressed)

    @staticmethod
    def truncate_history(
        messages: list[dict],
        max_tokens: int = 4000,
        keep_system: bool = True
    ) -> list[dict]:
        """대화 이력 잘라내기 - 최근 메시지 우선"""
        system_msgs = []
        other_msgs = []

        for m in messages:
            if m["role"] == "system" and keep_system:
                system_msgs.append(m)
            else:
                other_msgs.append(m)

        result = list(system_msgs)
        token_count = sum(len(m["content"]) // 4 for m in result)

        for m in reversed(other_msgs):
            msg_tokens = len(m["content"]) // 4
            if token_count + msg_tokens > max_tokens:
                break
            result.insert(len(system_msgs), m)
            token_count += msg_tokens

        return result

3. 비용 모니터링 대시보드 쿼리

-- 일별 팀별 비용 집계
SELECT
    DATE(created_at) AS date,
    team_id,
    model,
    COUNT(*) AS request_count,
    SUM(input_tokens) AS total_input_tokens,
    SUM(output_tokens) AS total_output_tokens,
    SUM(cost_usd) AS total_cost_usd,
    AVG(latency_ms) AS avg_latency_ms
FROM llm_requests
WHERE created_at >= CURRENT_DATE - INTERVAL '30 days'
GROUP BY DATE(created_at), team_id, model
ORDER BY total_cost_usd DESC;

-- 캐시 히트율
SELECT
    DATE(created_at) AS date,
    COUNT(CASE WHEN cache_hit THEN 1 END)::float
        / NULLIF(COUNT(*), 0) AS cache_hit_rate,
    SUM(CASE WHEN cache_hit THEN estimated_cost_usd ELSE 0 END) AS cost_saved
FROM llm_requests
WHERE created_at >= CURRENT_DATE - INTERVAL '7 days'
GROUP BY DATE(created_at)
ORDER BY date;

실무 구성 예시: LiteLLM

# litellm_config.yaml
model_list:
  - model_name: gpt-4o
    litellm_params:
      model: openai/gpt-4o
      api_key: os.environ/OPENAI_API_KEY
      rpm: 60
      tpm: 100000

  - model_name: gpt-4o
    litellm_params:
      model: azure/gpt-4o-deployment
      api_key: os.environ/AZURE_API_KEY
      api_base: https://my-resource.openai.azure.com
      rpm: 120
      tpm: 200000

  - model_name: claude-sonnet
    litellm_params:
      model: anthropic/claude-3.5-sonnet-20241022
      api_key: os.environ/ANTHROPIC_API_KEY

  - model_name: local-llama
    litellm_params:
      model: openai/meta-llama/Llama-3-8B
      api_base: http://vllm-server:8000/v1
      api_key: dummy

router_settings:
  routing_strategy: least-busy
  num_retries: 3
  timeout: 60
  fallbacks:
    - gpt-4o: [claude-sonnet, local-llama]
    - claude-sonnet: [gpt-4o]
  allowed_fails: 3
  cooldown_time: 30

general_settings:
  master_key: os.environ/LITELLM_MASTER_KEY
  database_url: os.environ/DATABASE_URL
  cache: true
  cache_params:
    type: redis
    host: redis-server
    port: 6379
    ttl: 3600
# LiteLLM 실행
litellm --config litellm_config.yaml --port 4000

# 호출 (OpenAI 호환)
curl -X POST http://localhost:4000/v1/chat/completions \
  -H "Authorization: Bearer sk-litellm-master-key" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "gpt-4o",
    "messages": [{"role": "user", "content": "Hello"}]
  }'

관측성 (Observability)

핵심 메트릭

메트릭 설명 알림 임계값
Request Latency (p50/p95/p99) 응답 지연 시간 p95 > 5s
Error Rate 4xx/5xx 비율 > 5%
Token Throughput 분당 처리 토큰 TPM 한도 90%
Cache Hit Rate 캐시 적중률 < 10% (비효율)
Cost per Request 요청당 평균 비용 일일 예산 80%
Provider Health 제공자별 성공률 < 95%
TTFT Time to First Token > 2s

OpenTelemetry 연동

from opentelemetry import trace, metrics
from opentelemetry.sdk.trace import TracerProvider
from opentelemetry.sdk.metrics import MeterProvider

tracer = trace.get_tracer("llm-gateway")
meter = metrics.get_meter("llm-gateway")

request_counter = meter.create_counter(
    "llm.requests.total",
    description="Total LLM requests"
)
latency_histogram = meter.create_histogram(
    "llm.request.duration",
    unit="ms",
    description="Request latency distribution"
)
token_counter = meter.create_counter(
    "llm.tokens.total",
    description="Total tokens processed"
)

async def handle_request(request: dict) -> dict:
    with tracer.start_as_current_span("llm.request") as span:
        span.set_attribute("llm.model", request.get("model", ""))
        span.set_attribute("llm.provider", "openai")

        start = time.monotonic()
        try:
            response = await call_provider(request)

            duration = (time.monotonic() - start) * 1000
            latency_histogram.record(duration, {"model": request["model"]})
            request_counter.add(1, {"model": request["model"], "status": "success"})
            token_counter.add(
                response["usage"]["total_tokens"],
                {"model": request["model"]}
            )

            span.set_attribute("llm.tokens.input", response["usage"]["prompt_tokens"])
            span.set_attribute("llm.tokens.output", response["usage"]["completion_tokens"])

            return response
        except Exception as e:
            request_counter.add(1, {"model": request["model"], "status": "error"})
            span.set_status(trace.StatusCode.ERROR, str(e))
            raise

참고 자료

자료 URL
LiteLLM 공식 문서 https://docs.litellm.ai/
Portkey AI Gateway https://portkey.ai/docs
Cloudflare AI Gateway https://developers.cloudflare.com/ai-gateway/
Kong AI Gateway https://docs.konghq.com/gateway/latest/ai-gateway/
OpenTelemetry LLM SIG https://github.com/open-telemetry/semantic-conventions/tree/main/docs/gen-ai
LLM Gateway Design Patterns (블로그) https://martinfowler.com/