12. 일반 Agent와 SandboxAgent 비교
이 장에서 처음 나오는 말4개
A/B fixture- 비교하려는 한 변수 외에는 model·instructions·요청을 같게 둔 두 시험 대상이다.
suspend- idle actor의 상태를 snapshot으로 내리고 worker slot을 반납하는 전환이다.
rehydration- snapshot을 worker에 복원해 actor 실행을 다시 시작하는 과정이다.
capacity saturation- 동시 actor 수가 WorkerPool slot을 넘어 대기·실패가 생기는 상태다.
11장에서 같은 cluster에 Substrate를 추가했다. 이제 Agent와 SandboxAgent의 이름만 다르게 하고
runtime·modelConfig·instructions는 같게 둔다. 첫 비교에서는 MCP tool을 빼서 snapshot lifecycle과
tool egress policy를 한 번에 바꾸지 않는다.
비교 질문과 통제 변수
섹션 제목: “비교 질문과 통제 변수”| 고정할 것 | 바꿀 것 | 측정할 것 |
|---|---|---|
| cluster·namespace·kagent version | Agent ↔ SandboxAgent | resource Ready까지 걸린 시간 |
| Go runtime·ModelConfig·system message | Deployment ↔ Substrate actor | 최초·복원 invoke latency |
| controller A2A route·request text | 상주 ↔ suspendable | idle Pod·worker slot 수 |
| backend client code·timeout | pod lifecycle ↔ snapshot lifecycle | saturation·worker 장애 결과 |
답변 품질이나 token 수는 model 변동이 섞이므로 주된 판정값으로 쓰지 않는다. 각 구간을 최소 5회 반복하고
p50·최댓값·실패 수를 함께 남긴다. 한 번 빨랐던 값만 적지 않는다.
같은 선언 두 개 만들기
섹션 제목: “같은 선언 두 개 만들기”일반 Agent를 먼저 만든다.
kubectl apply -f - <<'EOF'apiVersion: kagent.dev/v1alpha2kind: Agentmetadata: name: resident-echo namespace: kagent labels: study.upggu.com/fixture: substrate-abspec: type: Declarative description: Resident side of the Substrate A/B fixture. declarative: runtime: go modelConfig: default-model-config systemMessage: |- You are the A/B runtime fixture. Answer in one short sentence and do not use tools.EOF같은 선언을 SandboxAgent와 kagent-default WorkerPool에 연결한다.
kubectl apply -f - <<'EOF'apiVersion: kagent.dev/v1alpha2kind: SandboxAgentmetadata: name: sandbox-echo namespace: kagent labels: study.upggu.com/fixture: substrate-abspec: type: Declarative description: Substrate side of the A/B fixture. declarative: runtime: go modelConfig: default-model-config systemMessage: |- You are the A/B runtime fixture. Answer in one short sentence and do not use tools. substrate: workerPoolRef: name: kagent-defaultEOF공식 0.9.9 walkthrough 기준으로 Substrate Declarative 경로는 Go runtime을 사용한다. SandboxAgent API가
다른 실행 유형도 표현한다고 해서 이 pinned 실습의 지원 조합을 넓혀 추측하지 않는다.
준비 시간과 실행 형태 확인
섹션 제목: “준비 시간과 실행 형태 확인”-
적용 시각부터 condition까지 각각 잰다
터미널 창 time kubectl -n kagent wait agent/resident-echo \--for=condition=Ready --timeout=3mtime kubectl -n kagent wait sandboxagent/sandbox-echo \--for=condition=Ready --timeout=5m첫
SandboxAgent는 golden snapshot을 만드는 데 공식 walkthrough 기준 약 60~90초가 걸릴 수 있다. 이미 Ready인 resource에서wait를 다시 실행한 시간은 생성 시간이 아니므로, 재측정하려면 새 이름을 쓴다. -
resource와 workload를 나란히 본다
터미널 창 kubectl -n kagent get agent resident-echo -o widekubectl -n kagent get sandboxagent sandbox-echo -o widekubectl -n kagent get deploykubectl -n kagent get workerpool kagent-defaultkubectl -n kagent get pods -o wideresident-echo쪽에는 전용 Deployment가 생기고,sandbox-echo는 전용 상주 Deployment 대신 WorkerPool의 actor를 사용해야 한다. 이름 문자열만 찾지 말고 owner reference·resource status·UI를 함께 본다. -
UI에서 actor lifecycle을 본다
http://localhost:8001의 View → Substrate에서sandbox-echo를 연다. 요청을 처리한 뒤 actor가Suspended가 되고 worker가 다시 idle로 돌아오는 화면을 캡처하거나 관찰 기록에 적는다.
같은 backend A2A client로 호출하기
섹션 제목: “같은 backend A2A client로 호출하기”controller port-forward를 유지한 상태에서 /tmp/kagent-backend-probe/compare.mts를 만든다. 이 code는
resource kind를 알지 못한다. 이름만 받아 두 대상의 같은 Agent Card와 message/send route를 호출한다.
import { randomUUID } from 'node:crypto';
const target = process.env.TARGET ?? 'resident-echo';const burst = Number(process.env.BURST ?? '1');const base = `http://localhost:8083/api/a2a/kagent/${target}/`;
async function invoke(index: number) { const id = randomUUID(); const started = performance.now(); const response = await fetch(base, { method: 'POST', headers: { 'content-type': 'application/json' }, body: JSON.stringify({ jsonrpc: '2.0', id, method: 'message/send', params: { id, message: { role: 'user', parts: [{ kind: 'text', text: `Reply with: runtime probe ${index}` }], }, }, }), signal: AbortSignal.timeout(120_000), }); const payload = await response.json() as { error?: { code: number; message: string }; result?: { status?: { state?: string } }; }; if (!response.ok) throw new Error(`HTTP ${response.status}`); if (payload.error) throw new Error(`${payload.error.code}: ${payload.error.message}`); return { index, elapsedMs: Math.round(performance.now() - started), state: payload.result?.status?.state ?? 'unknown', };}
const cardResponse = await fetch(new URL('.well-known/agent.json', base));if (!cardResponse.ok) throw new Error(`Agent Card HTTP ${cardResponse.status}`);
const results = await Promise.all( Array.from({ length: burst }, (_, index) => invoke(index + 1)),);console.log(JSON.stringify({ target, burst, results }, null, 2));먼저 순차 호출을 반복한다.
cd /tmp/kagent-backend-probeTARGET=resident-echo npx tsx compare.mtsTARGET=sandbox-echo npx tsx compare.mtsUI에서 sandbox-echo actor가 Suspended가 된 것을 확인한 다음 같은 command를 다시 실행한다. 다음 세 값을
구분해 기록한다.
- 상주 Agent의 warm invoke
- SandboxAgent의 첫 session invoke
Suspended에서 복원된 SandboxAgent invoke
두 대상 모두 같은 /api/a2a/kagent/{name}/ 계약으로 성공해야 한다. backend에 SandboxAgent 전용 chat
client가 필요해졌다면 adapter 경계가 잘못 새고 있는 것이다.
WorkerPool 포화 보기
섹션 제목: “WorkerPool 포화 보기”replica 1에서 세 요청을 겹쳐 보낸다.
cd /tmp/kagent-backend-probeTARGET=sandbox-echo BURST=3 npx tsx compare.mtskubectl -n kagent get workerpool kagent-default -o yamlkubectl -n kagent get pods -o widekubectl -n ate-system get pods -o wide성공 수·대기 시간·timeout과 관련 Event를 남긴다. 다음으로 live CR만 replica 2로 늘려 같은 시험을 반복한다.
kubectl -n kagent scale workerpool kagent-default --replicas=2kubectl -n kagent get workerpool kagent-default -o yamlkubectl -n kagent get pods -o wideTARGET=sandbox-echo BURST=3 npx tsx compare.mts두 번째 worker Pod가 Ready이고 WorkerPool status가 원하는 replica를 반영한 뒤 burst를 실행한다. 설치한
Substrate CRD가 제공하지 않는 condition 이름을 추측해 kubectl wait에 넣지 않는다.
kubectl scale은 다음 Helm upgrade에서 chart value 1로 돌아갈 수 있는 임시 변경이다. 비교가 끝나면 release의
desired state와 맞춘다.
kubectl -n kagent scale workerpool kagent-default --replicas=1production sizing은 등록 Agent 수가 아니라 동시 active session 수·요청 시간·Harness slot을 기준으로 한다.
Worker 하나를 잃어 보기
섹션 제목: “Worker 하나를 잃어 보기”이 절차는 kind-kagent-lab 전용 failure injection이다. 먼저 kubectl -n kagent get pods -o wide에서
WorkerPool이 소유한 worker Pod 이름 하나를 확인한다. 다른 Pod를 추측해서 지우지 말고, 그 정확한 이름을
두 번째 terminal의 kubectl -n kagent delete pod <확인한-worker-pod-이름>에 넣는다.
첫 terminal에서는 BURST=3 호출을 실행하고, 다음을 기록한다.
- 진행 중 요청이 성공·재시도·실패 중 어디로 끝나는가
- WorkerPool이 replacement Pod를 Ready로 만드는 데 걸린 시간
sandbox-echo가 다시 호출 가능한가- 기존
resident-echo와backend-reader에는 영향이 없는가
이 결과는 gVisor 격리 자체보다 shared worker failure domain을 보여 준다. 호출 재시도 정책은 controller에 맡겼다고 가정하지 말고 backend의 timeout·idempotency 설계로 가져간다.
MCP tool은 egress policy와 함께 두 번째로 붙인다
섹션 제목: “MCP tool은 egress policy와 함께 두 번째로 붙인다”SandboxAgent.spec.sandbox.network.allowedDomains가 비었거나 없으면 sandbox execution의 outbound는 기본
거부다. 그래서 lab-reader의 MCP 설정을 그대로 복사해 첫 A/B 결과를 흐리지 않는다.
tool 비교가 필요하면 다음 순서를 지킨다.
ModelConfig와RemoteMCPServer에서 실제 HTTPS·Service DNS host를 확인한다.- 그 host만
spec.sandbox.network.allowedDomains에 명시한 별도sandbox-reader를 만든다. lab-reader와 같은toolNames를 넣고 허용 조회와 금지 변경 요청을 반복한다.- 허용 목록 밖 DNS·IP 호출이 차단되는지 runtime log와 NetworkPolicy에서 확인한다.
wildcard나 전체 인터넷 허용으로 빨리 통과시키지 않는다. 모델 endpoint와 MCP endpoint가 달라 어느 host가 빠졌는지 식별하는 과정 자체가 온프렘 egress 설계의 증거다.
비교 기록표
섹션 제목: “비교 기록표”| 항목 | 일반 Agent | SandboxAgent | 판정 |
|---|---|---|---|
| resource Ready | |||
| 최초 invoke p50 / max / 실패 | |||
| warm 또는 restore invoke p50 / max / 실패 | |||
| idle 전용 Pod 수·memory | |||
| burst 3, worker 1 결과 | |||
| burst 3, worker 2 결과 | |||
| worker Pod 교체 시간·요청 영향 | |||
| 새 운영 구성 요소·로그 위치 | Deployment | Substrate control/data/storage |
측정 host·cluster resource·version pair·ModelConfig alias·시각을 표와 함께 남긴다. 비용 비교에서 Substrate control plane·Valkey·snapshot storage의 고정비를 빼지 않는다.
완료 체크
섹션 제목: “완료 체크”- 같은 model·instructions·A2A client로
Agent와SandboxAgent를 비교했다. SandboxAgent의 golden snapshot 준비,Suspended, restore를 각각 확인했다.- WorkerPool replica 1·2의 burst 결과와 worker Pod failure 결과를 기록했다.
- MCP 비교를 sandbox egress allowlist 검증과 묶어 별도 단계로 분리했다.
- “등록 수가 많다”가 아니라 측정된 idle 절감·복원 지연·운영비로 Substrate 가치를 설명할 수 있다.
참고 자료
섹션 제목: “참고 자료”- kagent Agent Substrate walkthrough —
SandboxAgentmanifest, golden snapshot과Suspended관찰 - kagent Agent Substrate 개념 — actor·WorkerPool·snapshot/restore lifecycle
- kagent API reference —
SandboxAgent,SandboxConfig,allowedDomainsschema