TL;DR
Supervised Fine-Tuning의 성능 한계는 학습 인프라보다 데이터 준비 품질에서 먼저 결정됩니다. Continued Pre-training은 도메인 지식을 넓히고 SFT는 응답 행동을 바꾸며 RFT는 보상 신호로 최적화하므로, foundation model에 필요한 지식이 이미 있다면 SFT와 RFT를 순차 적용하는 구성이 적합합니다. 학습 전에는 정답성·다양성·일관성·중복·안전성을 점검하고, 추론 시 사용할 system prompt와 Chat Template을 JSONL 데이터에 동일하게 반영해야 하며, reasoning trace와 tool calling도 각 역할과 필드 규칙에 맞춰 구성해야 합니다. 데이터의 10–20%를 운영 분포를 대표하는 평가 세트로 분리하고 수정하지 않은 모델의 baseline과 비교해야 SFT의 개선과 overfitting을 구분할 수 있습니다.
섹션별 상세
{
"schemaVersion": "bedrock-conversation-2024",
"system": [{"text": "You are a helpful coding assistant."}],
"messages": [
{
"role": "user",
"content": [{"text": "Write a Python function to check if a string is a palindrome."}]
},
{
"role": "assistant",
"content": [
{"text": "def is_palindrome(s):
cleaned = s.lower().replace(' ', '')
return cleaned == cleaned[::-1]"}
]
}
]
}Amazon Nova의 Converse API 형식으로 한 대화 예제를 하나의 JSON 객체에 담는 구조입니다.
{
"schemaVersion": "bedrock-conversation-2024",
"system": [{"text": "You are a financial analyst. Provide data-driven answers with supporting calculations."}],
"messages": [
{
"role": "user",
"content": [{"text": "Calculate YoY revenue growth. 2024: $4.2M, 2025: $5.1M"}]
},
{
"role": "assistant",
"content": [
{
"reasoningContent": {
"reasoningText": {
"text": "The user asks for year-over-year revenue growth. I need to calculate the percentage change: (new - old) / old x 100. That gives (5.1 - 4.2) / 4.2 x 100 = 21.43%."
}
}
},
{"text": "YoY revenue growth is approximately 21.4 percent: ($5.1M - $4.2M) / $4.2M x 100."}
]
}
]
}reasoningContent 필드에 계산 과정을 넣고 최종 응답과 연결하는 추론 Trace 예제입니다.
{
"schemaVersion": "bedrock-conversation-2024",
"system": [{"text": "You are an expert in composing function calls."}],
"toolConfig": {
"tools": [
{
"toolSpec": {
"name": "getItemCost",
"description": "Retrieve the cost of an item from the catalog",
"inputSchema": {
"json": {
"type": "object",
"properties": {
"item_id": {
"type": "string",
"description": "The ASIN of item to retrieve cost for"
}
},
"required": ["item_id"]
}
}
}
}
]
},
"messages": [
{
"role": "user",
"content": [{"text": "How much does item id-456 cost?"}]
},
{
"role": "assistant",
"content": [
{
"toolUse": {
"toolUseId": "getItemCost_0",
"name": "getItemCost",
"input": {"item_id": "id-456"}
}
}
]
},
{
"role": "user",
"content": [
{
"toolResult": {
"toolUseId": "getItemCost_0",
"content": [
{"text": "{"name": "getItemCost", "results": {"cost": "$29.99"}}"}
]
}
}
]
},
{
"role": "assistant",
"content": [
{"text": "Item id-456 costs $29.99."}
]
}
]
}toolUse와 toolResult를 서로 다른 역할에 배치하고 동일한 toolUseId로 호출과 결과를 연결하는 형식입니다.
{
"schemaVersion": "bedrock-conversation-2024",
"system": [{"text": "You are a document analysis assistant."}],
"messages": [
{
"role": "user",
"content": [
{
"document": {
"format": "pdf",
"name": "quarterly_report",
"source": {"s3Location": {"uri": "s3://<your-bucket-name>/report.pdf"}}
}
},
{"text": "Summarize the key findings from this quarterly report."}
]
},
{
"role": "assistant",
"content": [
{"text": "The quarterly report highlights three key findings..."}
]
}
]
}Amazon S3 위치의 PDF 문서와 텍스트 요청을 함께 넣어 multimodal 학습 사례를 구성합니다.
용어 해설
- 지도 Fine-tuning(Supervised Fine-Tuning)
- — 이미 알고 있는 지식을 새로 주입하기보다 입력과 정답의 curated pair를 반복 학습해 모델의 응답 방식을 바꾸는 방법입니다. 지시 준수, 출력 형식, 말투, 구조화된 응답을 조정하는 데 쓰입니다.
- Continued Pre-training
- — 대규모 비정형 도메인 텍스트를 추가로 학습해 foundation model의 용어, 개념, 데이터 패턴에 대한 지식 기반을 넓히는 방식입니다. 기본 모델이 특정 분야의 핵심 지식을 모를 때 필요합니다.
- 강화 Fine-tuning(Reinforcement Fine-tuning)
- — 명시적인 모범 답안 대신 보상 신호를 사용해 모델의 행동을 최적화하는 방식입니다. 출력 품질을 프로그램으로 평가할 수 있지만 대규모로 추론 과정을 직접 작성하기 어려운 작업에 적합합니다.
- 의미적 커버리지(Semantic Coverage)
- — Fine-tuning 데이터가 얼마나 다양한 작업 영역과 사용자 표현 방식을 포함하는지를 나타내는 기준입니다. 실제 요청의 의도와 표현 범위를 충분히 포함해야 특정 사례에만 맞춘 모델을 피할 수 있습니다.
- 정보 깊이(Information Depth)
- — 개별 학습 예제가 담고 있는 정보와 문제 해결 과정의 풍부함을 뜻합니다. 단순한 사례만 모으기보다 복잡한 문제, 여러 단계의 처리, 적절한 답변 근거를 포함해야 일반화에 도움이 됩니다.
- Chat Template
- — 모델이 system, user, assistant 메시지를 구분하기 위해 사용하는 정확한 토큰 배열과 구분자 형식입니다. 학습 데이터와 추론 시 구조가 다르면 모델이 system prompt를 무시하거나 잘못된 형식의 출력을 만들 수 있습니다.
- 추론 Trace(Reasoning Trace)
- — 문제의 입력에서 최종 답변에 이르는 중간 사고 단계를 기록한 학습 데이터입니다. 단계가 실제 답변을 뒷받침하고 문제 난이도에 비례해야 하며, reasoning 기능을 사용하는 학습과 추론 환경에서 일관되게 다뤄야 합니다.
기술
- Amazon Nova
- Amazon Bedrock
- Amazon SageMaker HyperPod
- Llama Guard
- Hugging Face Transformers
- Amazon S3
활용 사례
- 출력 schema를 준수하는 구조화된 응답
- 도메인 분류 taxonomy에 맞춘 분류
- 지정된 말투를 유지하는 대화 시스템
- Tool calling과 function calling
- PDF·이미지·비디오 기반 multimodal 이해
- 다중 turn 고객 지원 대화
AI 요약 · 북마크 · 개인 피드 설정 — 무료
출처 · 인용 안내
인용 시 "요약 출처: AI Trends (aitrends.kr)"를 표기하고, 사실 확인은 원문 보기 기준으로 진행해 주세요. 자세한 기준은 운영 정책을 참고해 주세요.