TL;DR
이번 기간에는 Muse Spark 1.3과 Gemini 3.8 Flash가 Agentic Work와 Coding 성능을 앞세워 연이어 출시되면서 모델 경쟁의 기준이 장기 작업 수행력과 비용으로 이동했습니다. Muse Spark 1.3은 Artificial Analysis Intelligence Index에서 xhigh 61점, max 62점을 기록했고, xhigh의 작업당 비용은 $0.55로 집계됐습니다. Gemini 3.8 Flash는 DeepSWE 1.1에서 73.7%를 기록했으며, Cyber 변형은 CyberGym 86.2%와 CWE-Bench 47.2%를 기록했습니다. 동시에 Linear와 Cursor는 오류 수정과 인프라 접근을 자동화하는 실행형 도구를 확장했고, white-box 모니터링과 오픈 웨이트 모델의 생태계 확장이 별도 흐름으로 부상했습니다.
𝕏 실시간 트렌드 토픽
🔥 Muse Spark 1.3의 에이전트 성능과 비용 경쟁포스트 13
Muse Spark 1.3이 Agentic Work와 Coding 평가에서 상위권에 진입하면서 성능뿐 아니라 작업당 비용과 추론 토큰량이 함께 비교됐습니다.
세부 내용 보기
- Meta가 Muse Spark 1.3을 Muse Code와 Meta Model API에 출시했고, 한 스레드에서 더 긴 작업을 이어가며 명확화 질문, 중단 신호, 중요 작업 전 확인을 수행하도록 개선했습니다. 내부 비교에서는 Muse Spark 1.2보다 도구 호출이 약 20% 줄고 토큰 사용량이 약 25% 감소했습니다.
- Artificial Analysis Intelligence Index에서 Muse Spark 1.3 (xhigh)은 61점으로 Muse Spark 1.2의 57점보다 4점 높았고, Muse Spark 1.3 (max)은 62점을 기록했습니다. xhigh는 Tau3-Bench Banking 47%, Terminal-Bench 2.1 85%, GDPval-AA v2 1,709 Elo를 기록했으며, max는 Tau3-Bench Banking 52%, GDPval-AA v2 1,754 Elo로 더 높은 추론량을 사용했습니다.
- xhigh의 작업당 비용은 Meta의 $1.25/$4.25 per 1M token 가격과 캐시 입력 $0.15 per 1M 조건에서 $0.55로 집계됐습니다. GPT-5.6 Sol (max)의 $0.95, Grok 4.6 (high)의 $0.94보다 낮지만, Muse Spark 1.2의 $0.40보다 오른 수치이며 입력 토큰이 약 57% 늘어난 것이 주요 원인으로 제시됐습니다.
원문 트윗 2개 보기
Artificial Analysis
Meta has released Muse Spark 1.3, their fourth Muse Spark model release in five months. Muse Spark 1.3 (max), which is in limited preview for Meta’s partners, scores 62 on the Artificial Analysis Intelligence Index, behind only Claude Fable 5.1 and Claude Opus 5. The variant available now, Muse Spark 1.3 (xhigh), scores 61 and ties with GPT-5.6 Sol (max) and Grok 4.6 (high). Both variants’ gains come primarily from improvements in agentic work and scientific capabilities Muse Spark 1.3 (xhigh) enters the Artificial Analysis Intelligence Index at 61, up 4 points from Muse Spark 1.2 (57, August) and 8 points from Muse Spark 1.1 (53, July). It enters tied with GPT-5.6 Sol (max), Grok 4.6 (high), and Claude Opus 5 (high), and behind Claude Fable 5.1 (max, 66), Claude Opus 5 (max, 63), and Claude Fable 5 (max, 62) Muse Spark 1.3 (max), which is in a limited preview stage, lands at 62. This higher index score is enabled by gains vs. Muse Spark 1.3 (xhigh) in Tau3-Bench Banking (52% vs. 47%) and GDPval-AA v2 (1,754 Elo vs. 1,709). Muse Spark 1.3 (max) is second only to Claude’s Fable and Opus variants in total score Congratulations to @AIatMeta , @finkd , and @alexandr_wang on the release! Key Takeaways: ➤ Continued improvement on agentic knowledge work tasks. At the launch of Muse Spark 1.2, we noted its significant gains in agentic knowledge work performance vs. Muse Spark 1.1. The latest iteration continues this trend, with Muse Spark 1.3 (xhigh) demonstrating a notable 12-point gain vs. Muse Spark 1.2 in Tau3-Bench Banking (35% to 47%), a 5-point gain in Terminal-Bench 2.1 (80% to 85%), and a new GDPval-AA v2 Elo of 1709 against its predecessor’s 1615. Muse Spark 1.3 (max) improves further on Tau3-Bench Banking (52%) and GDPval-AA v2 (1,754 Elo). This Tau3-Bench Banking score is #1 among all models. Muse Spark 1.3 (max) achieves these higher agentic work scores by using more turns and total reasoning tokens, reasoning 62% more on GDPval-AA v2 and 28% more on Tau3-Bench Banking compared to Muse Spark 1.3 (xhigh) ➤ The lowest cost per task for any model at 59+ on the Artificial Analysis Intelligence Index. Muse Spark 1.3 (xhigh) costs $0.55 per Intelligence Index task at Meta's unchanged $1.25/$4.25 per 1M token pricing ($0.15 for cached input), with its peers GPT-5.6 Sol (max) and Grok 4.6 (high) costing $0.95 and $0.94 respectively, a 70%+ premium. This places Muse Spark 1.3 (xhigh) on the Pareto frontier for Intelligence vs. Cost per Task. Its cost per task is higher than Muse Spark 1.2 ($0.40 per task), driven by ~57% more input tokens per task on agentic evaluations, with output tokens up only ~8%. Pricing for Muse Spark 1.3 (max) is not yet publicly available ➤ Scientific Reasoning results rose across the board, led by CritPt. CritPt was the standout non-agentic score gain vs. Muse Spark 1.2, with a material +8 points for the xhigh variant (18% to 26%), and GPQA Diamond achieved +4 points (90% to 94%), while Humanity’s Last Exam and SciCode each gained a more modest 2-3 points (45% to 47% and 56% to 59%, respectively). Muse Spark 1.3 (max) achieved roughly similar scores to the xhigh variant, gaining 2 points in Humanity’s Last Exam, tying on GPQA Diamond, and losing a point on CritPt vs. Muse Spark 1.3 (xhigh) ➤ Minor regressions in only two evaluations. Both Muse Spark 1.3 (xhigh) and Muse Spark 1.3 (max) dropped 4 points in AA-LCR (83% to 79%) when compared to Muse Spark 1.2, and AA-Omniscience (Accuracy) fell 3 points for xhigh and 1 point for max. The drops in AA-Omniscience (Accuracy) are due to a higher abstention rate (not answering questions when unsure), which also lowered the hallucination rate for Muse Spark 1.3 (xhigh) Other model details (xhigh variant): ➤ Context window: 1M tokens, unchanged from Muse Spark 1.2 ➤ Pricing: unchanged from Muse Spark 1.2: $1.25/$4.25 per 1M input/output tokens, with cache hits discounted to $0.15 per 1M ➤ Input modalities: text, image, video ➤ Availability: Meta's first-party API and Muse Code

AI at Meta
We’re excited to release Muse Spark 1.3 with improved performance on agentic and coding tasks, and a focus on real-world usability. Key capabilities: → Sustains longer-horizon work across multiple workflows in a single thread → More actively collaborates with users: it asks clarifying questions, flags when it's stuck, confirms before consequential actions → Better calibrated on its own limits instead of hallucinating outcomes → ~20% fewer tool calls and ~25% fewer tokens vs. Muse Spark 1.2 in internal comparisons
🔥 Gemini 3.8 Flash의 코딩·보안 특화 확장포스트 10
Gemini 3.8 Flash가 장기 Coding과 다단계 추론을 겨냥해 출시됐고, Cyber 변형은 취약점 탐지와 패치 평가를 별도로 끌어올렸습니다.
세부 내용 보기
- Google은 Gemini 3.8 Flash를 Gemini 3.7 Flash 이후 Software Engineering, Agentic Tasks, Multi-Step Reasoning을 개선한 모델로 출시했습니다. Gemini App, Cursor, Gemini API, Android Studio, Antigravity, StitchbyGoogle 등 여러 경로에 제공되며 연말까지 입력 $0.75/1M, 출력 $3.75/1M 토큰의 도입 가격이 적용됩니다.
- Gemini 3.8 Flash는 DeepSWE 1.1에서 73.7%를 기록했고, Google은 복잡한 프로젝트를 탐색하며 코드를 작성하는 작업과 Native Video Understanding을 결합한 에이전트 루프를 함께 제시했습니다. Effort Controls로 작업별 추론량을 조절해 난이도와 토큰 지출을 맞추는 구조입니다.
- Gemini 3.8 Flash Cyber는 CyberGym 86.2%, CWE-Bench 패치 47.2%, 내부 20개 프로그래밍 언어 취약점 탐지 성공률 70% 이상을 기록했습니다. Chrome 보안 팀에서는 대형 상용 모델보다 취약점의 올바른 패치를 2.6배 더 많이 만들었고, Google Cloud 팀은 일반적으로 수개월 걸리는 핵심 취약점을 2시간 이내에 찾았다고 밝혔습니다.
원문 트윗 2개 보기
Today we’re introducing a new 3.8 Flash model, our 3rd Flash release in just 6 wks. It delivers significant leaps from 3.7 Flash across software engineering, agentic tasks, and multi-step reasoning. On DeepSWE v1.1, it outperforms most larger frontier models in autonomously

Sundar Pichai
We’re also introducing Gemini 3.8 Flash Cyber, our most capable cybersecurity model. It shows frontier-level performance in discovering vulnerabilities and patching them at scale, with Flash-level speed & pricing. That includes achieving 86.2% on the important CyberGym industry benchmark, plus 47.2% on CWE-Bench for patching. We saw a 70%+ success rate in discovering vulnerabilities across 20 programming languages on our internal benchmark.
📈 개발 환경에 연결된 실행형 코딩 에이전트포스트 4
Coding Agent가 코드 작성만 수행하는 단계를 넘어 오류 추적, 패치 배포, 사내 서비스 접근까지 이어지는 실행 루프로 확장됐습니다.
세부 내용 보기
- Linear의 버그 자동 수정 루프는 Datadog와 Sentry에서 이슈를 조사한 뒤 Linear의 Coding Agent에 수정을 할당하는 흐름으로 구성됐습니다. Triage 과정에 이 연결을 넣어 최근 30일 동안 300개가 넘는 버그를 수정했습니다.
- Claude Cowork와 Claude Code는 사용자가 데스크톱 작업을 맡기면 백그라운드에서 클릭, 입력, 앱 실행을 수행합니다. Cursor Cloud Agents는 수요에 따라 자동 확장되는 머신 풀과 사내 서비스·특수 하드웨어에 대한 접근을 제공하면서 Agent Loop 자체는 Cursor에 유지합니다.
- GPT-5.6 Sol은 PloyAI 팀의 복잡한 Marketing Campaign 운영에서 Subagents를 조율하는 데 활용됐습니다. 각 사례는 모델의 응답 생성보다 외부 도구 연결, 작업 배분, 실제 환경의 변경 실행을 중심에 둔 구조입니다.
원문 트윗 2개 보기

Karri Saarinen
Our bug autofix loop in @linear has fixed 300+ bugs in the last 30 days. The team set it up as part of triage: it connects to Datadog and Sentry to investigate issues, then assigns them to Linear’s coding agent to fix

Cursor
You can now run Cursor cloud agents on your infrastructure, including pools of machines that automatically scale with demand. This lets you give agents access to internal services or specialized hardware, while the agent loop stays in Cursor.
➖ CoT 의존도를 둘러싼 AI 위험 모니터링포스트 1
White-Box Techniques가 모델의 비공개 계획을 포착하는 대안으로 연구되는 가운데, 현재는 CoT보다 낮은 모니터링 신뢰도가 핵심 제약으로 남아 있습니다.
세부 내용 보기
- Jack W. Lindsey는 White-Box Techniques가 현재 일부 비언어화된 생각과 계획을 포착하지만 CoT만큼 안정적으로 감시하지는 못한다고 밝혔습니다. 최근 몇 달 사이 Production Monitoring에 활용되는 Activation Decoding 기법이 발표될 만큼 진전 속도는 빠르지만, 1년 안에 CoT 수준에 도달할지는 불확실하다는 평가입니다.
- 그가 언급한 방어 방식은 CoT 모니터링을 당장 약화시키지 않으면서 White-Box Techniques를 개발하고 스트레스 테스트하는 Defense in Depth입니다. 기법이 작동하더라도 비용이 높거나 개발이 어려우면 널리 적용되지 않을 수 있다는 점도 함께 제기됐습니다.
- 현재 접근은 Mechanistic Interpretability보다 모델의 내부 표현을 읽는 Mind Reading에 가깝고, 내부 생각을 읽더라도 그 생각을 만든 메커니즘까지 파악하지는 못한다는 구분이 제시됐습니다.
📈 오픈 웨이트 AI 생태계의 중심 이동포스트 1
NVIDIA와 대형 기술 기업의 저장소 확장이 이어지는 가운데 OpenRouter에서 오픈 모델이 차지하는 토큰 비중과 폐쇄형 모델과의 성능 근접도가 함께 높아졌습니다.
세부 내용 보기
- NVIDIA는 지난 12개월 동안 Hugging Face의 기업 오픈소스 AI 저장소를 500개 이상 추가해 Alibaba Cloud, Hugging Face, Tencent를 앞섰습니다. Nemotron, Cosmos, GR00T와 로보틱스 데이터셋 및 최적화 모델이 증가분에 포함됐습니다.
- OpenRouter에서는 오픈 모델이 2025년 대부분 약 25%의 라우팅 토큰을 차지했지만 2026년 봄 50%를 넘었고 5월 이후 다수를 유지했습니다. 동시에 최고 오픈 웨이트 모델과 폐쇄형 모델의 벤치마크 성능 격차가 수년에서 수개월 수준으로 줄었다고 집계됐습니다.
- 해당 흐름의 주요 모델로 Kimi K3, GLM 5.3, Qwen 3.8 Max, DeepSeek v4 Pro가 거론됐고, Qwen과 DeepSeek는 Hugging Face 다운로드에서도 앞섰습니다. 접근성이 높아질수록 기업이 Agent, Robot, 특화 AI 시스템을 직접 구축할 여지가 커지고 기술 주권을 둘러싼 선택지도 넓어집니다.
용어 해설
- 에이전트형 작업(Agentic Work)
- — 모델이 한 번의 답변에 그치지 않고 여러 단계의 추론과 도구 호출을 이어가며 업무를 수행하는 방식입니다. 긴 작업에서는 계획 수립, 실행, 오류 대응을 반복해 최종 결과를 완성합니다.
- 장기 작업(Long-Horizon Task)
- — 여러 단계와 긴 실행 시간이 필요한 작업을 뜻합니다. 모델이 하나의 대화 흐름 안에서 맥락을 유지하고 중간 상태를 관리해야 하므로 단순 질의응답보다 지속적인 계획과 자기 점검이 중요합니다.
- 추론 노력 제어(Effort Controls)
- — 작업 난이도에 따라 모델이 사용하는 추론량을 조절하는 기능입니다. 더 많은 토큰을 투입해 복잡한 문제를 처리하거나, 간단한 요청에서는 계산량을 낮춰 비용과 지연을 줄이는 방식으로 작동합니다.
- 화이트박스 모니터링(White-Box Monitoring)
- — 모델 내부 활성값이나 표현을 관찰해 외부 출력만으로 드러나지 않는 계획과 위험 신호를 찾는 감시 방식입니다. CoT 의존도를 낮추는 대안으로 연구되지만 현재는 탐지 신뢰도와 적용 비용이 쟁점입니다.
- 오픈 웨이트 모델(Open-Weight Model)
- — 모델 가중치를 공개해 기업과 연구자가 직접 배포하거나 수정할 수 있는 모델입니다. 폐쇄형 API에만 의존하지 않고 에이전트, 로봇, 특화 시스템을 구축할 수 있는 기반으로 활용됩니다.
AI 요약 · 북마크 · 개인 피드 설정 — 무료
출처 · 인용 안내
인용 시 "요약 출처: AI Trends (aitrends.kr)"를 표기하고, 사실 확인은 원문 보기 기준으로 진행해 주세요. 자세한 기준은 운영 정책을 참고해 주세요.