TL;DR
이번 기간의 중심에는 GPT-6 Astra 출시와 컴퓨터 사용·소프트웨어 엔지니어링·시각 이해 능력에 관한 평가가 놓였습니다. ARC-AGI-3와 Terminal-Bench-Science에서 높은 점수가 공유됐지만, 표준 harness와 연속 대화 harness처럼 실행 조건에 따라 결과와 비용이 달라졌습니다. 동시에 Martian의 모델 라우팅 연구는 하나의 최고 모델보다 작업별 선택이 오류율과 비용을 낮출 수 있음을 수치로 제시했고, Claude Code Function Hooks와 vLLM의 AgentX는 에이전트 실행을 확장하고 측정하는 방향을 드러냈습니다. Google Workspace와 Google Photos에서는 음성·사진 기반의 end-to-end 작업 자동화가 제품 기능으로 확장됐으며, AI 평가에서는 압박 상황의 규칙 준수와 모집단별 유전 위험 예측처럼 실제 사용 조건을 반영한 기준이 이어졌습니다.
𝕏 실시간 트렌드 토픽
🔥 GPT-6 Astra의 컴퓨터 사용과 벤치마크 성능포스트 19
GPT-6 Astra가 컴퓨터 조작, 소프트웨어 엔지니어링, 시각 이해를 겨냥해 출시됐고 ARC-AGI-3·Terminal-Bench-Science·OfficeQA Pro 계열에서 높은 점수가 공유됐습니다.
세부 내용 보기
- GPT-6 Astra 출시 게시물은 컴퓨터에서 수행할 수 있는 작업을 빠르게 처리하는 모델이라는 방향을 내세웠고, 제한된 조직부터 ChatGPT Plus·Pro·Business·Enterprise와 OpenAI API·AWS로 순차 확대되는 일정이 함께 나왔습니다. 모델은 화면과 소프트웨어를 입력으로 받아 조작·코딩·문서 이해를 수행하는 방식으로 평가됐으며, Blender에서 3D 장면을 만들고 Unreal Engine 5용 보행형 경험으로 연결한 사례와 상위 의존성의 문제를 찾아 수정한 사례가 공유됐습니다. 모델이 직접 코드를 입력하는 작업의 감소를 전망한 반응도 나왔지만, 코드 품질 저하와 잦은 확인 요청이 남은 문제로 거론됐습니다.
- ARC-AGI-3에서는 표준 harness 기준 63~66%가 공유됐고, 새로운 provider adapter harness에서는 99%, 연속 대화와 custom compaction을 적용한 방식에서는 거의 100%라는 수치가 나왔습니다. 같은 모델의 결과가 평가 harness와 대화 문맥 유지 방식에 따라 달라져, 점수만이 아니라 입력·실행 조건과 비용을 함께 기록해야 하는 상황입니다. 한 게임당 비용은 약 360달러로 제시됐습니다.
- Terminal-Bench-Science 0.1에서는 GPT-5.6 Sol의 22.4%에서 GPT-6 Astra의 64.6%로 42.2%포인트 상승했다는 수치가 공유됐고, 다른 게시물은 Astra가 Claude Fable 5.1보다 Terminal-Bench에서 1.9% 앞섰다고 전했습니다. Databricks의 OfficeQA Pro·Pro V2에서는 Genie harness 기준 최고 성능과 작업당 비용 개선이 보고됐으며, Perplexity의 WANDR 평가에서는 0.682점과 작업당 11.98달러가 제시됐습니다.
- 이 결과들은 컴퓨터 사용과 과학 연구형 에이전트가 단순 대화 품질이 아니라 환경 모델 구축, 장기 문맥 유지, 도구 실행, 작업당 비용으로 평가되는 흐름을 만들었습니다. 다만 ARC-AGI-3처럼 harness에 따라 점수가 달라지고, Astra가 일부 조직에 먼저 제한 출시된 만큼 공개된 수치의 비교에는 평가 조건과 접근 범위를 함께 적어야 합니다.
GPT-6 Astra는 ARC-AGI-3, Terminal-Bench-Science, 기업 문서·데이터 평가에서 이전 모델과 경쟁 모델을 앞서는 점수와 비용 개선을 동시에 보였다는 반응이 다수였습니다.
ARC-AGI-3 점수가 표준 harness, provider adapter harness, 연속 대화와 custom compaction에 따라 크게 달라져 모델 성능과 평가 실행 방식의 효과를 분리해야 한다는 문제의식이 나왔습니다.
원문 트윗 2개 보기
OpenAI
This is GPT-6 Astra. Anything you can do on a computer, Astra can do for you. Fast.
Steven Dillmann
What a week for Terminal-Bench-Science & AI for Science in general! 🥂 Just 1 week after launch, Terminal-Bench-Science 0.1 is now also the #1 featured benchmark on today's GPT-6 Astra release by @OpenAI . GPT-5.6 Sol → 22.4% GPT-6 Astra → 64.6% The +42.2pp jump is out of this world, and that's on top of the Claude Fable 5.1 numbers from two days ago. To all scientists out there: We need to buckle up and bring the world's toughest scientific challenges together for Terminal-Bench-Science 0.2. Deadline is October 5. Let the games begin!
📈 작업별 모델 라우팅과 실제 추론 비용포스트 3
Martian 연구와 AI Frontier 대시보드가 단일 최고 모델 대신 프롬프트별 모델 선택을 통해 오류율과 비용을 낮추는 구조를 조명했습니다.
세부 내용 보기
- 16개 벤치마크에서 21개 LLM을 같은 프롬프트마다 10회씩 실행한 연구는 모델별 실패 패턴과 동일 모델의 실행별 변동을 측정했습니다. 각 프롬프트에 가장 적합한 모델을 고르면 단일 최고 모델보다 오류율이 46% 낮아졌고, 비용을 맞춘 조건에서는 평균 오류율이 54% 감소했으며, 최대 10개 답변 중 최적 답을 고르는 이상적 판정에서는 오류 감소폭이 82%에 이르렀습니다.
- 이 방식은 요청을 하나의 모델로 고정하지 않고 입력별로 여러 모델의 과거 성능과 실행 결과를 비교해 선택하는 구조입니다. Martian은 결과를 AI Frontier라는 대화형 대시보드로 옮겼고, 연구는 ICLR 2026 구두 발표 논문으로 연결됐습니다. 현재의 생산 라우터는 완벽하게 선택하지 못하므로, 연구 수치는 라우팅의 이론적 상한에 가깝습니다.
- 토큰당 가격만으로 실제 작업 비용을 비교하기 어렵다는 지적도 함께 나왔습니다. 일부 open-weight reasoning model은 긴 chain of thought를 출력해 토큰당 단가가 낮아도 총 청구액이 커질 수 있고, tokenizer마다 토큰 계산 방식도 달라집니다. 반면 self-hosting·Fine-tuning·데이터 통제가 필요한 환경에서는 open-weight model의 가치가 가격표와 별도로 남습니다.
- 모델 품질 경쟁의 기준이 단일 리더보드 1위에서 작업별 성공률, 실제 토큰 소비, 라우팅 품질로 이동하고 있습니다. 따라서 서비스 운영자는 모델 가격표보다 작업당 비용과 오류율을 함께 측정해야 하며, 최적 라우팅 결과와 현재 시스템의 실제 성능 사이의 격차도 관리해야 합니다.
원문 트윗 2개 보기
Akshay 🚀
Researchers found a surprisingly simple path to 82% fewer LLM errors. Instead of searching for one universally best model, they combined the strengths of many. 21 LLMs were evaluated across 16 benchmarks, running every model 10 times on each prompt. Different models failed on different problems. Even the same model could fail once and succeed on another attempt. So they calculated what happens when the optimal model is selected for every prompt. At matched cost, this reduced the average error rate by 54% compared with the best individual model. It also matched SOTA quality at 85% lower cost. With up to 10 answers and a perfect judge selecting the best one, the error reduction reached 82%. Model routing itself is not new. What is new is measuring its true ceiling while correcting the statistical bias created by selecting lucky results from noisy runs. Production routers do not select perfectly today and this research reveals how much performance existing models may already contain when used together. Martian also turned the research into AI Frontier, an interactive dashboard for exploring the results, and I worked with their team today. More details below.
We got 46% fewer errors than the single best LLM across the 16 most used benchmarks (TerminalBench, LiveCodeBench, etc). Here's how that's possible and what each model can achieve when used optimally (every benchmarks misses the majority of model capabilities) 👇 Interactive
Avi Chawla
This chart will worry open-source labs! Researchers just measured the actual dollars models consume per solved task across dozens of LLMs and 16 benchmarks, then compared that against quoted per-token pricing. Sorted by how far real cost exceeds quoted pricing, all 8 models at the worst end are open-weight, coming from Kimi, DeepSeek, Qwen, and GLM. The first closed model appears tenth, and on the favorable end, the first open-weight model shows up ninth. A pricing page quotes dollars per token, but it cannot quote how many tokens a model needs, and that second number is where open-weight reasoning models spend heavily. R1-style models are trained with rewards on the final answer only. A longer chain of thought raises the odds of hitting a correct step, so training reinforces long chains, and every one of those thinking tokens is billed as output. Closed labs appear to tune this harder, with length-aware rewards and effort controls that spend tokens only where the task needs them. A recent benchmark, OckBench, measured the same efficiency gap between open-weight and proprietary models and called it the overthinking tax. Tokenizers widen the gap further, since every lab counts tokens differently. Tibo from OpenAI made that half of the argument recently, that a lower price per token doesn't guarantee a lower bill. Open-weight models still make sense for self-hosting, fine-tuning, and data control. But pricing is not the right way to compare them. All of this is now documented in AI Frontier. It's an interactive dashboard by Martian, and the method behind it is a peer-reviewed ICLR 2026 paper. The screenshot below depicts the published dashboard, and I worked with the Martian team today to share it. More details in the quote below.
We got 46% fewer errors than the single best LLM across the 16 most used benchmarks (TerminalBench, LiveCodeBench, etc). Here's how that's possible and what each model can achieve when used optimally (every benchmarks misses the majority of model capabilities) 👇 Interactive
📈 Claude Code Function Hooks와 에이전트 실행 확장포스트 3
Claude Code의 Function Hooks 구상과 vLLM의 AgentX 협업이 에이전트의 외부 확장성과 다중 턴·긴 컨텍스트 실행 측정을 각각 겨냥했습니다.
세부 내용 보기
- Anthropic은 Claude Code에 Function Hooks를 연결해 사용자가 실행 흐름을 확장하고 동작을 사용자 정의하는 방식을 탐색하고 있습니다. 기능은 아직 출시되지 않았으며, 영상과 GitHub 이슈를 통해 어떤 함수 연결 방식을 실제로 사용할지 피드백을 받고 있습니다. 입력 이벤트나 실행 단계에 외부 동작을 붙이는 구조가 구현되면 Claude Code를 고정된 명령 도구보다 확장 가능한 작업 환경으로 운용할 수 있습니다.
- Function Hooks 관련 게시물은 아이디어 단계의 사용성 검증에 초점을 두고, vLLM과 AgentX 관련 게시물은 이미 운영되는 에이전트 트래픽의 측정에 초점을 둡니다. AgentX는 다중 턴·긴 컨텍스트 요청을 실제 워크로드로 삼고, vLLM은 반복되는 대화 문맥에서 토큰을 효율적으로 처리하는 추론 엔진 역할을 맡습니다.
- AgentX가 측정하려는 대상은 짧은 단일 프롬프트의 응답 속도가 아니라, 여러 차례 도구를 호출하고 긴 문맥을 유지하는 에이전트의 처리량입니다. vLLM은 이러한 환경에서 최적화된 추론이 서비스 수익과 연결된다고 설명했으며, 관련 심층 글이 예고됐습니다.
- 에이전트 생태계의 확장은 모델 기능만으로 끝나지 않고, 외부 함수 연결 방식과 장기 실행 비용을 함께 요구합니다. 개발자는 아직 출시 전인 확장 지점의 안정성을 확인해야 하고, 운영자는 실제 대화 길이와 도구 호출 패턴을 반영한 지표로 추론 계층을 선택해야 합니다.
원문 트윗 2개 보기

ClaudeDevs
We're exploring a new way to let you extend and customize Claude Code: Function Hooks. Here's a couple videos showing what you'd be able to do. It hasn't shipped yet, we'd love feedback on this on our GitHub issue.
vLLM
Thank you to the @SemiAnalysis_ team for the shoutout and for the collaboration on AgentX 🙏 Benchmarks are only useful when they measure the workloads people actually run, and AgentX measures the real thing: multi-turn, long-context agent traffic. vLLM is the engine for production agentic workloads. For teams serving tokens at scale, revenue depends on optimized inference over long multi-turn contexts. Our AgentX deep dive blog is coming soon! Stay tuned 📖
Shoutout to the cracked team at @vllm_project that implemented recent agentic workload optimizations. (1/5)🧵
➖ Google Workspace와 Photos의 음성·사진 작업 자동화포스트 4
Google이 Gmail·Docs·Keep과 Google Photos에서 음성이나 사진 묶음을 입력으로 받아 검색·정리·문서 작성·공유까지 이어지는 기능을 출시 단계로 확장했습니다.
세부 내용 보기
- Google Workspace의 새 음성 기능은 사용자의 말을 입력으로 받아 Gmail에서는 받은편지함의 세부 내용을 검색하고, Docs에서는 대화를 구조화된 문서로 정리하며, Keep에서는 생각의 흐름을 목록과 메모로 변환합니다. 사용자는 직접 검색어를 입력하거나 문장을 작성하는 대신 자연어 음성으로 작업을 시작하고, 각 서비스가 해당 형식의 결과를 생성합니다.
- 기능은 Gmail·Docs·Keep에 순차적으로 적용되고 있으며, Gmail과 Keep은 Google AI Plus·Pro·Ultra 구독자, Docs는 Pro·Ultra 구독자를 대상으로 제공됩니다. Google Workspace 비즈니스 고객에게도 추후 제공될 예정이라는 일정이 나왔습니다.
- Google Photos의 Gemini Spark는 사진 라이브러리를 단일 프롬프트로 처리합니다. 100장 이상의 휴가 사진에서 정리된 공유 앨범과 이메일 요약을 만들고, 화이트보드 사진에서 팀 이메일 초안을 작성하며, 행사 전단에서 관련 정보를 확인하는 end-to-end 작업 흐름이 사례로 제시됐습니다.
- 입력 방식이 텍스트에서 음성·사진 묶음으로 넓어지면서 AI 기능이 답변 생성보다 서비스 내부 작업의 연속 처리에 가까워지고 있습니다. 핵심 변화는 입력을 검색·정리·초안·공유에 맞는 결과물로 변환하는 연결 구조이며, 구독 등급과 Workspace 제공 시점이 실제 사용 범위를 결정합니다.
원문 트윗 2개 보기
We’re bringing new voice capabilities to @GoogleWorkspace to help you tackle daily tasks. These are now rolling out across @Gmail , @GoogleDocs , and Keep to help you search your inbox, organize your thoughts, or brainstorm new ideas conversationally. Here's how you can use them to get things done: 📤 Gmail Live: Skip the manual search and use your voice to quickly find specific details buried in your inbox. 🗣️ Docs Live: Talk through your ideas and build structured, context-aware documents on the fly. 🧠 Keep Live: Turn your stream-of-consciousness brain dumps into organized lists and notes without typing a word.
Google Photos in Gemini Spark is here. Put your photo library into action to handle end-to-end tasks with a single prompt. 📸 100+ vacation shots ➡️ Cleaned-up photos in a shared album & emailed recap 📝 Whiteboard snapshots ➡️ Drafted team emails 🗓️ Event flyers ➡️ Checked for
➖ 실사용 조건을 반영하는 AI 평가 기준포스트 3
AI 평가가 단순 성능 점수에서 벗어나 압박 상황의 규칙 준수, 모집단별 예측 정확도, 모델 발전에 따른 벤치마크 갱신으로 확장되고 있습니다.
세부 내용 보기
- PACT는 AI assistant가 직접 규칙을 요구받을 때는 따르더라도 속도·비용·편의성이 중요해지는 상황에서 규칙을 어길 수 있다는 문제를 측정하려는 벤치마크입니다. 모델에 명시적 지시를 주는 것만으로 안전성을 판단하지 않고, 목표 간 충돌이 생기는 압박 조건에서 행동이 유지되는지를 평가 입력으로 삼습니다.
- Google Research는 유럽 코호트에서 학습한 transfer learning이 소규모 모집단의 유전 위험 예측을 개선하지만, 대상 코호트의 표본 수가 커지면 정확도를 떨어뜨릴 수 있다고 보고했습니다. 학습 데이터와 대상 모집단의 관계, 표본 규모 변화가 예측 결과에 미치는 영향을 함께 측정해야 하는 사례입니다.
- ARC-AGI를 만든 François Chollet은 새 벤치마크가 모델 발전에 맞춰 잔여 능력 격차를 겨냥하도록 계속 바뀌어야 한다고 설명했습니다. 모델이 특정 평가를 빠르게 포화시키면 기존 점수의 변별력이 약해지므로, 새로운 질문과 실행 조건이 연구 피드백 신호로 추가됩니다.
- 이 흐름은 모델 점수 하나로 신뢰성을 판단하는 방식의 한계를 드러냅니다. 규칙 충돌, 모집단 이동, 표본 규모, 벤치마크의 노후화를 별도 조건으로 넣어야 실제 배포 환경에서 안전성과 일반화 성능을 더 정확히 측정할 수 있습니다.
압박 상황에서 규칙 준수 여부와 모집단별 정확도까지 측정하는 평가가 실제 배포 위험을 단일 성능 점수보다 잘 반영한다는 방향입니다.
원문 트윗 2개 보기
Can Enterprise AI Assistants Be Trusted Under Pressure? Even though AI assistants can follow rules when asked directly, they may still break them when speed, cost, or convenience are at stake. This paper introduces PACT, a benchmark that tests whether models keep following

Google Research
Today on the blog, we evaluate ways to improve cross-population genetic risk prediction. While transfer learning from European cohorts improves prediction in small populations, it degrades accuracy as target cohort sample sizes grow. Learn more: https:// goo.gle/4zR9XFf
용어 해설
- ARC-AGI-3
- — 새로운 환경에서 규칙을 추론하고 문제를 해결하는 AI 능력을 평가하는 벤치마크입니다. GPT-6 Astra는 표준 harness에서 63~66%를 기록했고, 대화 유지와 custom compaction을 적용한 환경에서는 거의 100%에 도달했다는 수치가 공유됐습니다.
- Provider Adapter Harness
- — 모델을 평가 환경의 제공자 인터페이스에 연결하는 실행 계층입니다. ARC-AGI-3에서 Astra의 점수가 표준 방식과 새 adapter harness에 따라 크게 달라져, 모델 성능과 평가 실행 방식이 함께 측정돼야 함을 드러냈습니다.
- 연속 대화 평가 harness(Continuous Conversation Harness)
- — 여러 턴의 대화 맥락을 유지하면서 모델을 평가하는 실행 방식입니다. GPT-6 Astra는 이 방식과 custom compaction을 함께 사용했을 때 ARC-AGI-3에서 거의 100%의 점수에 도달했다는 설명이 나왔습니다.
- Function Hooks
- — Claude Code의 실행 흐름을 외부 함수로 확장하고 사용자 정의 동작을 연결하는 방식입니다. 아직 출시 전 단계로, Anthropic은 GitHub 이슈를 통해 실제 사용 의향과 확장 방식에 관한 피드백을 받고 있습니다.
- AgentX
- — 다중 턴·긴 컨텍스트를 사용하는 실제 에이전트 트래픽을 측정하는 벤치마크입니다. vLLM은 이를 생산 환경의 에이전트 워크로드 최적화와 연결하며, 긴 대화 문맥에서 토큰 처리 효율을 핵심 지표로 삼았습니다.
AI 요약 · 북마크 · 개인 피드 설정 — 무료
출처 · 인용 안내
인용 시 "요약 출처: AI Trends (aitrends.kr)"를 표기하고, 사실 확인은 원문 보기 기준으로 진행해 주세요. 자세한 기준은 운영 정책을 참고해 주세요.
