TL;DR
이번 기간에는 대화형 AI가 답변을 만드는 단계를 넘어 도구를 호출하고, 브라우저와 파일을 조작하며, 반복 작업을 스킬로 저장하는 흐름이 두드러졌습니다. Grok @Bot은 Google Workspace 연동, 영상 생성·편집, 업무 자동화 사례로 확장됐고 Grok 4.6은 Agentic Index에서 Claude Opus 5와 공동 1위를 기록했습니다. ChatGPT와 Codex에는 Computer History, Exa 연결, Record & Replay가 추가돼 작업 맥락과 외부 자료를 재사용하는 경로가 마련됐습니다. 동시에 Hawkeye의 하드웨어 인식 커널 최적화, Chroma Foundation의 자기 개선형 메모리, Muse Spark 1.2의 로봇 제어가 에이전트 실행을 뒷받침하는 기술 축으로 나타났습니다.
𝕏 실시간 트렌드 토픽
🔥 Grok @Bot의 업무 자동화와 멀티모달 실행포스트 6
Grok @Bot이 Google Workspace 계정과 외부 데이터베이스를 연결하는 업무 에이전트로 쓰이고, 개념 설명 영상의 생성·편집까지 한 대화 흐름에서 처리하는 사례가 이어졌습니다.
- Grok @Bot 활용 사례가 단순 질의응답에서 개인 업무와 영상 제작으로 넓어지면서, 사용자는 채팅창을 클라이언트로 두고 @bot을 서버처럼 다루는 흐름을 제시했습니다. Google Workspace 계정을 인증해 Mac과 iPhone에서 활용하거나, gh·sheets·SaaS를 데이터베이스로 연결하는 구조가 언급됐고, 복잡한 개념 설명 영상에서는 rank·spectral power·optimization을 기하학적으로 풀어낸 뒤 영상 자산 생성과 편집까지 이어졌습니다. 같은 대화 인터페이스에서 입력 해석, 외부 서비스 호출, 콘텐츠 생성, 결과 편집을 연결한다는 점이 핵심입니다.
원문 트윗 2개 보기
matt palmer
@mattyp
My workflows changed when I started treating each Grok Bot like an application Client: chat Server: @bot DB: anything (gh, sheets, SaaS) I've been white-pilled on "agents for everything" - thin client, thick server
Yun-Ta Tsai
@yunta_tsai
Grok @bot visualized human migration using @imagine to generate assets, then wrote the visualization of human migration. It also cut and edited the videos, making sure the message is coherent. Pretty good video editing tool just using chat.
If human history were a Git repository, influence would be the cherry-pick, and war the brutal way of settling a merge conflict.
🔥 Grok 4.6의 Agentic Index 공동 1위포스트 2
Grok 4.6이 Artificial Analysis Agentic Index에서 Claude Opus 5와 나란히 59점을 기록하며 공동 1위로 언급됐습니다.
- Agentic Index는 일반적인 답변 품질보다 도구 사용, 계획, 자율성, 복합 문제 해결처럼 실제 에이전트 업무에 필요한 능력을 측정하는 지표로 제시됐습니다. Grok 4.6은 high 설정에서 59점, Claude Opus 5는 max 설정에서 59점을 기록했고, 게시물은 이 결과를 Grok Build와 Grok Bot이 질문 응답을 넘어 도구를 사용하고 작업을 완료하는 흐름과 연결했습니다. 다만 제공된 포스트에서는 세부 평가 항목별 점수나 재현 조건이 추가로 제시되지 않았습니다.
포함된 두 포스트 모두 Grok 4.6의 59점 공동 1위를 실제 에이전트 업무 능력을 평가하는 결과로 긍정적으로 받아들였습니다.
원문 트윗 2개 보기

Elon Musk
@elonmusk
Grok 4.6
Grok 4.6 just tied for #1 on the Artificial Analysis Agentic Index • Grok 4.6 (high) — 59 • Claude Opus 5 (max) — 59 Outperforming Claude Fable 5 and GPT-5.6 Sol We’re entering the agentic era, and this is exactly the kind of benchmark that matters so much: tool use,
Tesla Owners Silicon Valley
@teslaownersSV
BREAKING: Grok 4.6 just tied for #1 on the Artificial Analysis Agentic Index. • Grok 4.6 (High) — 59 • Claude Opus 5 (Max) — 59 Grok 4.6 is now at the top of the benchmark, ahead of other leading models. The Agentic Index focuses on capabilities that matter for real-world AI agents — tool use, planning, autonomy, and complex problem-solving. This is especially significant as Grok expands into Grok Build and Grok Bot, where AI needs to do more than answer questions — it needs to use tools and complete tasks. Source: Artificial Analysis
📈 ChatGPT·Codex의 컴퓨터 기록과 재사용 스킬포스트 5
ChatGPT와 Codex가 최근 컴퓨터 작업 맥락을 기억하고, 사용자의 업무 시연을 편집 가능한 스킬로 바꾸며, Exa를 통해 100B개 이상의 웹 자료와 문서에 접근하는 기능이 추가됐습니다.
- 반복 업무를 매번 처음부터 설명해야 하는 문제가 작업 자동화의 마찰로 남아 있었고, 새 기능들은 기록된 작업 맥락과 외부 검색을 입력으로 받아 에이전트 실행에 재사용하는 경로를 만들었습니다. Computer History는 Mac 앱에서 최근 앱과 웹사이트 활동을 대화 맥락으로 제공하며 EEA·영국·스위스의 Pro·Business·Enterprise 사용자에게 제공됩니다. Record & Replay는 지출 보고나 휴가 신청 같은 시연을 inspectable·editable skill로 변환하고, Exa 플러그인은 Codex와 ChatGPT에 100B개 이상의 웹사이트·논문·문서·인물·기업 자료를 연결합니다. 브라우저 기반 회계 처리, Gmail 정리, Stripe Radar 규칙 설정, PR 품질 확인 같은 작업 사례가 함께 제시됐습니다.
원문 트윗 2개 보기

OpenAI Developers
@OpenAIDevs
Europe, two updates for you. Computer History is now available in the EEA, UK, and Switzerland for ChatGPT Pro, Business, and Enterprise users on Mac. ChatGPT can remember activity across your apps and websites, so future conversations feel more personalized and require less explanation.
Codex and ChatGPT can now understand the context of your recent work. Opt into Computer History to give ChatGPT richer context, so it can pick up where you left off, understand patterns in your work, and suggest skills or scheduled tasks for work you repeat.

OpenAI Developers
@OpenAIDevs
Record & Replay is also available in the ChatGPT app for macOS across the EEA, UK, and Switzerland. Turn your go-to workflows into reusable skills by showing, not just telling, ChatGPT Work and Codex what to do.
Show Codex a workflow once. Reuse it as a skill. Record & Replay lets you show Codex a recurring task, like filing an expense report or submitting a time-off request. Codex turns that demo into an inspectable, editable skill. You control when recording starts and stops.
📈 Hawkeye와 전용 칩으로 좁혀지는 AI 컴퓨팅 최적화포스트 4
Hawkeye는 코딩 에이전트가 GPU별 최적화 커널을 작성하도록 만들고, Waymo는 자율주행용 자체 칩으로 외부 가속기 의존도를 낮추는 방향을 제시했습니다.
- 새로운 ML 가속기가 늘어날수록 하드웨어별 커널 최적화와 전용 연산 장치 지원이 소프트웨어 병목이 됩니다. Hawkeye는 최적화 전략마다 10개의 단위 테스트와 solution kernel을 최소 감독 신호로 제공해 코딩 에이전트가 테스트 시점 계산을 확장하도록 하며, Ampere·Hopper·Blackwell과 NVIDIA·AMD, FP8·NVFP4·MXFP4 사이의 커널 이식을 수행합니다. 별도 게시물에서는 Waymo가 TSMC 5nm 공정의 1,000 TOPS 초과 자체 ASIC을 최신 robotaxi에 양산 적용했다고 전했습니다. 두 사례 모두 모델 규모만 키우는 대신 실행 하드웨어에 맞춘 소프트웨어와 칩 설계가 처리 성능의 조건이 된다는 점을 드러냅니다.
원문 트윗 2개 보기
Arya Tschand
@AryaTschand
We’ve seen an explosion of new ML chips with unique architectural features, but software support remains the critical bottleneck Achieving peak performance increasingly relies on hardware-specific optimizations in the kernels, but we observe that coding agents are particularly weak at this Introducing Hawkeye, a framework that brings hardware-awareness to coding agents by grounding them in a minimal and comprehensive taxonomy of optimization strategies For new GPU or ML accelerator architectures, you only need to write 10 unit tests and solution kernels (one per optimization strategy), and we show that coding agents can effectively scale test-time compute with this minimal supervision to write hardware-aware kernels Hawkeye can port kernels across architectures (Ampere, Hopper, Blackwell), vendors (NVIDIA, AMD), and precisions (FP8, NVFP4, MXFP4) while consistently leveraging hardware features and approaching expert kernel performance Work co-led with @keramakr and done in collaboration with Alexander Ingare @simonguozirui @18jeffreyma @ZishenW @simran_s_arora @Azaliamirh @profvjreddi
TLDR Newsletter
@tldrnewsletter
Waymo has built its own custom chip for its robotaxis, an application-specific integrated circuit made by TSMC on a 5-nanometer process that is capable of more than 1,000 TOPS. The chip is in production in Waymo's latest generation of vehicles and reduces its reliance on outside suppliers including NVIDIA.
📈 Chroma Foundation의 에이전트 세션 메모리포스트 2
Chroma가 에이전트 세션에서 스스로 개선되는 메모리를 구축하는 연구 프리뷰 Foundation을 공개했습니다.
- 장시간 에이전트 작업에서는 이전 세션의 실행 기록을 다음 작업에 활용하는 메모리 구조가 필요합니다. Foundation은 에이전트 세션을 입력으로 받아 경험을 축적하고 자기 개선형 메모리로 구성하는 Chroma의 연구 프리뷰이며, 게시물은 이를 3년간 준비한 Chroma의 메모리 솔루션으로 설명했습니다. 세부 저장 구조나 정량 평가 수치는 제공된 포스트에 없지만, 에이전트가 매번 초기 상태에서 출발하지 않도록 세션 단위 경험을 재사용하려는 방향은 분명하게 제시됐습니다.
원문 트윗 2개 보기
Jeff Huber
@jeffreyhuber
I’ve been looking forward to today for 3 years Today we’re announcing Foundation - Chroma’s solution to memory Our research preview of this technology builds self-improving memory from your agent sessions. Try it out at https:// trychroma.com/foundation

Harrison Chase
@hwchase17
very cool launch we did a webinar with jeff on "wiki" style memory and it's clear he'd thought about this problem a lot (webinar here: https:// youtube.com/watch?v=Lsut4T Cfygw …)
I’ve been looking forward to today for 3 years Today we’re announcing Foundation - Chroma’s solution to memory Our research preview of this technology builds self-improving memory from your agent sessions. Try it out at https:// trychroma.com/foundation

📈 Muse Spark 1.2의 멀티모달 로봇 오케스트레이션포스트 3
Muse Spark 1.2가 시각 정보와 도구 호출을 결합해 로봇의 하위 작업을 계획하고, Muse Image·Muse Video와 연결되는 멀티모달 처리 흐름을 확장했습니다.
- 실제 환경의 로봇과 영상 업무는 텍스트만 처리해서는 해결되지 않으므로 시각 관찰, 도구 호출, 행동 결과 확인을 한 루프로 연결해야 합니다. Muse Spark 1.2는 사용자 지시를 받아 도구 호출을 해석하고 결과를 관찰하면서 작업을 반복하며, 양손 로봇이 책상 위 물건을 정리할 때 hair brush와 makeup brush를 구별하고 lipstick을 서랍에 넣도록 하위 작업을 계획합니다. 또 Muse Image와 함께 agentic media generation을 수행하고 학습 데이터용 상세 캡션을 생성하며, 텍스트·이미지·영상에서 후속 애플리케이션용 신호와 인사이트를 추출합니다.
원문 트윗 2개 보기

AI at Meta
@AIatMeta
Muse Spark 1.2 supports a broad range of multimodal tasks, from turning visuals into working code to translating perception into physical action. It also brings robust audio-visual understanding to enable video-heavy workflows common in real-world enterprise use. Today, we’re sharing new evals and demos that illustrate the breadth of the model’s visual understanding and reasoning capabilities. Let’s start with a demo that shows how Muse Spark parses multimodal observations and calls tools to guide a robot to navigate in an unstructured environment to find a rubber duck.

AI at Meta
@AIatMeta
Muse Spark brings spatial intelligence to robotics. A specialized variant of Muse Spark acts as the robot brain and orchestrator: it takes a user instruction, decodes tool calls, observes the results, and loops until the task is complete. As shown in this demo, Muse Spark 1.2 can plan sub-tasks for a bimanual robot tidying a desk, distinguishing a hair brush from a makeup brush and placing the lipstick in a drawer.
📈 장시간 에이전트 작업의 자기 평가와 전략 고정포스트 2
에이전트가 긴 작업 중 진행 상황을 평가하거나 다른 에이전트를 학습시키는 과정에서 초기 전략을 고정하는 문제가 새 연구와 dots3-note Preview를 통해 부각됐습니다.
- 긴 작업에서는 실행 도중 전략을 다시 선택하고 현재 결과를 점검하는 능력이 필요하지만, 한 연구는 에이전트가 첫 단계에서 학습 전략을 고정한 뒤 남은 예산을 같은 전략 안의 국소 조정에 사용한다고 관찰했습니다. experience-driven scaffold는 GSM8K에서 12.6점, HumanEval에서 40.8점을 높였지만 전략 자체는 고정됐고, 추가 추론 계산은 쉬운 작업에는 효과가 있었으나 가장 어려운 작업에는 거의 영향을 주지 않았습니다. 별도로 dots3-note Preview는 280B total·16B active parameters, 512K context, text·vision·speech 입력을 바탕으로 장시간 작업이 끝나기 전에 자신의 진행 상황을 평가하도록 학습한 open-weight model로 소개됐습니다.
원문 트윗 2개 보기
elvis
@omarsar0
Finally a good paper testing whether agents can really post-train other agents. (bookmark it) They analyzed a large corpus of publicly released post-training trajectories. Across tasks, the agent locks in its training strategy at the very first step and spends the entire remaining budget on local adjustments inside it. They then tried three escalating fixes. An experience-driven scaffold lifted execution broadly, worth 12.6 points on GSM8K and 40.8 on HumanEval, and the strategy stayed frozen. Human guidance redirected the opening choice, and the agent slid back into local loops once training began. Extra inference compute paid off on easy tasks and did almost nothing on the hardest one. What agents lack here is a way to reconsider strategy while execution is still running. Paper: https:// arxiv.org/abs/2608.19072 Track more trending AI papers in our academy: https:// academy.dair.ai
elvis
@omarsar0
dots3-note Preview is an open-weight model built for tasks that run for hours—sometimes days. → 280B total / 16B active parameters → 512K context → Text, vision, and speech → Reasoning, coding, and tool use But the most interesting part isn’t the model size. It’s how the model learns to evaluate its own progress before a long task is finished.
➖ Embedding 모델과 LLM의 비용·성능 선택포스트 1
새 연구는 LLM이 embedding model보다 높은 성능을 낼 수 있지만 비용이 더 크다는 선택 기준을 제기했습니다.
- 검색이나 유사도 계산에서 embedding model과 LLM 중 무엇을 쓸지는 품질과 비용의 균형에 달려 있습니다. 공유된 연구는 LLM이 embedding model을 능가하는 경우를 확인했지만 훨씬 높은 비용이 따른다고 전했으며, Stanford AI Lab 게시물은 이를 두 접근법의 선택 문제로 압축했습니다. 제공된 포스트에는 모델별 점수, 비용 배수, 데이터셋 조건이 없으므로 특정 사용 사례에서 어느 쪽이 우월한지는 확인되지 않았습니다.
용어 해설
- 에이전트 성능 지수(Agentic Index)
- — AI 에이전트가 단순히 답변하는 수준을 넘어 도구 사용, 계획 수립, 자율 실행, 복합 문제 해결을 얼마나 수행하는지 평가하는 지표입니다. Grok 4.6과 Claude Opus 5의 비교에 쓰였습니다.
- 기록 및 재생(Record & Replay)
- — 사용자가 반복 업무를 한 번 수행하는 과정을 기록하면 ChatGPT Work와 Codex가 이를 편집 가능한 스킬로 바꿔 이후 같은 작업에 재사용하는 기능입니다.
- 하드웨어 인식 커널(Hardware-aware Kernel)
- — GPU나 ML 가속기의 구조와 정밀도에 맞춰 연산 커널을 최적화하는 방식입니다. Hawkeye는 최적화 전략별 테스트와 커널을 바탕으로 여러 아키텍처에 코드를 이식합니다.
- 자기 개선형 메모리(Self-improving Memory)
- — 에이전트 세션에서 축적된 실행 기록을 활용해 기억 구조를 스스로 개선하는 방식입니다. Chroma의 Foundation은 에이전트 작업 흐름에서 이 메모리를 구축하는 연구 프리뷰입니다.
- 멀티모달 추론(Multimodal Reasoning)
- — 텍스트·이미지·영상·음성 같은 여러 입력을 함께 처리해 상황을 해석하고 다음 행동을 결정하는 능력입니다. Muse Spark 1.2는 도구 호출과 로봇 제어까지 연결합니다.
AI 요약 · 북마크 · 개인 피드 설정 — 무료
출처 · 인용 안내
인용 시 "요약 출처: AI Trends (aitrends.kr)"를 표기하고, 사실 확인은 원문 보기 기준으로 진행해 주세요. 자세한 기준은 운영 정책을 참고해 주세요.