TL;DR
이번 포스트 묶음에서는 Grok Bot의 데스크톱 작업 자동화와 물리적 생산 실행 가능성, AI agent의 컨텍스트 확보·도구 호출 신뢰성, Agent Harness의 운영 계층화가 이어졌다. 두 연구는 각각 25~300개 후보를 대상으로 한 Optimal Question Asking과 507개 업무 흐름을 평가한 Thinkingbox로 agent의 추론 선택과 실제 상태 변경을 측정했다. Transformer의 Attention과 FFN을 손계산으로 따라가는 학습 자료는 위치 간 혼합과 feature 차원 변환을 분리해 설명하며, Honor의 휴머노이드 로봇은 100m를 9.32초에 주파했다. Claude Code CLI 수정, OpenCode inference service 안정화, Runway의 LATAM 진출과 Pixel 11 Pro XL의 AI 기능도 제품·운영 변화로 함께 나타났다.
𝕏 실시간 트렌드 토픽
📈 Grok Bot의 데스크톱 자동화와 공장 실행 구상포스트 2
Grok Bot을 활용해 데스크톱 파일 정리 같은 반복 작업을 처리한 사례와 이미지에서 출발해 복잡한 공학 문제를 해결하려는 실험이 공유됐다.
세부 내용 보기
- Grok Bot 관련 포스트에서는 일주일 동안 쌓인 수백 개의 데스크톱 파일을 특정 기준에 맞게 정리하는 작업에 화면 녹화를 입력으로 활용했고, 사용자는 반복적인 파일 분류를 맡긴 뒤 결과를 확인했다. 다른 사례에서는 네 장의 이미지를 설계 단서로 삼아 여러 bot이 복잡한 공학 문제를 처음부터 끝까지 해결하는 실험이 언급되면서, 소프트웨어 작업 자동화에서 공장 단위 실행으로 범위가 넓어지는 경로가 나타났다.
- 첫 사례의 입력은 작업 화면과 음성 기록이고, Grok Bot은 파일을 분류하는 행동을 수행하는 구조다. 두 번째 사례는 이미지 단서에서 설계 정보를 추론하고 bot 팀이 문제 해결 단계를 이어가는 흐름으로 제시됐지만, 완료율이나 생산 현장 성능 수치는 출처에 없다.
- 이 흐름의 의미는 Grok Bot이 대화 응답을 넘어 컴퓨터 조작과 물리적 생산 과정까지 맡을 수 있는지에 관심이 이동했다는 점이다. 다만 포스트에 나온 내용은 개별 실험과 기대를 담고 있어 실제 공장 운영 성과로 확대할 근거는 없다.
원문 트윗 2개 보기

Elon Musk
Try Grok @Bot
Alright, it happened. After a slow start with Grok Bot I just had my mind blown. Twice I did a screen recording with audio of me cleaning up my desktop after a week that resulted in hundreds of random files that all needed particular sorting It's a monotonous task I've always
Yun-Ta Tsai
Maybe one day Grok @bot can run the factory end-to-end.
Grok @bot is incredible - and they can even manufacture real physical objects! Here is a little experiment I did last night: I created a team of bots and asked them to solve a complex engineering problem end to end - starting from four images as design cues, inferring
📈 에이전트의 질문 선택과 업무 상태 신뢰성포스트 2
AI agent가 부족한 컨텍스트를 보완하는 행동을 선택하는 문제와, 도구 호출 뒤 실제 업무 상태가 올바르게 바뀌었는지를 측정하는 벤치마크가 함께 다뤄졌다.
세부 내용 보기
- 첫 연구는 사용자가 제약 조건을 빠뜨릴 때 agent가 기본값을 추측하거나 질문·검색·도구 호출을 선택해야 한다는 문제에서 출발한다. Optimal Question Asking은 잠재 작업 상태에 대한 믿음을 내부 단계에서 갱신하고, 외부 단계에서 다음 컨텍스트 행동·작업 행동·종료 행동을 예상 free energy와 비용을 기준으로 고르는 목적 함수를 두며, 결정론적 환경에서는 정보 이득을 토큰 비용으로 정규화할 수 있다.
- 이 연구는 정확한 posterior와 dynamic programming oracle을 만들고 25~300개 후보를 가진 이진·다중 선택 과제에서 frontier model을 측정했다. 별도 Thinkingbox 연구는 retail, hospitality, auto insurance, neobank IT, consulting support의 507개 policy-conditioned workflow를 격리된 MCP 호환 도구 세션에서 실행하고, 응답 문장이나 도구 호출이 아니라 agent가 남긴 백엔드 상태를 executable check로 판정했다.
- Thinkingbox에서 가장 강한 모델의 pass@1은 65.36%, pass^20은 25.25%였고, 일부 실패는 유효한 상태 변경 도구 호출 뒤에도 발생했다. 따라서 작업이 끝난 듯한 출력만으로는 완료 여부를 판단하기 어렵고, 허용된 상태·누락된 효과·불필요한 효과를 직접 검사하는 평가가 필요하다.
원문 트윗 2개 보기
elvis
What a fascinating paper on AI agents. A lot of the issues we see with AI agents today revolve around wrong assumptions the LLMs make. This leads to problems like hallucination, cost inefficiencies, unreliable tool calls and much more. I think if we can solve this problem, even current LLMs would significantly improve in terms of performance and efficiency. The problem is that context acquisition is treated as afterthought, but it shouldn't be that way. Users tend to leave out constraints when prompting. So the agent agent needs to guess the default, or spend tokens on a clarifying question, a retrieval call, a tool call, or a prompt trial. This new work gives this problem an objective function. Context acquisition becomes active inference over a latent task state. An inner step updates beliefs, and an outer step picks the next context action, task action, or stop action to minimize expected free energy under cost. In deterministic settings the epistemic term reduces to expected information gain, optionally normalized by token cost. That is directly implementable today as a scoring rule. They coin it as Optimal Question Asking, with exact posteriors and a dynamic programming oracle, then benchmark frontier models on binary and multiway tasks from 25 to 300 candidates. So you can measure the gap between your agent and the true optimum. Paper: https:// arxiv.org/abs/2608.19202 Track more trending AI papers in our academy: https:// academy.dair.ai
DAIR.AI
Banger paper from Microsoft. It's on agent reliability in real business workflows. (bookmark it) Thinkingbox is a sandbox with isolated MCP-compatible tool sessions, plus a benchmark of 507 policy-conditioned workflows across retail, hospitality, auto insurance, neobank IT, and consulting support. Every attempt is graded on the backend state the agent leaves behind. Executable checks accept valid trajectories and reject wrong, missing, or extra effects, so collateral damage counts against you. The strongest model reaches 65.36% pass@1 and 25.25% pass^20. Many failed trials terminate cleanly with valid state-changing tool calls. Watching the response or the tool call tells you very little about whether the task actually completed. Paper: https:// arxiv.org/abs/2608.19741 Track more trending AI papers in our academy: https:// academy.dair.ai
➖ Agent Harness와 Claude Code CLI의 운영 계층포스트 5
여러 agent를 연결하고 컨텍스트·정책·비용을 관리하는 Agent Harness의 필요성과 Claude Code 2.1.240의 CLI 안정화 변경이 이어졌다.
세부 내용 보기
- Agent Harness 관련 포스트는 agent마다 실행 환경을 다시 만들지 않도록 공유 계층을 두는 방향을 짚었다. Databricks의 Omnigent는 orchestration, control, collaboration을 위한 open-source meta-harness로 소개됐고, agent를 바꿔도 컨텍스트를 유지하며 작업 라우팅, contextual policy, spend control, human approval을 적용하는 구조를 제공한다.
- Claude Code 쪽에서는 2.1.240에서 CLI crash와 오류 메시지를 줄이는 bug fix 및 reliability improvement가 반영됐고, 별도 changelog에는 Claude-Code-Plugin-Manager 추가와 두 모델 항목 제거가 기록됐다. 릴리스 간격은 1일 0시간 27분 9초였으며 bundle file size는 69.3 kB 감소했고 prompt files는 3개, prompt tokens는 11,921개 늘었다.
- 이 변화는 agent 성능을 모델 응답만으로 결정하기보다 실행 환경, 도구 연결, 정책, 승인 절차, CLI 안정성을 함께 관리해야 한다는 실무 흐름을 드러낸다. Claude Code harness의 open-source 여부를 바라는 의견과 Omnigent의 공유 계층은 이 운영 부분을 독립적인 개발 자산으로 보는 시각을 뒷받침한다.
원문 트윗 2개 보기
Databricks
Working with multiple AI agents shouldn’t mean rebuilding the same setup every time. Omnigent is Databricks’ open-source meta-harness that gives agents a shared layer for orchestration, control, and collaboration. Switch agents without losing context, route tasks intelligently, set contextual policies and spend controls, and add human approval where it matters. @YoussefMrini and Quentin Ambard sit down with Databricks co-founder and CTO @matei_zaharia to show how it all comes together. https:// youtube.com/watch?v=sk6HBd mVmL8 …
Claude Code Changelog
Claude Code 2.1.240 has been released. 1 CLI change Highlights: • Bug fixes and reliability improvements to the CLI to reduce crashes and improve error messages • CLI fixes reduce crashes and improve command consistency, leading to fewer interruptions and steadier results All details in thread ↓
➖ Transformer의 Attention·FFN 손계산포스트 1
Transformer block을 다섯 개 입력 feature로 직접 계산하며 Attention의 위치 간 혼합과 FFN의 feature 차원 변환을 분리한 학습 자료가 공유됐다.
세부 내용 보기
- 학습 자료는 embeddings, positional encoding, self-attention, multi-head attention, layer norm 등 여러 구성 요소 가운데 Attention weighting과 feed-forward network가 핵심적인 두 부분이라고 보고, 이전 block에서 들어온 다섯 위치의 입력 feature를 한 block 안에서 처리한다. 먼저 query-key 모듈이 attention weight matrix A를 만들고, 입력과 A를 곱해 위치 사이 정보를 섞은 Z를 산출한다.
- 이후 다섯 개의 weighted feature를 FFN 첫 계층에 넣어 weight와 bias를 곱하고, 각 위치의 feature를 3개 숫자에서 4개로 확장한다. ReLU가 음수를 0으로 만든 뒤 두 번째 계층이 4차원을 3차원으로 줄이며, 결과는 다음 block으로 전달된다. 모든 위치가 같은 weight matrix를 통과하는 점이 position-wise 연산의 핵심이다.
- 이 계산에서 Attention은 가로 방향으로 이웃 위치 정보를 결합하고 FFN은 세로 방향으로 각 위치의 feature 차원을 결합한다. 두 과정이 차례로 반복되고 block마다 별도 parameter를 사용하므로, Transformer의 반복 구조가 위치 정보 혼합과 내부 표현 변환을 함께 수행한다.
📈 Honor Lightning 휴머노이드 로봇의 100m 주행포스트 1
스마트폰 제조사 Honor의 중국산 휴머노이드 로봇 Lightning이 100m를 9.32초에 달렸다는 기록과 충돌 뒤 기체 안정성에 관한 관찰이 공유됐다.
세부 내용 보기
- 포스트는 Honor의 휴머노이드 로봇 Lightning이 100m를 9.32초에 주파했다고 전한다. 비교 대상으로 Usain Bolt의 기록 9.58초가 함께 언급됐으며, 로봇의 달리기 속도를 사람의 세계 기록과 같은 단위로 놓고 평가하는 방식이다.
- 영상 속 로봇은 주행 뒤 안전벽에 충돌했지만 허리가 안쪽으로 무너지거나 불꽃이 튀는 현상은 나타나지 않았다고 포스트 작성자가 적었다. 따라서 입력은 주행 명령과 트랙이고 출력은 100m 기록 및 충돌 뒤 기체 상태이며, 보행 안정성이나 반복 주행 성능에 관한 추가 수치는 없다.
- 속도 기록과 충돌 장면이 한 사례에 묶이면서 휴머노이드 로봇 평가에서 단순한 최고 속도뿐 아니라 충돌 이후 구조적 안정성도 관찰 대상이 됐다. 다만 제공된 자료만으로 Lightning의 실제 양산성이나 장기 내구성을 판단할 수는 없다.
➖ AI 안전 대비와 California SB 53 입장 변화포스트 2
주요 AI 연구소의 rogue model 통제 계획 공개 수준과 OpenAI의 California SB 53 강화 요구가 안전 대비 이슈로 묶였다.
세부 내용 보기
- 한 포스트는 주요 AI 연구소가 rogue model을 통제하기 위한 공개 문서화 계획을 거의 갖추지 않았다는 새 연구를 전하며, AI system이 예상 밖의 잠재적 위험 행동을 보일 때 대비 수준에 의문을 제기했다. 다른 포스트는 OpenAI가 과거 반대했던 California AI safety bill SB 53을 강화하라고 요구했다는 변화를 전한다.
- 두 내용의 공통점은 모델의 예기치 않은 행동을 사전에 통제하는 내부 계획과 외부 규칙의 관계다. 첫 사례는 공개적으로 기록된 containment plan의 부족을 문제로 삼고, 두 번째 사례는 특정 법안에 대한 기업 입장 변화로 안전 기준을 둘러싼 제도적 대응을 드러낸다.
- 출처에는 해당 연구의 세부 방법, 연구소별 계획 수, SB 53의 구체 조항이나 OpenAI가 요구한 수정 내용이 담겨 있지 않다. 따라서 이 포스트 묶음에서 확인되는 범위는 안전 대비 문서화와 법안 입장 변화라는 두 신호다.
원문 트윗 2개 보기
TechCrunch
A new study finds leading AI labs have few publicly documented plans for containing rogue models, raising questions about preparedness as AI systems increasingly demonstrate unexpected and potentially dangerous behavior.
TechCrunch
OpenAI is calling for California to strengthen SB 53, an AI safety bill that the company previously opposed.
➖ AI 기능을 앞세운 제품·서비스의 지역 확장포스트 3
Runway의 LATAM 진출, Pixel 11 Pro XL의 AI 기능, AI-native 회계 플랫폼 Rillet의 대규모 자금 조달이 제품 확장 사례로 나타났다.
세부 내용 보기
- Runway는 Brazil 방문 이후 Chile에서 고객 및 기업, GobiernodeChile와 만남을 진행하며 LATAM 진출을 공식화했고, Japan·France·UK에 이어 사용자 기반을 넓히겠다고 밝혔다. Pixel 11 Pro XL은 Rambler를 포함한 AI 기능과 더 빠른 카메라를 내세웠지만, 최근 Pixel 사용자에게 업그레이드 동기를 줄 만큼 변화가 큰지는 의문으로 남았다.
- Rillet은 미국의 회계사 부족을 배경으로 AI-native accounting platform의 성장이 이어졌고, CEO Nicolas Kopp가 48시간 만에 1억 달러를 조달했다는 TechCrunch 포스트가 공유됐다. 세 사례는 각각 지역 시장 확대, 단말 기능 추가, 특정 직무를 겨냥한 플랫폼 자금 조달이라는 입력과 결과를 가진다.
- 제공된 포스트에는 Runway의 LATAM 매출이나 사용자 수, Pixel 11 Pro XL의 AI 기능별 성능, Rillet의 고객·매출 지표가 없다. 확인 가능한 변화는 제품 및 서비스가 새로운 지역과 직무 영역으로 확장되고 있다는 점이다.
원문 트윗 2개 보기

Cristóbal Valenzuela
Runway is officially coming to LATAM. Last month, we visited our customers in Brazil, and now it’s Chile We’re meeting with some of the most important companies and customers in Chile and also with @GobiernodeChile . We also have a few community events, so if you’re there, make sure to join us. LATAM is becoming an incredibly important user base for Runway. We’re already in Japan, France, and the UK, and we’ll continue our global expansion to make sure Runway reaches every corner of the world. Dos chilenos de vuelta a su país @matamalaortiz
TechCrunch
Google’s Pixel 11 Pro XL brings snappier cameras and genuinely useful AI features like Rambler, but its iterative upgrades may not be enough to tempt recent Pixel owners.
용어 해설
- 에이전트 하네스(Agent Harness)
- — AI agent가 여러 도구와 작업 흐름을 실행하도록 조정하는 운영 계층이다. 에이전트 전환, 컨텍스트 유지, 정책 적용, 비용 통제, 사람 승인 같은 기능을 한곳에서 관리해 반복적인 실행 환경 구성을 줄인다.
- 능동 추론(Active Inference)
- — 에이전트가 불완전한 작업 상태에 대한 믿음을 갱신하면서 다음 질문, 검색, 도구 호출, 작업 또는 종료 행동을 선택하는 방식이다. 예상 정보 이득과 비용을 함께 계산해 필요한 컨텍스트를 확보한다.
- pass@1
- — 에이전트가 한 번의 시도에서 작업을 올바르게 완료한 비율이다. Thinkingbox 벤치마크에서는 백엔드 상태를 기준으로 유효한 결과를 판정하며, 잘못된 효과나 누락된 효과도 실패에 포함한다.
- 피드포워드 네트워크(Feed-Forward Network)
- — Transformer의 각 위치에서 feature 차원을 변환하는 신경망 블록이다. Attention이 위치 사이 정보를 섞은 뒤, FFN은 각 위치의 feature 차원을 확장하고 ReLU를 거쳐 다시 축소한다.
AI 요약 · 북마크 · 개인 피드 설정 — 무료
출처 · 인용 안내
인용 시 "요약 출처: AI Trends (aitrends.kr)"를 표기하고, 사실 확인은 원문 보기 기준으로 진행해 주세요. 자세한 기준은 운영 정책을 참고해 주세요.
