TL;DR
최근 X 상의 논의는 에이전트·엣지·캐싱 최적화 중심으로 전개됐다. Grok 4.5는 Excel 내부 작동을 목표로 'Grok for Excel'을 공개해 재무 모델링·데이터 분석 워크플로에 모델을 직접 연결했다. Agent 성능 지표에서는 Kimi K3가 Agent Arena에서 상위권으로 도약해 실사용 에이전트 업무 성공률과 사용자 만족 지표에서 두드러진 결과를 보였다. NVIDIA는 Cosmos 3 Edge(4B 파라미터·엣지 실행 가능)를 오픈 소스로 공개해 로컬·물리 기반 월드 모델과 MCP 연결을 통한 창작·로보틱스용 에이전트 적용을 확장했다. 한편 KV 캐시의 한계를 겨냥한 CacheBlend/LMCache 논의와 LangChain·LangSmith·허깅페이스 툴체인의 에이전트 개발·평가 도구 발전이 맞물려, 실무에서의 처리량·비용·검증 루프 개선 흐름이 가속화되고 있다.
𝕏 실시간 트렌드 토픽
📈 Grok 4.5의 Excel 통합: 모델을 스프레드시트에 직접 연결포스트 3
Grok이 'Grok for Excel' 기능을 공개해 Excel 내부에서 Grok 4.5로 재무 모델링·시장 데이터 분석·차트 생성을 수행할 수 있게 됐다. 서비스 링크와 공개 발표가 동시에 올라 사용자 접근 경로가 명확해졌다.
- Grok for Excel은 모델 호출을 스프레드시트 워크플로에 직접 연결해 셀 기반 입력과 결과 생성을 통합한다.
- 엘론 머스크(Grok)와 공식 계정의 공지가 동시에 올라 접근성이 강조됐다.
- 커뮤니티 반응은 React 코드 작성 능력과 비용·토큰 효율성 관련 긍정적 평가가 섞여 있다.
스프레드시트 내부에서 모델을 직접 호출하면 데이터 이동·포맷 변환 비용이 줄고 실무자 워크플로가 단축된다.
모델 통합은 생산성 향상을 기대하게 하나, 비용·프라이버시·토큰 사용량 관리는 별도 검토가 필요하다.
원문 트윗 2개 보기

Elon Musk
@elonmusk
Grok for Excel is now live
Grok for Excel is live. Use Grok 4.5 to build financial models, analyze market data, and generate charts and graphs. Try it now https:// x.ai/grok/excel

Grok
@grok
Grok for Excel is live. Use Grok 4.5 to build financial models, analyze market data, and generate charts and graphs. Try it now https:// x.ai/grok/excel
🔥 Agent Arena와 Kimi K3: 에이전트 성능 지표의 이동포스트 2
Agent Arena의 최근 집계에서 Kimi K3가 프론트엔드·에이전트 작업에서 상위권으로 도약해 Confirmed Success와 Praise vs. Complaint 같은 사용자 만족 지표에서 강점을 보였다. 성능 분석은 수만 건 이상의 장기 에이전트 세션을 기반으로 했다.
- Kimi K3는 Frontend Web App Arena 등에서 Elo 상승과 다수 도메인 1위 등급을 기록했다.
- Arena 지표는 'Confirmed Success', 'Praise vs. Complaint', 'Steerability', 'Bash Recovery', 'Tool Hallucination'의 5가지 신호로 구성된다.
- 순위는 작업 성공률과 도구 사용 복구 능력 등 실사용 행태 기반 지표를 반영한다.
Kimi K3의 상위권 등장은 에이전트용 대규모 실사용 평가에서 실무적 성과가 개선되고 있음을 시사한다.
벤치마크 결과는 특정 작업·도구셋에 종속될 가능성이 있어 전 범위 일반화에는 추가 검증이 필요하다.
원문 트윗 2개 보기
Arena.ai
@arena
Exciting update: Kimi K3 has landed at #4 on the Agent Arena leaderboard, matching Claude Opus 4.8 and GPT-5.6 Sol. If Kimi K3's weights are released on schedule by July 27, it will become the #1 open-weight model. This release marks a major leap in agentic performance over Kimi K2.7 Code (#23 to #4). Based on 8K+ live agentic sessions, Kimi K3 leads on confirmed task success rate (#1). It also posts a strong +20.6% on praise vs. complaint (#3). It currently lags the field in steerability (#14) and bash recovery (#17). Agent Arena measures models on millions of real-world, long-horizon agentic tasks. Models get web search, filesystem, and terminal tools to complete complex workflows: writing code, creating slide decks, researching the web, building apps, and analyzing documents. We use causal tracing methodology to measure a model's net improvement, which indicates how much it improves outcomes relative to the average model. Here's a primer on the 5 signals: User-satisfaction proxies - Confirmed Success: an explicit "yes that worked" feedback from the user - Praise vs. Complaint: implicit sentiment in users reactions - Steerability: can the model course-correct when you push back? Tool-use proxies - Bash Recovery: how it recovers from CLI errors (primary signal for tool use) - Tool Hallucination: does it call tools that don't exist Below we break down how Kimi K3 scored across the 5 signals, drawn from tasks submitted by a global community of users. Congrats @Kimi_Moonshot on another big milestone!
Big news: Kimi-K3 by @Kimi_Moonshot is now #1 in the Frontend Code Arena with 1679 pts, surpassing Claude Fable 5. This is a 17-place jump from Kimi-k2.6 (#18 -> #1). In Frontend, Kimi-K3 ranked #1 in 6 of 7 domains: Brand & Marketing, Reference-Based Design, Data & Analytics, x.com/Kimi_Moonshot/…

Design Arena
@DesignArena
BREAKING: Kimi K3 by @Kimi_Moonshot is officially 1st on Frontend Web App Arena by DesignArena With an Elo of 1326, this open-weight model leads the way, ahead of Fable 5, Sonnet 5, and Opus 4.8 by @AnthropicAI Huge congrats to the @Kimi_Moonshot team for this achievement!
📈 NVIDIA의 Cosmos 3 Edge와 MCP: 엣지에서 구동 가능한 월드 모델포스트 8
NVIDIA는 Cosmos 3 Edge(4B 파라미터, Nemotron 기반 reasoner 포함)를 공개해 로컬·엣지 디바이스에서 월드 모델을 실행할 수 있게 했고, Omniverse·Omniverse 라이브러리 연동과 MCP(Model Context Protocol)로 창작 도구 내 에이전트 통합을 목표로 하고 있다.
- Cosmos 3 Edge는 로봇·자율주행·인프라용 월드 모델로 설계되어 DGX, Jetson 등에서 구동 가능한 것을 표방한다.
- MCP 연결을 통해 장면·에셋·타임라인 정보를 모델과 주고받아 창작자가 제어권을 유지한 채 에이전트를 도입한다.
- NVIDIA는 모델 가중치·레시피·코드를 오픈으로 제공한다고 발표해 현장 실험·검증이 가능하다.
엣지에서 실행 가능한 월드 모델과 MCP 연동은 로컬 제어·실시간 반응성이 필요한 로봇·크리에이티브 워크플로에서 실무적 가치를 제공한다.
오픈 가중치 공개는 투명성과 실험을 촉진하나, 엣지 배포를 위한 최적화·안전성 검증이 병행되어야 한다.
원문 트윗 2개 보기

NVIDIA AI
@NVIDIAAI
Introducing Cosmos 3 Edge: our open frontier world model built to run on-device. Cosmos 3 Edge helps robots learn and act, autonomous vehicles understand road scenes and predict intent, and vision AI agents reason across live video for smart infrastructure. With 4B parameters and a 2B Nemotron-based reasoner, you can run it on DGX Spark, NVIDIA Jetson, and more.
NVIDIA
@nvidia
NVIDIA’s research isn’t just focused on creating worlds that look real — but that behave realistically and respond in real time. Whether the output is a game, film, robot or factory digital twin, the goal is the same: expand the canvas of creativity with AI-generated worlds that are grounded in 3D, governed by physics and directed by creators. https:// nvda.ws/3TbQLkD #SIGGRAPH2026
➖ KV 캐시·CacheBlend·LMCache: 멀티문서·접두사 캐시 한계와 대안포스트 1
기존 접두사 기반 캐시는 문서 순서나 결합 변화에 민감해 재계산이 빈번했다. CacheBlend는 문서별 캐시를 재사용하고 경계 토큰만 선택적으로 재계산해 멀티문서 쿼리에서 처리량을 2~4배 향상시키는 접근을 제시했고, LMCache로 실무 통합이 진행되고 있다.
- 접두사 캐시는 정확한 바이트 단위 일치가 필요해 문서 순서 변경·RAG 조합에 취약하다.
- CacheBlend는 국지적 문맥 의존성만 재계산해 전체 재처리 비용을 크게 낮춘다.
- LMCache는 인퍼런스 엔진 외부에서 캐시를 관리해 다양한 런타임(vLLM, SGLang, TensorRT-LLM)과 연동된다.
문서 단위 재사용 전략은 멀티문서 RAG 시나리오에서 캐시 효율성과 응답 속도를 동시에 개선한다.
경계 토큰 식별·정합성 보장 등 구현 복잡도가 남아 있어 시스템 통합·테스트가 필요하다.
📈 에이전트 개발·평가 도구의 확장: LangChain·LangSmith·허깅페이스 연계포스트 8
LangChain의 도구·샌드박스 확장과 LangSmith의 Engine(트레이스 기반 클러스터링·수정 제안) 등 에이전트 개발·평가 인프라가 확장됐다. 허깅페이스·Novita 등과의 통합으로 모델·데이터 검색·배포 파이프라인이 터미널·CLI 중심으로 발전하고 있다.
- LangSmith Sandboxes 무료 제공과 IssueBench 같은 평가 스위트가 에이전트 반복 개선을 지원한다.
- 허깅페이스 확장(hf-find)과 Novita 모델 검색 커맨드는 에이전트가 모델·데이터를 자동으로 찾아 활용하도록 돕는다.
- LangSmith Engine은 트레이스 기반 클러스터링으로 문제를 식별하고 수정안을 제안하는 장기 평가 루프를 운영한다.
개발·평가 툴의 고도화는 에이전트 실무화에 필요한 반복적 디버깅·모니터링·모델 발견 체인을 강화한다.
툴 체인의 확장은 운영 복잡성을 높일 수 있어 조직별 통합 전략과 비용·프라이버시 고려가 필요하다.
원문 트윗 2개 보기

Harrison Chase
@hwchase17
LangSmith Engine is our in product agent that runs over traces, clusters them into issues, and proposes fixes It's a long running, complex, ambiguous process We wrote about how we evaluate it!
LangChain
@LangChain
We’re making LangSmith Sandboxes available to try for free. Start today: https:// smith.langchain.com
➖ Palantir AIP Evolve 사례: 비용·호출 횟수 최적화포스트 2
Palantir의 데모에서 AIP Evolve는 모델 스크리닝·프롬프트 재작성·결정론적 코드 전환을 통해 비용과 호출 횟수를 크게 낮추는 결과를 보고했다. 발표 수치는 '70% 낮은 비용·84% 적은 GPT 호출' 등으로 제시됐다.
- 모델 공급자 간 스크리닝과 전문가 피드백을 결합해 불필요한 모델을 배제하고 최적 후보를 선별했다.
- 프롬프트 재작성과 일부 작업을 결정론적 코드로 대체해 품질을 유지하면서 비용을 절감했다.
- 발표에는 의사결정 검증을 위해 전문가 선호도 측정 결과(예: 90% 전문가 선호)가 함께 제시됐다.
모델 선택·프롬프트 최적화·코드 전환의 조합은 비용과 호출 횟수를 동시에 줄이면서 출력 품질을 유지하는 실무적 접근이다.
사례 수치들은 특정 환경·워크플로에 기반하므로 타 환경 일반화를 위해선 추가 검증이 필요하다.
원문 트윗 2개 보기
Palantir
@PalantirTech
70% lower cost. 84% fewer GPT calls. 90% expert-preferred output. At DevCon 6, Dr. David Zihr, Medical Director for Physician Advising Services at Tampa General Hospital, and Palantir Forward Deployed Engineer Colton Rusch show how AIP Evolve achieved these optimizations by screening models across providers, rewriting prompts, and swapping AI for deterministic code on Tampa General's AI-generated utilization reviews.
Palantir
@PalantirTech
See how Palantir Software Engineer Colton Rusch uses AIP Evolve to prove that cheaper AI doesn't have to mean worse AI: "Evolve screened a whole suite of models across providers, then used evals and expert feedback to rule out certain options and keep others for further optimization." "We got to the point where evolve produced a vastly cheaper version — 68% lower compute costs." "We got lower costs and better quality at the same time."
AI 요약 · 북마크 · 개인 피드 설정 — 무료
출처 · 인용 안내
인용 시 "요약 출처: AI Trends (aitrends.kr)"를 표기하고, 사실 확인은 원문 보기 기준으로 진행해 주세요. 자세한 기준은 운영 정책을 참고해 주세요.