본문으로 건너뛰기
X (Twitter)조회 1

Grok 4.5·Kimi K3·Cosmos 3 Edge 중심의 에이전트·엣지 AI 최신 동향

Grok 4.5의 Excel 통합, Kimi K3의 Agent Arena 도약, NVIDIA의 Cosmos 3 Edge 공개로 에이전트·엣지·캐시 최적화 논의가 동시에 가속됐다.

이 요약은 AI가 원문을 분석해 생성했습니다. 정확한 내용은 원문 기준으로 확인하세요.

TL;DR

최근 X 상의 논의는 에이전트·엣지·캐싱 최적화 중심으로 전개됐다. Grok 4.5는 Excel 내부 작동을 목표로 'Grok for Excel'을 공개해 재무 모델링·데이터 분석 워크플로에 모델을 직접 연결했다. Agent 성능 지표에서는 Kimi K3가 Agent Arena에서 상위권으로 도약해 실사용 에이전트 업무 성공률과 사용자 만족 지표에서 두드러진 결과를 보였다. NVIDIA는 Cosmos 3 Edge(4B 파라미터·엣지 실행 가능)를 오픈 소스로 공개해 로컬·물리 기반 월드 모델과 MCP 연결을 통한 창작·로보틱스용 에이전트 적용을 확장했다. 한편 KV 캐시의 한계를 겨냥한 CacheBlend/LMCache 논의와 LangChain·LangSmith·허깅페이스 툴체인의 에이전트 개발·평가 도구 발전이 맞물려, 실무에서의 처리량·비용·검증 루프 개선 흐름이 가속화되고 있다.

𝕏 실시간 트렌드 토픽

📈 Grok 4.5의 Excel 통합: 모델을 스프레드시트에 직접 연결포스트 3

Grok이 'Grok for Excel' 기능을 공개해 Excel 내부에서 Grok 4.5로 재무 모델링·시장 데이터 분석·차트 생성을 수행할 수 있게 됐다. 서비스 링크와 공개 발표가 동시에 올라 사용자 접근 경로가 명확해졌다.

세부 내용 보기
  • Grok for Excel은 모델 호출을 스프레드시트 워크플로에 직접 연결해 셀 기반 입력과 결과 생성을 통합한다.
  • 엘론 머스크(Grok)와 공식 계정의 공지가 동시에 올라 접근성이 강조됐다.
  • 커뮤니티 반응은 React 코드 작성 능력과 비용·토큰 효율성 관련 긍정적 평가가 섞여 있다.
찬성다수

스프레드시트 내부에서 모델을 직접 호출하면 데이터 이동·포맷 변환 비용이 줄고 실무자 워크플로가 단축된다.

중립소수

모델 통합은 생산성 향상을 기대하게 하나, 비용·프라이버시·토큰 사용량 관리는 별도 검토가 필요하다.

원문 트윗 2개 보기

🔥 Agent Arena와 Kimi K3: 에이전트 성능 지표의 이동포스트 2

Agent Arena의 최근 집계에서 Kimi K3가 프론트엔드·에이전트 작업에서 상위권으로 도약해 Confirmed Success와 Praise vs. Complaint 같은 사용자 만족 지표에서 강점을 보였다. 성능 분석은 수만 건 이상의 장기 에이전트 세션을 기반으로 했다.

세부 내용 보기
  • Kimi K3는 Frontend Web App Arena 등에서 Elo 상승과 다수 도메인 1위 등급을 기록했다.
  • Arena 지표는 'Confirmed Success', 'Praise vs. Complaint', 'Steerability', 'Bash Recovery', 'Tool Hallucination'의 5가지 신호로 구성된다.
  • 순위는 작업 성공률과 도구 사용 복구 능력 등 실사용 행태 기반 지표를 반영한다.
찬성다수

Kimi K3의 상위권 등장은 에이전트용 대규모 실사용 평가에서 실무적 성과가 개선되고 있음을 시사한다.

중립소수

벤치마크 결과는 특정 작업·도구셋에 종속될 가능성이 있어 전 범위 일반화에는 추가 검증이 필요하다.

원문 트윗 2개 보기

Arena.ai

@arena

1달 전

Exciting update: Kimi K3 has landed at #4 on the Agent Arena leaderboard, matching Claude Opus 4.8 and GPT-5.6 Sol. If Kimi K3's weights are released on schedule by July 27, it will become the #1 open-weight model. This release marks a major leap in agentic performance over Kimi K2.7 Code (#23 to #4). Based on 8K+ live agentic sessions, Kimi K3 leads on confirmed task success rate (#1). It also posts a strong +20.6% on praise vs. complaint (#3). It currently lags the field in steerability (#14) and bash recovery (#17). Agent Arena measures models on millions of real-world, long-horizon agentic tasks. Models get web search, filesystem, and terminal tools to complete complex workflows: writing code, creating slide decks, researching the web, building apps, and analyzing documents. We use causal tracing methodology to measure a model's net improvement, which indicates how much it improves outcomes relative to the average model. Here's a primer on the 5 signals: User-satisfaction proxies - Confirmed Success: an explicit "yes that worked" feedback from the user - Praise vs. Complaint: implicit sentiment in users reactions - Steerability: can the model course-correct when you push back? Tool-use proxies - Bash Recovery: how it recovers from CLI errors (primary signal for tool use) - Tool Hallucination: does it call tools that don't exist Below we break down how Kimi K3 scored across the 5 signals, drawn from tasks submitted by a global community of users. Congrats @Kimi_Moonshot on another big milestone!

Arena.ai

Big news: Kimi-K3 by @Kimi_Moonshot is now #1 in the Frontend Code Arena with 1679 pts, surpassing Claude Fable 5. This is a 17-place jump from Kimi-k2.6 (#18 -> #1). In Frontend, Kimi-K3 ranked #1 in 6 of 7 domains: Brand & Marketing, Reference-Based Design, Data & Analytics, x.com/Kimi_Moonshot/…

인용 트윗 보기
💬 11 7 41👁 3726

Design Arena

@DesignArena

1달 전

BREAKING: Kimi K3 by @Kimi_Moonshot is officially 1st on Frontend Web App Arena by DesignArena With an Elo of 1326, this open-weight model leads the way, ahead of Fable 5, Sonnet 5, and Opus 4.8 by @AnthropicAI Huge congrats to the @Kimi_Moonshot team for this achievement!

💬 9 15 51👁 1120

📈 NVIDIA의 Cosmos 3 Edge와 MCP: 엣지에서 구동 가능한 월드 모델포스트 8

NVIDIA는 Cosmos 3 Edge(4B 파라미터, Nemotron 기반 reasoner 포함)를 공개해 로컬·엣지 디바이스에서 월드 모델을 실행할 수 있게 했고, Omniverse·Omniverse 라이브러리 연동과 MCP(Model Context Protocol)로 창작 도구 내 에이전트 통합을 목표로 하고 있다.

세부 내용 보기
  • Cosmos 3 Edge는 로봇·자율주행·인프라용 월드 모델로 설계되어 DGX, Jetson 등에서 구동 가능한 것을 표방한다.
  • MCP 연결을 통해 장면·에셋·타임라인 정보를 모델과 주고받아 창작자가 제어권을 유지한 채 에이전트를 도입한다.
  • NVIDIA는 모델 가중치·레시피·코드를 오픈으로 제공한다고 발표해 현장 실험·검증이 가능하다.
찬성다수

엣지에서 실행 가능한 월드 모델과 MCP 연동은 로컬 제어·실시간 반응성이 필요한 로봇·크리에이티브 워크플로에서 실무적 가치를 제공한다.

중립소수

오픈 가중치 공개는 투명성과 실험을 촉진하나, 엣지 배포를 위한 최적화·안전성 검증이 병행되어야 한다.

원문 트윗 2개 보기

KV 캐시·CacheBlend·LMCache: 멀티문서·접두사 캐시 한계와 대안포스트 1

기존 접두사 기반 캐시는 문서 순서나 결합 변화에 민감해 재계산이 빈번했다. CacheBlend는 문서별 캐시를 재사용하고 경계 토큰만 선택적으로 재계산해 멀티문서 쿼리에서 처리량을 2~4배 향상시키는 접근을 제시했고, LMCache로 실무 통합이 진행되고 있다.

세부 내용 보기
  • 접두사 캐시는 정확한 바이트 단위 일치가 필요해 문서 순서 변경·RAG 조합에 취약하다.
  • CacheBlend는 국지적 문맥 의존성만 재계산해 전체 재처리 비용을 크게 낮춘다.
  • LMCache는 인퍼런스 엔진 외부에서 캐시를 관리해 다양한 런타임(vLLM, SGLang, TensorRT-LLM)과 연동된다.
찬성다수

문서 단위 재사용 전략은 멀티문서 RAG 시나리오에서 캐시 효율성과 응답 속도를 동시에 개선한다.

중립소수

경계 토큰 식별·정합성 보장 등 구현 복잡도가 남아 있어 시스템 통합·테스트가 필요하다.

원문 트윗 1개 보기

Akshay

@akshay_pachaar

1달 전

90% of your KV cache never gets reused. (prompt caching was never meant to fix it) if your system prompt and tool definitions are stable, prompt caching is the single highest-leverage optimization available today. cached input tokens get up to 90% cheaper, and hit rates of 60 to 85% are realistic. but it comes with one rigid rule. the cached portion must be an exact, byte-for-byte prefix of the new request. change a single character in that region and you get a full cache miss. that rule breaks in three situations you hit constantly: → RAG with multiple documents. you cached document A alone and document B alone. a query now needs both. document B's cached state was computed without any awareness of A, so it's invalid and gets recomputed from scratch. → document order changes. the same three documents appear in a different order across requests. every permutation is a cache miss, even though the content is identical. → growing conversation history. each new turn changes everything after the stable prefix, so earlier cached states beyond it become useless. Alibaba Cloud's production data shows how bad this gets: 10% of KV cache blocks serve 77% of all hits. the rest sits in storage and never gets reused, because prefix matching won't allow it. CacheBlend, a research paper from the LMCache team (EuroSys 2025 Best Paper Award), attacks exactly this. the insight is that in modern transformers, tokens overwhelmingly attend to their own local context. only a small fraction of tokens carry real connections across document boundaries. so instead of recomputing everything after the first cached document, CacheBlend reuses every document's cache as-is and selectively recomputes just those few boundary tokens. those are the small orange fixes between documents in the diagram, and they are the entire cost. the result is 2 to 4x faster processing on multi-document queries with no quality loss. the order problem disappears with it. shuffle the same documents however you like, and every permutation stays cached, where prefix caching recomputes all of them every time. the bottom of the diagram shows that side by side. that's the real shift: from caching prefixes to caching knowledge. every document in your knowledge base becomes a reusable cached asset, regardless of what order it appears in or what sits next to it. CacheBlend ships inside LMCache, the open-source cache management layer that runs outside the inference engine and integrates with vLLM, SGLang, and TensorRT-LLM, on both NVIDIA and AMD GPUs. check it out on GitHub: https:// github.com/LMCache/LMCache (don't forget to star ) i wrote the full breakdown of the architecture, including why cache management should never live inside your inference engine. the article is quoted below. stay tuned for more on this!

💬 2 2 5👁 1285

📈 에이전트 개발·평가 도구의 확장: LangChain·LangSmith·허깅페이스 연계포스트 8

LangChain의 도구·샌드박스 확장과 LangSmith의 Engine(트레이스 기반 클러스터링·수정 제안) 등 에이전트 개발·평가 인프라가 확장됐다. 허깅페이스·Novita 등과의 통합으로 모델·데이터 검색·배포 파이프라인이 터미널·CLI 중심으로 발전하고 있다.

세부 내용 보기
  • LangSmith Sandboxes 무료 제공과 IssueBench 같은 평가 스위트가 에이전트 반복 개선을 지원한다.
  • 허깅페이스 확장(hf-find)과 Novita 모델 검색 커맨드는 에이전트가 모델·데이터를 자동으로 찾아 활용하도록 돕는다.
  • LangSmith Engine은 트레이스 기반 클러스터링으로 문제를 식별하고 수정안을 제안하는 장기 평가 루프를 운영한다.
찬성다수

개발·평가 툴의 고도화는 에이전트 실무화에 필요한 반복적 디버깅·모니터링·모델 발견 체인을 강화한다.

중립소수

툴 체인의 확장은 운영 복잡성을 높일 수 있어 조직별 통합 전략과 비용·프라이버시 고려가 필요하다.

원문 트윗 2개 보기

Palantir AIP Evolve 사례: 비용·호출 횟수 최적화포스트 2

Palantir의 데모에서 AIP Evolve는 모델 스크리닝·프롬프트 재작성·결정론적 코드 전환을 통해 비용과 호출 횟수를 크게 낮추는 결과를 보고했다. 발표 수치는 '70% 낮은 비용·84% 적은 GPT 호출' 등으로 제시됐다.

세부 내용 보기
  • 모델 공급자 간 스크리닝과 전문가 피드백을 결합해 불필요한 모델을 배제하고 최적 후보를 선별했다.
  • 프롬프트 재작성과 일부 작업을 결정론적 코드로 대체해 품질을 유지하면서 비용을 절감했다.
  • 발표에는 의사결정 검증을 위해 전문가 선호도 측정 결과(예: 90% 전문가 선호)가 함께 제시됐다.
찬성다수

모델 선택·프롬프트 최적화·코드 전환의 조합은 비용과 호출 횟수를 동시에 줄이면서 출력 품질을 유지하는 실무적 접근이다.

중립소수

사례 수치들은 특정 환경·워크플로에 기반하므로 타 환경 일반화를 위해선 추가 검증이 필요하다.

원문 트윗 2개 보기

용어 해설

KV 캐시(KV cache)
Transformer 기반 모델에서 이전 토큰의 키(key)·값(value) 행렬을 저장해 재사용하는 메모리 구조다. 동일한 컨텍스트를 반복할 때 전체 어텐션 계산을 줄여 처리 지연과 비용을 낮춘다. 캐시 관리는 캐시 적중률·일관성(문서 순서·접두사 변화)에 민감해 설계 난제가 된다.
CacheBlend
문서별로 계산된 KV 캐시를 그대로 재사용하고 문서 경계에서만 일부 토큰을 선택적으로 재계산하는 기법이다. 접두사 기반 캐시의 순서·조합 취약성을 완화해 멀티문서 쿼리 처리 속도를 크게 높인다. LMCache 구현체와 함께 배포되어 실무 적용 사례가 보고됐다.
Model Context Protocol(Model Context Protocol (MCP))
에디터·DCC 툴과 모델 간에 컨텍스트(장면, 에셋, 타임라인 등)를 주고받는 표준 인터페이스다. 에이전트나 모델이 창작 툴 내부 자산을 직접 조작하면서도 제작자가 제어권을 유지하도록 설계되어 창작 워크플로의 자동화·상호운용성을 개선한다.
월드 모델(World model)
환경의 상태·물리 법칙·행동 결과를 예측·시뮬레이션하는 모델 계층이다. 텍스트·이미지·비디오·사운드·액션을 통합해 현실 세계의 변화와 상호작용을 예측하며 로봇·자율주행·디지털 트윈에서 의사결정과 계획에 사용된다.
에이전트 하니스(Agent harness)
에이전트가 외부 도구·데이터·검증 루프와 상호작용하도록 연결·조율하는 실행 인프라다. 입력 파이프라인, 도구 호출 관리, 롤백·재시도 정책, 추적(trace) 및 평가를 포함해 에이전트의 실무적 신뢰성과 일반화를 높인다.
AI 분석 전체 내용 보기

AI 요약 · 북마크 · 개인 피드 설정 — 무료

출처 · 인용 안내

원문 발행 2026. 07. 21.수집 2026. 07. 21.출처 타입 TWITTER

인용 시 "요약 출처: AI Trends (aitrends.kr)"를 표기하고, 사실 확인은 원문 보기 기준으로 진행해 주세요. 자세한 기준은 운영 정책을 참고해 주세요.