본문으로 건너뛰기

오픈 웨이트 확산과 Grok·Opus 경쟁, Kimi K3의 delta attention

산업 리더들이 open-weight 공개를 지지하고 Grok·Opus 5의 문서 처리·통합 기능과 Kimi K3의 delta attention이 긴 컨텍스트 처리를 가속화하고 있다.

이 요약은 AI가 원문을 분석해 생성했습니다. 정확한 내용은 원문 기준으로 확인하세요.

TL;DR

최근 산업 리더들이 'open-weight' 모델 공개와 협력을 강조하며 오픈 중심의 생태계가 확산되는 흐름이 확인됐다. 동시에 Grok와 Opus 5가 문서·스프레드시트 처리 능력과 생산성 통합(예: Google Workspace 연동)으로 주목을 받았고, Kimi K3의 'delta attention'은 KV 캐시를 고정 크기 행렬로 바꿔 수십만~백만 토큰 수준의 컨텍스트를 선형 비용으로 처리하는 작동 원리를 보였다. ChatGPT는 사용자 맞춤 'Pet' 공유 기능과 지속적 작업환경·클라우드 컴퓨팅 자원(게시물 내 인용) 같은 에이전트형 기능을 확장했고, Perplexity·Hermes·Neon 등 도구·인프라 레이어에서도 에이전트 연동과 보안(credential proxy) 변화가 관측됐다. 전반적으로 성능·통합·개방성 추세가 맞물리며 혁신과 보안·정책의 균형이 핵심 쟁점으로 떠올랐다.

𝕏 실시간 트렌드 토픽

🔥 오픈 웨이트(모델 가중치 공개) 연대와 투명성 요구포스트 6

Satya Nadella·Jensen Huang 등 주요 리더들이 open-weight 모델의 중요성을 천명했고, Elon Musk는 X 시스템 코드 전라인 오픈소스화 계획을 발표했다. 이 움직임은 모델 접근성·확산·검증을 촉진하는 동시에 국가 안보와 책임 문제를 병행 고려해야 한다는 논의로 이어졌다.

세부 내용 보기
  • Satya·Jensen의 서한·발언은 오픈 모델이 경쟁력·확산·사이버보안 측면에서 이득을 준다고 지적했다.
  • Elon Musk는 X 플랫폼 코드의 전면 오픈과 서드파티 감사를 예고해 투명성 강화를 주장했다.
  • 일부 참여자는 오픈이 안전성과 주권을 강화한다고 본 반면, 안보·악용 리스크 관리 필요성이 함께 제기됐다.
찬성다수

오픈웨이트 공개는 검증·확산·생태계 경쟁을 촉진해 혁신과 주권을 강화한다는 주장

중립분열

오픈이 이득을 주나 국가 안보·악용 위험을 낮추기 위한 규제·감사 메커니즘 병행이 필요하다는 견해

원문 트윗 2개 보기

📈 Grok 4.x와 Opus 5의 성능·통합 경쟁포스트 7

Grok 4.5는 실전 청구서(invoice) 추출 벤치마크에서 상위 성과를 보고했고, Grok는 Google Workspace(문서·시트·슬라이드) 사이드패널 통합으로 직접 문서 내 작업을 실행하는 흐름을 보였다. Opus 5는 스프레드시트·슬라이드 생성 능력과 컨설턴트 수준 산출물 성능을 강조하며 여러 채널로 배포되고 있다.

세부 내용 보기
  • Grok는 Sheets에서 셀 기반 근거 인용·수식 작성·시나리오 재빌드, Slides·Docs에서 콘텐츠 생성·구조화 등의 워크스페이스 통합 기능을 제공한다.
  • Grok 4.5는 대규모 실제 청구서 집합 테스트에서 완전 추출(perfect-extraction) 비율이 높았다는 사례가 공개되었다.
  • Opus 5는 단기간에 스프레드시트·슬라이드 품질이 크게 향상되었다는 사용자·업계 반응과 함께 OpenCode·에이전트 허브로의 배포가 보고되었다.
찬성다수

문서·스프레드시트 처리에서 높은 정확도와 워크스페이스 통합은 실제 업무 생산성을 높인다는 주장

중립소수

성능 주장은 초기 벤치·사례 기반이며 비용·일반화 한계는 추가 검증이 필요하다는 관점

원문 트윗 2개 보기

X Freeze

@XFreeze

1달 전

Grok is now available directly inside Google Workspace One add-on brings Grok into Google Sheets, Slides and Docs, where it works beside the file you already have open In Sheets, Grok can: • Read selected ranges and explain the data • Cite the exact cells behind its answers • Write and fill real spreadsheet formulas • Update values and rebuild financial scenarios In Slides, it can: • Generate a complete presentation from an outline • Add new slides that match the existing theme • Restructure entire sections • Rewrite titles so every slide communicates a clear takeaway In Docs, it can: • Draft directly inside the document • Turn rough notes into formatted sections • Correct grammar and rewrite content • Apply consistent headings and styles across the file Everything happens through a side panel, without constantly copying information between tabs or rebuilding the output manually One installation works across all three apps, and organizations can deploy it across their entire workforce Grok is rapidly becoming a native intelligence layer for the tools people already use every day

💬 56 53 239👁 16290

Alex Albert

@alexalbert__

1달 전

Just over 6 months later, Opus 5 now produces near-superhuman level spreadsheets and slide decks that match what a consultant would make. Things are changing fast.

Alex Albert

I'm hearing from many folks across finance industry that Claude for Excel is blowing their minds. The agentic coding takeoff but for other fields is coming in 2026.

트윗에 첨부된 이미지
💬 17 31 548👁 42061

📈 Kimi K3의 'delta attention' — 긴 컨텍스트 처리 방식포스트 1

Kimi K3는 기존 KV 캐시(토큰별 key·value 리스트)를 고정 크기 행렬로 압축하는 'delta attention'을 도입해 컨텍스트 메모리를 일정하게 유지하면서 수십만~백만 토큰을 다루는 방식을 제시했다. 작동은 새 토큰의 키로 먼저 읽어 기존 추정값을 확인한 뒤 원하는 값과의 차분(delta)만을 기록해 행렬을 보정하는 것이다.

세부 내용 보기
  • 전통적 KV 캐시는 토큰 수에 따라 저장과 스캔 비용이 증가해 쿼드라틱 비용이 발생한다.
  • delta attention은 읽기-차분쓰기 두 단계로 고정 크기 행렬을 갱신해 메모리 사용을 일정하게 유지한다.
  • 압축에 따른 단일 토큰 회상은 근사적이므로정확한 조회가 필요한 층과 혼합(일부 레이어는 전체 어텐션 유지)해 사용한다.
찬성소수

KV 캐시를 고정 행렬로 대체하면 긴 컨텍스트를 선형 비용으로 처리할 수 있어 메모리·연산 효율이 개선된다는 주장

원문 트윗 1개 보기

Akshay

@akshay_pachaar

1달 전

one matrix replaced the KV cache. (the technique is 100% open source) Kimi just dropped K3, an open model at frontier scale. it leans on a new mechanism called delta attention that does not keep a growing KV cache. that is how it holds a million tokens of context without the memory blowing up. before we can understand delta attention, we need to understand attention itself. it is a lookup. every token stores a key, which works like an address, and a value, which is the content at that address. to build its output, a token sends out a query, matches it against every key in the sequence, and pulls back a blend of the values whose keys matched. that is the picture at the top of the diagram. standard attention keeps every one of those key and value pairs as a list, one entry per token. that list is the KV cache, and it grows with the sequence, so each new token has to scan the whole thing to build its output. double the context and you double both the storage and the scanning, which is where the quadratic cost and the ballooning memory come from. delta attention keeps the lookup but throws away the list. the entire past collapses into one fixed-size matrix that still behaves like the lookup table. hand it a key, it returns the value the past associated with that key. one thousand tokens or one million, the matrix stays the same size. the hard part is writing into it. you cannot staple a new pair onto a fixed matrix, because new writes land on top of old ones and blur them together. the delta rule solves this in two moves per token, shown at the bottom of the diagram. → it reads before it writes. it hands the matrix the new token's key and sees what value the memory currently returns, its existing guess for that address. → it writes the difference, not the value. it compares that guess against the value it actually wants stored and writes only the gap. that gap is the delta, and it corrects the old association instead of stacking a new one on top. the matrix also lets old entries fade over time, so a fixed size can keep absorbing a long sequence without filling up. that is the shift the diagram draws. standard attention remembers by keeping everything and pays quadratic cost to re-scan it. delta attention remembers by rewriting one matrix and pays linear cost. you give up something real. a compressed matrix cannot store every token exactly, so recall of any single token becomes approximate. that is why production models interleave the two, a few full-attention layers for exact lookup and the rest running linear. the difference is one verb. standard attention appends, delta attention corrects. read more: https:// kimi.com/blog/kimi-k3 for anyone who wants the full picture of KV caching, i have written a detailed article, which is quoted below. it explains how KV caching works from first principles, with illustrations for every step.

💬 1 0 2👁 953

ChatGPT의 'Pet' 공유 기능과 에이전트형 작업 환경 확장포스트 3

ChatGPT에서 만든 맞춤형 'Pet'을 웹에서 공유하는 기능이 활성화되었고, ChatGPT Work 관련 게시물은 클라우드 컴퓨팅(예: 15GB RAM), 지속적 작업공간·파일 저장 같은 에이전트형 사용 흐름을 부각시켰다. Pet 공유는 웹에서 링크 복사로 입양·공유가 가능하다는 운영 지침이 게시되었다.

세부 내용 보기
  • OpenAIDevs는 웹에서 custom pet의 공유 링크 생성 방법과 스프라이트시트 다운로드 옵션을 안내했다.
  • 관련 게시물은 모바일에서의 지속적 워크스페이스·클라우드 자원 제공이 에이전트 활용을 촉진한다고 언급했다.
  • 사용자 생성 에이전트의 확산은 에이전트 상태·보안·접근 제어 이슈를 동반한다.
찬성소수

공유 기능과 지속적 워크스페이스는 협업·재사용 측면에서 편의성을 높인다는 입장

원문 트윗 2개 보기

📈 개발자 도구·에이전트 인프라 변화: Perplexity CLI·IronProxy·Neon 업데이트포스트 3

Perplexity는 에이전트가 웹을 탐색하도록 설정하는 CLI 스킬을 공개했고, Hermes Agent에는 로컬 샌드박스용 credential firewall(IronProxy) 구현이 보고되었다. Neon은 Postgres 진단·스냅샷 등 CLI 기능 강화와 MCP 서버 로그 쿼리 확장을 발표했다.

세부 내용 보기

용어 해설

오픈 웨이트 모델(Open-weight models)
모델 파라미터(가중치)를 공개해 연구자·기업이 직접 다운로드하고 재사용할 수 있게 만드는 접근 방식이다. 공개는 모델 검증·재현성·확산을 촉진하고 생태계 경쟁을 유도하지만, 악용 위험·안보 이슈와 병행 정책이 필요하다는 논의가 따라붙는다.
델타 어텐션(Delta Attention)
KV 캐시를 길이 가변 리스트가 아닌 고정 크기 행렬로 압축해 컨텍스트를 유지하는 메커니즘이다. 새 토큰 처리 시 '읽기-차분 쓰기' 두 단계로 기존 연관성을 보정해 메모리 크기를 고정시키며, 이로써 긴 컨텍스트(예: 수십만~백만 토큰)를 선형 비용으로 다룰 수 있다.
KV 캐시(KV cache)
Transformer의 어텐션 연산에서 과거 토큰의 key·value 쌍을 저장해 다음 토큰 생성에 재사용하는 구조다. 전통적 구현은 시퀀스 길이에 비례해 저장용량과 스캔 비용이 증가해 컨텍스트 확장에 병목이 된다.
에이전틱 역량(Agentic capabilities)
모델이 도구 호출·워크플로 오케스트레이션·상태 유지 같은 연속적·자율적 작업을 수행하는 능력이다. 에이전트 기능은 생산성 향상과 복잡한 작업 자동화에 기여하나, 보안·검증·통제 문제를 수반한다.
AI 분석 전체 내용 보기

AI 요약 · 북마크 · 개인 피드 설정 — 무료

출처 · 인용 안내

원문 발행 2026. 07. 25.수집 2026. 07. 25.출처 타입 TWITTER

인용 시 "요약 출처: AI Trends (aitrends.kr)"를 표기하고, 사실 확인은 원문 보기 기준으로 진행해 주세요. 자세한 기준은 운영 정책을 참고해 주세요.