TL;DR
이번 기간에는 SpaceX의 Starship 13차 시험비행과 Starlink V3 위성 20기 배치가 관측되며 실제 운용 텔레메트리가 해상에서 수신된 사례가 보고됐다. X(Elon) 생태계에서는 Grok Build의 신기능(가이드·검색 개선·신뢰성 패치)과 Grok 4.5의 Augment 내 채택이 사용성 측면 변화를 가져왔다. Anthropic의 Claude Opus 5가 Perplexity에 통합되어 '대체 모델 대비 비용 57% 절감, 특정 벤치에서 상위권 성능' 주장이 제기된 반면, 대형 시스템 프롬프트 유출은 라우팅·안전성 논쟁을 촉발했다. Databricks는 데이터 중심 에이전트(Genie Code)가 문맥 활용으로 정확도 76.6%·작업당 비용 $0.55를 달성했다고 보고했고, NVIDIA의 오픈 웨이트 공개 지지 선언은 오픈 모델·안전성 논의의 정치적·산업적 전환점을 드러냈다.
𝕏 실시간 트렌드 토픽
🔥 Starship 13차 시험비행과 Starlink V3 위성 배치포스트 13
Flight 13에서 Starship은 해상에 부양한 채 텔레메트리를 송신했고, 같은 날 20기의 Starlink V3 위성 배치가 완료되어 위성 관측·데이터 수신이 병행됐다. Raptor 엔진의 핫스테이징 분리에 따른 부스터 복귀 관측과 열 차폐 개선(heat shield) 지표들이 공개됐다. 라이브 영상·다중 관측이 진행 상황을 실시간으로 가시화했다.
- Elon Musk가 Starship이 해상에 부양하며 텔레메트리 전송을 지속했다고 보고했다(Flight 13).
- SpaceX는 Raptor 엔진 점화와 핫스테이징 분리 장면을 공개했고, Super Heavy 부스터가 회수 지점으로 복귀하는 장면이 관측됐다.
- Starlink V3 20기 배치가 완료되어 위성 네트워크 시험이 동시에 진행됐다.
- 열 차폐(heat shield) 설계는 반복 비행을 거치며 개선되는 것으로 지적됐다.
Flight 13은 해상 부양 상태로 텔레메트리 전송이 확인되어 비행·회수·위성 배치의 통합 운용이 가시화됐다.
원문 트윗 2개 보기
📈 Grok 플랫폼 업데이트와 Grok 4.5 채택 흐름포스트 2
X의 Grok 개발팀은 Grok Build의 GUI·CLI 개선과 신뢰성 패치, Augment 내 Grok 4.5 노출을 통해 개발자·사용자 도구 경험을 보강했다. 가이드 투어, 스마트 검색, 실패 재개 같은 기능이 워크플로 신뢰성에 직접적 영향을 주었다. 모델 선택권 확대는 사용 패턴 변동을 유도했다.
- Grok Build 업데이트는 가이드/튜토리얼, 스마트 검색 도구 오버라이드, 실패한 실행의 원클릭 재개 등 워크플로 신뢰성 개선에 집중했다.
- Grok 4.5가 Augment에서 사용 가능한 옵션으로 보고되어 모델 선택의 전환이 관측됐다.
- 업데이트는 CLI 정책 변경과 함께 배포되어 개발자 도구 사용성에 실무적 영향을 미쳤다.
Grok Build의 기능 추가는 워크플로 안정성과 검색·디버깅 편의성을 개선해 개발자 운영 부담을 줄였다.
원문 트윗 2개 보기

Elon Musk
@elonmusk
Grok Build update http:// X.ai/cli
Grok Build new update brings a guided /tutorial tour, smarter search tool overrides, live workflow progress with one-click failed-run resume, and a wide reliability pass across workflows, voice, and tools Release Notes: v0.2.112 Breaking Changes: • CLI version policy now has x.com/XFreeze/status…

Grok
@grok
Try Grok 4.5 in Augment
@grok 4.5 saw the largest week-over-week increase in usage of any model in our picker. When the frontier is moving this fast, having your pick of models is a huge advantage: route to your favorite model today, not the top model from last quarter. More soon on what we've been
🔥 Claude Opus 5의 Perplexity 통합과 시스템 프롬프트 유출 논란포스트 5
Perplexity는 Claude Opus 5를 자사 서비스에 통합하며 '일부 벤치에서 상위권, 비용 57% 절감' 주장을 제시했고, Claude 관련 시스템 프롬프트 유출은 라우팅·안전 필터 설계의 내부 구조를 드러내며 논쟁을 촉발했다. 시스템 프롬프트 내부의 라우팅 블록은 특정 쿼리의 자동 재지정 등 민감한 동작을 포함하는 것으로 지적됐다.
- Perplexity는 Opus 5를 탑재해 평가에서 'Fable 5를 제외한 경쟁모델 중 상위' 성능과 비용 절감(57%)을 보고했다.
- 유출된 시스템 프롬프트에서는 'Fable'·'Mythos' 같은 라우팅·안전 계층이 존재하며, 민감 쿼리 일부를 다른 모델로 전환하는 로직이 포착됐다.
- Claude Code 관련 CLI·버전 업데이트 로그가 연속적으로 공개되며 안정성 개선 흐름이 병행됐다.
Perplexity 통합 결과는 Opus 5가 비용 대비 실무 성능에서 경쟁력이 있음을 시사했다(Perplexity 측 수치: 비용 57% 절감 주장).
시스템 프롬프트 유출은 내부 라우팅·안전성 설계가 어떻게 동작하는지 보여주며, 프롬프트 보안·투명성 문제를 제기했다.
원문 트윗 2개 보기

Aravind Srinivas
@AravSrinivas
Opus 5 is now available as an orchestrator model inside the Perplexity Computer harness for all Pro and Max users. It's almost as good as Opus but half the price. Congrats to @AnthropicAI for delivering world-class frontier models that excel at real-world research tasks!
Claude Opus 5 is now available in Perplexity and Perplexity Computer. We evaluated it against six other models on WANDR. It outperformed all but Fable 5, while being 57% cheaper.
Md Ismail Šojal
@0x0SojalSec
Claude Opus 5 system prompt leaked the cybersecurity routing is more aggressive than expected and the Fable safeguards section is wild. The Fable Opus routing is real. Fable 5 has dedicated safeguards for cybersecurity (and biology LLM R&D). When those trigger, the query is redirected to Opus 5 instead. Pliny dropped the full Claude Opus 5 system prompt (200k characters). A few things that stand out: - New Mythos tier above Opus (Project Glasswing, limited access) - Fable 5 and Mythos 5 share the same base model, but Fable has extra safety layers for cyber, biology, and LLM R&D - Sensitive queries on Fable can get silently routed to Opus 5 Anthropic’s own note admits the filters are tuned conservatively and still catch some harmless requests. Meanwhile their Claude Code team just said they cut 80% of the system prompt for these new models with no performance drop.
SYSTEM PROMPT LEAK WOW, talk about verbose_mode: enabled! Weighing in at a whopping 200,000 characters, here's the full system prompt for Claude Opus 5 A few interesting new blocks this time around, like the "<fable_safeguards_routing>" section Link to full
📈 NVIDIA의 오픈 모델 지지와 산업·정책적 파장포스트 4
NVIDIA는 한국 파트너십 발표와 함께 공개적으로 오픈 모델(오픈 웨이트) 지지를 표명했으며, 업계 리더들(CEO 서명 포함)의 공개 입장은 오픈 웨이트에 대한 정책·안전·혁신 논의를 가속했다. 이 변화는 공개 가중치의 채택·검증·안전체계 설계에서 산업적 전환을 촉발했다.
- NVIDIA는 한국과의 협력 발표에서 국가적 AI 인프라·생태계 구축 의지를 표명했다.
- Jensen Huang 서명 연대 및 주요 인사들의 공개 지지는 오픈 웨이트 논의의 정치적 정당성을 강화했다.
- 오픈 가중치 확산은 연구 재현성·주권·보안 논쟁을 동시에 불러일으킨다.
NVIDIA 등 주요 업계 리더의 공개 지지는 오픈 웨이트 채택을 촉진하는 산업적 전환점으로 작용했다.
오픈 웨이트 확산은 투명성과 혁신을 촉진하는 반면, 보안·안전 통제 설계가 병행되어야 한다는 균형적 과제를 남겼다.
원문 트윗 2개 보기
NVIDIA
@nvidia
This is the Golden Age of Korea. NVIDIA is proud to partner with leaders across government, industry, and global technology to advance the K-AI vision — across chips, frontier AI, physical AI and robotics, and AI factories. Together, we’re building AI infrastructure for Korea and the world.

Md Ismail Šojal
@0x0SojalSec
Three weeks ago, advocating for open weights still felt risky. Today, @JensenHuang , @finkd , and @satyanadella are publicly supporting them. This NVIDIA-signed letter is a real turning point for American AI leadership. Open weights accelerate progress without sacrificing safety or sovereignty.
For my first post, I’m sharing a letter @NVIDIA signed on why open models matter. AI will transform every industry, power every company, and be built by every country. Open models strengthen safety and cybersecurity, accelerate innovation and diffusion, and enable sovereignty.
➖ 데이터 중심 에이전트(Genie Code) 평가: 정확도와 비용 경쟁력포스트 1
Databricks 보고서는 Genie Code가 문맥 활용(시맨틱 검색·영구 메모리·워크스페이스 이해)을 통해 테스트된 코딩 에이전트 중 가장 높은 정확도(76.6%)와 가장 낮은 평균 비용($0.55/task)을 달성했다고 보고했다. 평가 방식은 실제 사용 사례에서 추출한 401개 과업을 동일 예산·환경에서 비교한 결과다.
- Genie Code는 401개 실사용 유래 과업에서 정확도 76.6%를 기록했고 평균 비용은 과업당 $0.55로 보고됐다.
- 성능 우위의 원인으로는 문맥 기반 검색과 영구 메모리, 작업 공간 이해가 지목됐다.
- 비교 대상은 동급 최신 모델들을 사용한 일반 코딩 에이전트였고, 동일 20분 작업 예산 하에서 평가가 진행됐다.
데이터 중심 에이전트는 문맥·검색·영구 메모리 결합으로 정확도와 비용 효율을 동시에 개선했다고 보고됐다.
평가는 동일 예산·특정 과업군에서 수행되었으며, 다른 과업군·운영 환경에서 결과가 달라질 수 있음을 시사한다.
📈 에이전트 컨텍스트 관리·OpenForgeRL: 배포·학습 현실성 개선포스트 2
Agentic Context Management 연구는 컨텍스트 축적의 비용-정확도 트레이드오프를 5가지 원칙(설계·수집·범위화·예측·응축)으로 분해했고, 응축 검증이 선형 비용으로 충실도 유지에 핵심임을 보고했다. OpenForgeRL은 실제 에이전트 하네스에서 훈련 데이터를 수집·재현해 적은 데이터로도 실전 배포 성능을 높이는 방법을 제시했다.
- Agentic Context Management는 무분별한 대화 누적이 토큰 비용을 제곱적으로 증가시킨다고 지적하고, 응축(compaction)을 통해 선형 비용으로 전환하는 방법을 제시했다.
- 참고 구현은 LongMemEval과 LoCoMo 같은 장기 기억 벤치에서 높은 점수(예: 92%, 93.2%)를 보고했다.
- OpenForgeRL은 실제 하네스 호출을 프록시로 기록하고 표준 RL 파이프라인으로 재학습해 적은 수의 작업으로도 배포와 유사한 성능을 달성했다.
컨텍스트 관리를 구조화하면 토큰 비용을 줄이면서 핵심 정보 충실도를 유지할 수 있으며, 응축 검증이 핵심 역할을 한다.
하네스에서 직접 데이터 수집·훈련(OpenForgeRL)은 배포 환경 격차를 줄여 작은 데이터로도 실전 성능 향상을 가능하게 했다.
원문 트윗 2개 보기
elvis
@omarsar0
// Agentic Context Management // Great read for the weekend. (bookmark it) Production agents fail less on reasoning and more on what sits in their context. Conversation history, big prompts, huge tool definitions, and ballooning tool outputs pile up every turn. The common response is storage and retrieval, a place to stash memories and look them up. New research argues that framing is too narrow and names a fuller discipline, Agentic Context Management. It decomposes into five primitives, architecting, ingesting, scoping, anticipating, and compacting with consolidation. Naive accumulation grows token cost with the square of conversation length. Crude summarization buys linear cost but hits an accuracy cliff. Only compaction validated against fidelity gets you linear cost without losing what matters. A reference implementation reports 92% on LongMemEval and 93.2% on LoCoMo. Why does it matter? Context management is becoming a first-class production concern that operates across an entire organization. The five primitives give you a structured way to reason about where your token budget goes. Paper: https:// arxiv.org/abs/2607.21503 Learn to build effective AI agents in our academy: https:// academy.dair.ai
DAIR.AI
@dair_ai
New research with Microsoft's and colleagues on training agents inside the harnesses they actually run in. (bookmark it) Why it matters: Agents today live inside elaborate harnesses like Claude Code, Codex, and OpenClaw. Those harnesses are hard to train end to end because open RL stacks cannot express stateful, multi-process harness inference. Most training happens in stripped-down environments that look nothing like deployment. OpenForgeRL closes that gap. A lightweight proxy serves the harness model calls while recording them as training data for a standard RL codebase, and a Kubernetes orchestrator runs each rollout in its own container. You train directly in the real harness, at scale. Using only hundreds to a few thousand tasks, OpenForgeGUI reaches 72.3 on WebVoyager, 63.0 on Online-Mind2Web, and 37.7 on OSWorld-Verified, beating open baselines of similar size and matching models several times larger. Harness choice turns out to be a training variable. Some harnesses are much harder to learn than others, and RL improves self-verification and multi-step completion while error recovery stays weak. That is a useful map for anyone doing agentic RL. Paper: https:// arxiv.org/abs/2607.21557 Learn to build effective AI agents in our academy: https:// academy.dair.ai
AI 요약 · 북마크 · 개인 피드 설정 — 무료
출처 · 인용 안내
인용 시 "요약 출처: AI Trends (aitrends.kr)"를 표기하고, 사실 확인은 원문 보기 기준으로 진행해 주세요. 자세한 기준은 운영 정책을 참고해 주세요.