TL;DR
이번 기간에는 OpenAI가 보안·모니터링·정렬 기준을 강화하기 위해 일부 Frontier RL 학습을 멈춘 소식이 가장 큰 기술 논점으로 떠올랐습니다. GLM-5.3은 Artificial Analysis Intelligence Index에서 Kimi K3와 같은 60점을 기록하고 API를 출시했지만, 이전 모델보다 토큰 사용량과 작업 비용이 늘었습니다. Grok Bot은 알림과 CLI 사용 사례를 넓혔고, Claude는 Gmail·Google Drive 작업과 Claude Code 기능을 확장했습니다. ClawGym II, GEPA, LangSmith Tuned Evaluators는 에이전트 실행을 관찰하고 학습·평가 과정에 사용자가 개입하는 흐름을 드러냈습니다. Tesla Semi는 최대 1.2MW 충전과 최대 500마일 주행거리라는 운송 기능을 내세웠습니다.
𝕏 실시간 트렌드 토픽
🔥 OpenAI Frontier RL 학습 중단과 안전 검증 강화포스트 9
OpenAI가 최신 모델의 일부 Reinforcement Learning 학습과 대규모 Frontier RL 실행을 보류하고, 격리·지속 보안 테스트·다단계 모니터링으로 안전 기준을 검증하는 흐름입니다.
- 모델 능력의 증가 속도가 안전·정렬 검증 속도를 앞설 수 있다는 우려로 OpenAI가 최신 배포 모델의 RL 학습을 2주 동안 일시 중단했고, 더 큰 Frontier RL 실행도 계속 보류한 상태입니다. 같은 기간 소규모 학습과 평가를 진행해 보호 장치의 작동 여부와 정렬 근거를 추가로 확보하는 방식입니다.
- OpenAI는 연구 환경의 workload·network isolation을 강화하고 지속적인 보안 테스트와 고위험 학습·평가·도구 사용 추론을 위한 다단계 모니터링을 도입했습니다. 이 체계는 우려되는 행동을 빠르게 감지하고 모델이 접근하거나 영향을 줄 수 있는 범위를 제한하도록 설계됐습니다.
- Sam Altman은 모델 발전이 매우 빠르며 안전에 대한 확신이 AI 발전 속도를 결정하게 될 것이라고 밝혔습니다. OpenAI는 향후 모델을 계속 출시할 전망이지만 이번 조치는 더 먼 시점의 릴리스에 영향을 주며, 업계 전체의 공통 안전 기준 조율도 필요하다는 입장입니다.
모델 능력이 안전 검증보다 빠르게 커지는 상황에서 대규모 학습을 멈추고 보안 격리·레드팀·모니터링을 먼저 검증하는 접근이 필요하다는 입장입니다.
이번 조치는 배포 모델의 일부 학습을 2주간 늦추고 더 큰 실행을 보류하는 운영 판단이며, 소규모 학습과 평가를 통해 안전 근거를 쌓은 뒤 추가 실행 여부를 정하는 구조입니다.
원문 트윗 2개 보기
OpenAI
@OpenAI
As models become more capable, the risks associated with developing and testing them internally also grow. We temporarily paused reinforcement learning (RL) training on our latest models intended for deployment for two weeks while we hardened and red-teamed our research environments and expanded monitoring coverage. Our largest planned frontier RL run remains on hold while smaller-scale training and evaluations validate these safeguards and establish more evidence of alignment.
Jakub Pachocki
@merettm
We temporarily slowed some frontier training to strengthen security and monitoring. Our largest planned frontier RL run remains on hold while smaller-scale training and evaluations help us test safeguards and gather more evidence of alignment. I expect confidence in safety to increasingly set the pace of AI development. We urgently need tools for labs and countries to coordinate on this, which is why I signed Pacing the Frontier. In the meantime, we’re taking practical steps ourselves - and will continue to share what we learn as our approach evolves. https:// openai.com/index/pacing-m odel-development-cyber-capabilities/ … https:// pacingthefrontier.com
📈 GLM-5.3의 오픈 웨이트 성능·비용 경쟁포스트 3
GLM-5.3이 Artificial Analysis Intelligence Index에서 60점을 기록하고 에이전트 업무 평가와 디자인 평가에서 순위를 높인 가운데, 토큰 사용량·작업 비용·웨이트 공개 일정이 함께 주목받았습니다.
- GLM-5.3은 Artificial Analysis Intelligence Index에서 60점을 받아 Kimi K3와 동률을 기록했고, 총 753B·활성 40B 파라미터와 1M 토큰 context window를 유지했습니다. GDPval-AA v2의 Elo는 GLM-5.2의 1524에서 1770으로 올랐고, Claude Opus 5 다음인 전체 2위로 제시됐습니다.
- 성능 향상에는 비용 증가가 따랐습니다. Intelligence Index 작업당 출력 토큰은 GLM-5.2의 15,700개에서 약 18,700개로 20% 늘었고, 작업당 비용은 $0.44에서 $0.68로 상승했지만 Kimi K3의 $0.84와 GPT-5.6 Sol의 $1.23보다 낮았습니다.
- AA-Omniscience 점수는 4에서 14로 높아졌고 정확도는 24%에서 34%, 시도율은 46%에서 55%로 올랐지만 환각률은 26%에서 30%로 소폭 상승했습니다. GLM-5.3 API는 코딩·방어적 사이버보안·장기 에이전트 작업을 대상으로 출시됐으며, 웨이트는 다음 주 안에 공개될 예정이라고 전해졌습니다.
원문 트윗 2개 보기
Artificial Analysis
@ArtificialAnlys
GLM-5.3 achieves 60 on the Artificial Analysis Intelligence Index, on par with Kimi K3 and up 7 points from GLM-5.2. Once the weights are released it will be tied as the leading open weights model @Zai_org has just launched GLM-5.3, which ties Kimi K3 (60) for the most intelligent open weights models on the Artificial Analysis Intelligence Index (assuming weights are released as per Z AI’s guidance). Between Kimi K3 and GLM-5.3, the open weights frontier is closer than ever to the proprietary frontier. GLM-5.3 keeps GLM-5.2's size at 753B total parameters and 40B active parameters. The Z AI team has shared that GLM-5.3’s weights are expected to be released within the next week. Key Takeaways: ➤ GLM-5.3 sees the largest gain in agentic capabilities, placing it amongst frontier models. The model's Elo rating in GDPval-AA v2, our real-world agentic knowledge work evaluation, rises from 1524 to 1770, a 246-point jump. That places GLM-5.3 second among all models, behind only Claude Opus 5 (1855), and surpassing the previous open weights leader on this evaluation, Kimi K3 (1668), by more than 100 points. ➤ GLM-5.3 is less token efficient than its predecessor. Across the Artificial Analysis Intelligence Index v4.1, GLM-5.3 uses roughly 18,700 output tokens per task, up about 20% from GLM-5.2 (15,700) and 27% more than Kimi K3 (14,700). This would have implications for cost of deployment. ➤ At $0.68 per Intelligence Index task, GLM-5.3 costs 1.5x GLM-5.2's $0.44, but is still cheaper than other models in the same intelligence tier. GLM-5.3 is 19% cheaper per task compared to Kimi K3 ($0.84) and 45% cheaper than GPT-5.6 Sol ($1.23). GLM-5.3’s increase in cost is partially driven by a 20% token usage increase compared to its predecessor. ➤ GLM-5.3 makes an improvement in real-world knowledge accuracy, scoring 14 on AA-Omniscience up from 4 for GLM-5.2. GLM-5.3 is now the second-best open weights model on AA-Omniscience, behind only Kimi K3 (20). A more detailed breakdown of where improvements are made reflects a genuine gain in accuracy as opposed simply more cautious abstention, as accuracy rate (24% to 34%) and attempt rate (46% to 55%) both increased. However, GLM-5.3’s hallucination rate regressed up slightly, from 26% to 30%. Additional model details: ➤ Size: 753B total parameters, 40B active (MoE), unchanged from GLM-5.2 ➤ Context window: 1M tokens ➤ Pricing: $1.40 per 1M input tokens and $4.40 per 1M output tokens. An 81% cache hit discount is applied ($0.26 per 1M cached input tokens) ➤ License: MIT ➤ Accessibility: GLM-5.3 is currently accessibly through Z AI’s first party API. The Z AI team has shared they expect to release the weights shortly Check out the full analysis of GLM-5.3 on Artificial Analysis: https:// artificialanalysis.ai
Z.ai
@Zai_org
GLM-5.3 API is now live. - Built for coding, defensive cybersecurity, and long-horizon agentic tasks - Priced the same as GLM-5.2 - Available via the official API and partner model gateways Get started: https:// docs.z.ai/guides/llm/glm -5.3 …
🔥 Grok Bot의 알림·CLI·업무 자동화 확장포스트 4
Grok Bot이 모바일 알림 개선과 CLI 사용 사례를 통해 일상 업무 자동화 도구로 소비되는 흐름입니다.
- Grok Bot은 모바일 알림을 Bot별로 묶고 각 Bot 전용 아이콘을 적용하는 변경을 배포했습니다. 알림의 출처를 Bot 단위로 구분하는 입력·표시 개선으로 여러 Bot을 함께 쓰는 상황의 확인 절차를 줄이는 방식입니다.
- 사용자 포스트에서는 Grok을 CLI에서 사용하며 특유의 응답 성격을 경험한 사례가 나왔고, 다른 사용자는 한 시간 안에 일상 업무의 25%를 자동화했다고 전했습니다. 다만 이 수치는 개인 사용자의 경험으로 제시됐으며 표준화된 벤치마크 수치는 아닙니다.
- Grok Bot을 에이전트로 활용해 여러 작업을 상시 실행할 수 있다는 사용 가이드와 제품 평가가 함께 확산됐습니다. 제품 업데이트, CLI 접근, 개인 업무 자동화 사례가 결합되며 단순 대화형 기능보다 지속 실행과 작업 위임이 관심의 중심이 됐습니다.
원문 트윗 2개 보기

Elon Musk
@elonmusk
Improvements to Grok @Bot
We've shipped several quality-of-life improvements to Grok Bot. Mobile notifications are now grouped by Bot and use their specific icon.
Jakub Pachocki
@merettm
We temporarily slowed some frontier training to strengthen security and monitoring. Our largest planned frontier RL run remains on hold while smaller-scale training and evaluations help us test safeguards and gather more evidence of alignment. I expect confidence in safety to increasingly set the pace of AI development. We urgently need tools for labs and countries to coordinate on this, which is why I signed Pacing the Frontier. In the meantime, we’re taking practical steps ourselves - and will continue to share what we learn as our approach evolves. https:// openai.com/index/pacing-m odel-development-cyber-capabilities/ … https:// pacingthefrontier.com
➖ Claude의 외부 서비스 실행과 개발자 도구 변화포스트 6
Claude가 Gmail·Google Drive에서 승인 기반 작업을 수행하고, Claude Code와 관련 실행 환경은 CLI·도구 호출·프롬프트 입력을 다루는 기능을 넓혔습니다.
- Claude는 connectors 메뉴에서 Gmail과 Google Drive를 연결한 뒤 이메일 스레드에 답장을 작성하고 전송하거나 파일을 관리할 수 있게 했습니다. 사용자가 승인 시점을 통제하는 구조이며 유료 요금제 전체에서 제공된다고 안내됐습니다.
- Claude Code는 주간 사용 한도를 50% 높인 기간을 연장했지만, 모델 수요가 강해 향후 몇 주간 용량이 빠듯할 수 있다고 밝혔습니다. 한도 확대와 공급 제약이 동시에 제시되면서 코딩 에이전트 사용량 관리가 제품 운영의 일부가 됐습니다.
- Claude Code 2.1.235에는 프롬프트 오탈자를 로컬 spellcheck로 표시하는 기능, 세션 간 과대 메시지를 사전에 거부하는 기능, 권한 부여 문구와 ‘withhold - don't ask again’을 명확히 표시하는 기능이 포함됐습니다. 이런 변경은 입력 정확도와 출력 손실 방지, 도구 권한 확인을 직접 다룹니다.
원문 트윗 2개 보기

Claude
@claudeai
Claude can now send emails in Gmail and manage files in Google Drive. Ask Claude to reply to a thread, and it drafts and sends the response. You control when it needs your approval. Connect Gmail or Google Drive from the connectors menu to try. Available on all paid plans.
Claude Code Changelog
@ClaudeCodeLog
Claude Code 2.1.235 has been released. 19 CLI changes Highlights: • Optional spellcheck underlines typos in prompts via local aspell/hunspell/ispell, improving prompt accuracy • SendMessage now rejects messages too large for cross-session delivery up front, preventing lost output • Dialogs show exact grant text and add 'withhold - don't ask again' when hidden, preventing accidental grants Full details are in thread ↓
📈 에이전트 학습·평가의 관찰 가능성과 사용자 개입포스트 6
ClawGym II, GEPA, LangSmith Tuned Evaluators, Bot Mode와 Miles가 에이전트 실행을 기록하고 평가하며, 학습 과정에 사람이 개입하는 도구 흐름을 확장했습니다.
- ClawGym II는 OpenClaw와 Claude Code를 불투명한 실행 상자로 취급하고, 모델 경계의 serving proxy에서 모든 호출을 수집한 뒤 다중 턴 구조를 prefix tree로 재구성합니다. 이 구조를 PPO와 GRPO에 넣어 Qwen3-30A3B의 Pass@1을 OpenClaw에서 9.98점, Claude Code에서 14.81점 높였고, 200~400회의 최적화 단계에서 안정적이었다고 보고됐습니다.
- GEPA는 사용자가 에이전트 실행을 깊게 관찰하고 최적화 궤적에 개입할 수 있도록 여러 조정 레버를 제공하는 방향으로 재구성됐습니다. 자동 최적화 결과를 그대로 받는 대신 실행 로그에서 문제를 찾고 수정하는 인간 조정 구조가 핵심입니다.
- LangSmith Tuned Evaluators는 운영 중인 에이전트 행동을 자동 채점하고 Perceived Error를 사용자 경험의 대리 신호로 표시합니다. 게시된 사례에서는 특화 평가 모델을 사용해 평가 비용을 약 82% 줄였으며, Bot Mode와 Miles는 역할·모델·메모리·스킬을 갖춘 Bot 구성과 대규모 RL 실행 검증을 지원하는 도구로 제시됐습니다.
원문 트윗 2개 보기
elvis
@omarsar0
Really interesting paper. I recommend it to anyone interested in training agents using existing harnesses. (bookmark it) ClawGym II runs RL through OpenClaw and Claude Code as opaque boxes. A serving proxy sits at the model boundary and captures every call the harness makes, then those calls get organized into prefix trees so PPO and GRPO can optimize over the recovered multi-turn structure. Qwen3-30A3B gains 9.98 points of Pass@1 through OpenClaw and 14.81 through Claude Code, stable across 200 to 400 optimization steps. Mix-harness training pushes further. One model gets optimized jointly by heterogeneous harnesses, which points at policies that generalize across execution systems instead of overfitting to a single one. Paper: https:// arxiv.org/abs/2608.16798 Track more trending AI papers in our academy: https:// academy.dair.ai
Git Maxd
@GitMaxd
New Tuned Evaluators from @LangChain Versioned judges you just 'turn on' in a tracing project The 'Perceived Error' Judge flags conversations where the agent probably messed up (proxy for user feedback!) They used a specially trained judge model & cut eval cost by ~82%
Introducing LangSmith Tuned Evaluators They automatically score agent behavior in production, starting with Perceived Error. Perceived Error is one of the clearest signals that your agent is giving users a helpful experience. In our benchmark, our specialized model
➖ Tesla Semi의 고출력 충전과 장거리 운송포스트 1
Tesla Semi 업데이트 페이지가 디젤 대비 운영 비용, 최대 1.2MW Megacharger, 최대 500마일 주행거리를 운송 차량의 핵심 사양으로 내세웠습니다.
- Tesla Semi는 디젤 트럭보다 에너지 비용과 유지보수 부품 수를 줄이는 운영 구조를 전제로 하며, 일반적인 트럭 보유 기간 안에 차량 비용을 회수할 수 있다고 안내됐습니다. 비교 기준은 연료·에너지 비용과 유지보수 부담입니다.
- Megacharger는 최대 1.2MW로 전력을 공급해 30분 안에 주행거리의 약 60%를 회복하도록 설계됐습니다. 고출력 충전 입력이 짧은 정차 시간 동안 배터리 주행 가능 거리를 보충하는 방식입니다.
- 게시된 최대 주행거리는 500마일입니다. 충전 속도와 주행거리, 유지보수 구조를 함께 제시해 장거리 운송에서 운행 중단 시간과 운영 비용을 줄이는 조건을 설명했습니다.
용어 해설
- Frontier RL 학습(Frontier RL Training)
- — 최첨단 모델의 능력을 높이기 위해 Reinforcement Learning을 적용하는 학습 단계입니다. OpenAI는 이 단계에서 모델 능력이 안전·정렬 검증 속도보다 앞서는 위험을 줄이기 위해 일부 대규모 실행을 일시 중단했습니다.
- 정렬(Alignment)
- — 모델의 행동이 사람이 의도한 목표와 안전 기준에 맞도록 만드는 과정입니다. 이번 사례에서는 모델의 능력 증가에 맞춰 보안, 모니터링, 위험 행동 탐지 체계를 강화하는 기준으로 사용됐습니다.
- 오픈 웨이트 모델(Open Weights Model)
- — 모델의 학습 가중치를 공개해 사용자가 직접 내려받거나 자체 인프라에서 운용할 수 있는 모델입니다. GLM-5.3은 공개 웨이트 모델 가운데 성능과 비용을 함께 비교하는 대상으로 다뤄졌습니다.
- 에이전트 능력(Agentic Capability)
- — 모델이 단일 답변을 넘어 여러 단계의 작업을 계획하고 도구를 사용해 완료하는 능력입니다. GLM-5.3과 ClawGym II 관련 포스트에서는 실제 지식 업무와 실행 시스템에서 이 능력을 측정했습니다.
- 접두사 트리(Prefix Tree)
- — 여러 호출에서 반복되는 앞부분과 분기 구조를 트리 형태로 저장하는 자료 구조입니다. ClawGym II는 harness가 모델에 보낸 다중 턴 호출을 접두사 트리로 재구성해 PPO와 GRPO 학습에 활용했습니다.
- 조정형 평가기(Tuned Evaluator)
- — 운영 중인 에이전트의 대화와 행동을 특정 기준으로 자동 채점하는 평가 모델입니다. LangSmith Tuned Evaluators는 Perceived Error를 감지해 사용자 경험을 대신 측정하고 평가 비용을 약 82% 낮추는 방식으로 소개됐습니다.
- Megacharger
- — Tesla Semi의 배터리를 고출력으로 충전하는 충전 설비입니다. 게시된 설명에 따르면 최대 1.2MW로 충전하고 30분 안에 주행거리의 약 60%를 회복하는 장치입니다.
AI 요약 · 북마크 · 개인 피드 설정 — 무료
출처 · 인용 안내
인용 시 "요약 출처: AI Trends (aitrends.kr)"를 표기하고, 사실 확인은 원문 보기 기준으로 진행해 주세요. 자세한 기준은 운영 정책을 참고해 주세요.
