TL;DR
이번 포스트에서는 Grok Bot이 Gmail·캘린더·파일·사진을 연결해 반복 업무와 미완료 작업을 찾는 개인용 Agent로 쓰이는 사례가 가장 큰 비중을 차지했다. GLM-5.3은 GLM-5.2와 같은 기본 모델·구조·전체 및 활성 파라미터를 유지한 채 장기 작업 환경과 RL을 한 달간 확장했고, Terminal-Bench 3.0 수치가 4.6%에서 32.4%로 상승했다. Qwen3.8-27B는 Cline의 1위 Local Model과 Harvey Legal Agent Benchmark의 1위 Open-Weight Model로 언급됐으며, Claude Desktop은 백그라운드 부팅 방식을 바꿔 시작 속도를 약 2배 높였다. AI Video 쪽에서는 Wan 3.0의 30초 영상·CLI Workflow와 Miora의 3D Camera Stage처럼 생성 전 장면 제어를 입력하는 기능이 이어졌다. 연구 포스트에서는 Debate Training의 Reward Hacking 감소, Diffusion Model의 Compute-Optimal Scaling, 학생 행동과 잠재 추론을 함께 학습하는 INSIDE가 새 평가축으로 묶였다.
𝕏 실시간 트렌드 토픽
🔥 Grok Bot의 개인 업무 자동화포스트 7
Grok Bot에 Gmail·캘린더·파일·사진을 연결해 구독·회의·환불 기한·미완료 약속을 찾는 사용 사례가 모였다. 자연어 요청만으로 여러 서비스의 기록을 읽고 후속 작업을 정리하는 Agent 활용이 핵심이다.
세부 내용 보기
- 기존에는 사용자가 메일과 일정, 결제 기록을 각각 열어 장기간 미뤄둔 일을 찾아야 했지만, Grok Bot은 Gmail·캘린더·파일·사진을 연결한 뒤 보관된 마케팅 메일, 유료 구독, 삭제할 스크린샷, 환불 기한, 연체된 예약을 항목별로 검색한다. 사용자는 취소·삭제·예약 같은 후속 조치의 후보를 한 번에 받을 수 있어 서비스별 수동 점검을 줄이는 구조다.
- Mike P의 사례에서는 두 Gmail 계정의 90,000개 메일을 정리 대상으로 지정했고, @grok은 “90k is nothing”이라고 응답했다. 별도 포스트는 부모 세대가 skills·plugins·MCP·Prompt Engineering을 익히지 않아도 채팅으로 요청하고 결과를 받는 사용성을 지적하며, 중국어 응답도 자연스럽다고 평가했다.
- Grok 4.6과 Grok Build를 Git Commit 등 작업에 쓰는 사례, $300/month SuperHeavy Plan에 Grok Bot·Grok 4.6·Grok Build·Imagine·Voice·DeepSearch·Cursor Ultra가 포함된다는 설명도 함께 나왔다. 다만 이 수치와 제품 평가는 해당 포스트 작성자의 경험·평가로 제시됐다.
원문 트윗 2개 보기
Todd Saunders
I’m currently having @bot do some things for me that I’ve been putting off forever. It’s pretty incredible. 1/ Unsubscribe me from every marketing email I archived in the last 120 days 2/ Find every subscription I’m paying for and tell me which ones I should cancel 3/ Go through my calendar and flag recurring meetings I should probably delete 4/ Find anything I bought recently that’s still inside the return window and remind me before it expires 5/ Find every person I told “let’s grab coffee soon” and never followed up with 6/ Find all the gift cards, credits, airline credits, and random balances I have sitting around unused 7/ Go through my photos and find all the screenshots I can delete 8/ Find every bill that has gone up significantly in the last year and tell me where I should negotiate or switch 9/ Find appointments I’m overdue for and help me schedule them 10/ Find everything in my inbox, calendar, and files that I said I would do but apparently never did

Elon Musk
Clear your email with @Grok @Bot
Grok Bot is going through 90,000 emails in my two gmail accounts and purging the bullshit. Something I've never dared to pursue myself. Good luck in there bud, do whatever you want. Just clean it tf up.
📈 GLM-5.3의 사후 학습 중심 확장포스트 2
GLM-5.3은 모델 크기를 키우는 대신 장기 작업 환경과 RL을 확장하는 방식으로 GLM-5.2와 같은 구조·파라미터 조건에서 성능 변화를 측정했다. Terminal-Bench 3.0에서는 4.6%에서 32.4%로 상승했다는 수치가 제시됐다.
세부 내용 보기
- 모델 확장에서 파라미터 수만으로 최적점을 정할 수 없다는 문제의식이 출발점이다. 데이터 규모, Compute 사용처, 실제 실행 주체와 조건을 함께 봐야 하며, Inference 비용이 학습 이후의 전체 비용을 지배하면 작은 모델을 더 오래 학습하는 방향으로 최적점이 이동한다는 설명이 이어졌다.
- Mixture of Experts에서는 전체 파라미터가 지식과 긴 꼬리 정보를 담는 규모에 가깝고, 활성 파라미터와 유효 깊이가 인과 사슬을 이어가는 추론 능력에 가깝다고 구분했다. 취약점 탐색처럼 여러 단계의 추론을 끝까지 유지해야 하는 작업은 단순한 CVE 암기보다 유효 깊이와 Post-Training의 영향을 받는다는 맥락이다.
- GLM-5.3은 GLM-5.2와 같은 Base·Architecture·Total Parameters·Activated Parameters를 유지하고, Long-Horizon Environment와 RL을 한 달 동안 확장한 통제 실험으로 제시됐다. 별도 벤치마크 포스트에서는 Terminal-Bench 3.0 점수가 GLM-5.2의 4.6%에서 GLM-5.3의 32.4%로 바뀌며 Grok 4.6을 대체했다고 전했다.
파라미터 수를 고정한 상태에서 Long-Horizon Environment와 RL을 확장하는 방식이 GLM-5.3의 성능 향상과 연결됐다는 입장이다.
Base Model Size, Pretraining Data, Forward Pass당 Compute 등 다른 확장 축도 남아 있어 단일한 최적 확장 비율로 결론 내리기 어렵다는 입장이다.
원문 트윗 2개 보기

jietang
Thoughts About Scaling Law Scaling, but not only of parameters. Every model release now ends with the same question: how many parameters? It isn't a question that can be answered on its own. Parameter count is only meaningful alongside three others — how much data you have, where you intend to spend your compute, and who will run the model, under what conditions. The field learned this the hard way. Kaplan et al. (2020) fit an exponent that told everyone to grow parameters faster than data — roughly 2.7:1 — and the industry complied: GPT-3, Gopher, MT-NLG. Hoffmann et al. (2022) redid the experiment across four hundred models and found the compute-optimal split is closer to 20 tokens per parameter, and that with sufficient compute the two should grow at the same rate rather than drifting apart. The error in the earlier fit compounded with every order of magnitude of compute, which is why the largest models of that generation were the most misallocated. The trillion-parameter round was, in retrospect, a detour the whole field took together and then reversed. Chinchilla wasn't the end either. It optimized training compute for models that would be trained once and evaluated. Today a model is called billions of times a day and inference dominates lifetime cost. Put inference into the objective and the optimum moves toward smaller models trained far longer — deliberate over-training, which is what Llama-2-7B and Gemma-2-9B were doing at roughly 290 and 889 tokens per parameter. Sparsity moved the target again. In a MoE model two quantities have to be kept apart: total parameters govern roughly how much the model can hold — knowledge, facts, the long tail — while activated parameters and effective depth govern roughly how far it can think, how many steps of a causal chain it can carry before it comes apart. A dense 20:1 ratio does not transfer. And the ratio isn't a single number at all: Roberts et al. (2025) find the optimal tokens-per-parameter is task-dependent, with memorization favoring more parameters and reasoning favoring more data. Follow-up work on MoE observes that at fixed TPP, pushing total parameters higher actually degrades reasoning, while activating more experts reliably helps it. This matters for what we are building toward. Finding a vulnerability is not a retrieval problem. It doesn't come from having memorized more CVEs; it comes from carrying a twenty-step chain of inference to the end without losing the thread. That capability does not live in total parameter count. Which brings us to this release. Total parameters appear to matter up to a threshold — enough to hold the world — after which additional capability comes from scaling elsewhere: effective depth per forward pass, and above all post-training. GLM-5.3 is our controlled experiment on that claim. Same base, same architecture, same total and activated parameters as GLM-5.2. One month of scaling long-horizon environments and RL. The gains are not marginal. Well, scaling has more than one dial. We turned the post-training one this time because it had the most slack left in it — not because the others are finished. Base model size, pretraining data, compute spent per forward pass: all of them are still on the table, and we will come back to each. What this experiment taught us is that the dials do not have to be turned together, and that the one worth turning next is rarely the one that was worth turning last. We are not done scaling. Next time, maybe mid-training, pre-training, and even more.
Ryan Marten
GLM 5.3 is the newest member to the cost-accuracy pareto for Terminal-Bench 3.0, replacing Grok 4.6 Impressive improvement from GLM 5.2 (4.6%) to GLM 5.3 (32.4%)!
📈 Qwen3.8-27B의 로컬·법률 Agent 성능포스트 2
Qwen3.8-27B가 Cline의 Local Model 순위에서 Qwen2.5-Coder-7B의 4개월 연속 1위를 4일 만에 대체했다. Harvey Legal Agent Benchmark에서는 Fable 5와 11.3점으로 공동 1위를 기록했다는 평가가 나왔다.
세부 내용 보기
- Cline의 Local Model 순위는 Qwen3.8-27B가 공개된 뒤 4일 만에 1위에 올랐고, 이전 1위였던 Qwen2.5-Coder-7B의 4개월 연속 기록이 종료됐다. 모델 출시와 실제 개발 도구 사용 결과가 짧은 기간 안에 순위 변동으로 연결된 사례다.
- Harvey Legal Agent Benchmark에서는 Qwen3.8-27B가 Fable 5와 11.3점으로 공동 1위를 기록했고 Kimi K3, Qwen 3.8 Max, DeepSeek V4보다 앞섰다는 비교가 인용됐다. Alibaba_Qwen은 이를 Professional Task를 처리하면서도 Local Machine에서 실행할 수 있는 규모로 묘사했다.
- 두 포스트가 제시한 축은 서로 다르다. Cline 결과는 개발 Workflow 안의 Local Model 선호를, Harvey 결과는 Legal Agent 작업 점수를 측정하므로, 단일 종합 순위보다 실행 환경과 작업 유형별 성능 변화로 읽는 편이 적절하다.
원문 트윗 2개 보기
Qwen
4 days to the top.Thank you to every builder who pushed Qwen3.8-27B to #1 on Cline! @cline
Qwen3.8-27B is now the #1 local model in Cline after just 4 days. This ends a *4 month* streak by the previous winner Qwen2.5-Coder-7B which has been the top local model since April. x.com/cline/status/2…
Qwen
#1 open-weight model on Harvey's Legal Agent benchmark! Strong enough to handle professional tasks. Small enough to run on your local machine. Qwen3.8-27B is becoming part of your everyday workflows.
On @harvey's Legal Agent, it is tied with Fable 5 at 11.3pts and is ahead of Kimi K3, Qwen 3.8 Max, and DeepSeek V4.
➖ Claude Desktop 백그라운드 부팅 최적화포스트 1
Claude Desktop이 숨겨진 상태로 시작할 때 Timer Throttling과 JavaScript Engine의 Power-Saving Mode가 부팅 지연을 만들던 구조를 바꿨다. 창이 보이지 않아도 Full Speed로 부팅하고 추가 Performance Fix를 적용해 시작 속도가 약 2배 빨라졌다.
세부 내용 보기
- Desktop을 매일 쓰는 환경에서는 앱이 열리기 전의 지연도 전체 사용감에 영향을 준다는 문제가 제기됐다. Claude Desktop은 백그라운드에서 시작될 때 Timer가 제한되고 JavaScript Engine이 절전 모드로 들어가 초기화가 늦어지는 경로를 수정했다.
- 새 방식은 창이 아직 숨겨져 있어도 JavaScript Engine을 Full Speed로 부팅하고, 그 위에 소규모 Performance Fix를 추가하는 처리다. 사용자가 창을 확인하는 시점에 초기화가 이미 더 진행된 상태가 되도록 백그라운드 실행 조건을 바꾼 셈이다.
- ClaudeDevs 포스트는 한 달 전보다 시작 속도가 약 2배 빨라졌다고 전했고, Claude Desktop 담당자는 매일 사용하는 앱에서 이런 작은 Quality-of-Life 개선을 계속 적용한다고 밝혔다.
📈 Wan 3.0·Miora의 AI Video 장면 제어포스트 4
Wan 3.0은 30초 Music Video와 Web URL 기반 Product Commercial, CLI Workflow를 생성하는 사례로 소개됐다. Miora는 3D Camera Stage에서 경로·추적 대상·저장 카메라 위치를 지정해 AI Video의 장면 실행을 직접 제어한다.
세부 내용 보기
- 텍스트로 “camera pushes in from wide to close-up”처럼 지시하고 결과를 기다리는 방식의 불확실성이 문제로 제기됐다. Miora는 3D Camera Stage에서 편집기에 보이는 장면 제어를 AI가 실행하도록 만들어, 지면에 경로를 그리면 캐릭터 이동과 Timeline 동기화가 이어지게 한다.
- 사용자는 Follow Target을 지정해 카메라가 대상을 추적하게 하고, 최대 10개의 Camera Position을 저장한 뒤 한 번의 클릭으로 각도를 바꿀 수 있다. 3D Stage에서 Camera Move를 촬영하고 Reference Image와 Prompt를 함께 전달하는 입력 흐름이다.
- Wan 3.0 관련 포스트는 30초 Music Video 생성, Web URL을 Product Commercial로 변환, 새 CLI를 통한 Workflow 자동화를 사례로 들었다. Miora와 Wan 3.0 모두 단순한 텍스트 생성보다 장면 구성과 제작 단계의 제어 범위를 넓히는 방향이다.
원문 트윗 2개 보기
Tencent AI
Telling AI "camera pushes in from wide to close-up" and hoping it understands? yeah, we got tired of that too. Shipping a 3D camera stage in Miora. What you see in the editor is what the AI executes: — draw a path on the ground, the character walks it, timing syncs to your timeline automatically. — set a follow target, the camera locks on and tracks your subject through the scene. — up to 10 saved camera positions, switch angles in one click. Shoot a camera move in the 3D stage, hand Miora a reference image and a prompt. Try it → https:// miora.design
Wan
Huge shoutout to @towya_aillust for putting together this incredible deep-dive into Wan 3.0! If you want to see exactly what the future of AI video looks like, you really cannot miss this. From generating full 30-second music videos, to seamlessly turning a web URL into a product commercial, and even automating workflows using our new CLI—this video covers it all. Grab your popcorn and take notes!
📈 Debate Training의 Reward Hacking 완화포스트 1
Google DeepMind의 새 연구는 Generator와 Critic이 대립하는 2인 Debate Game을 RL 학습에 넣어 RLAIF 기준선보다 Reward Hacking을 줄이는 방법을 실험했다. 약한 LLM Judge가 양측의 결과를 판정하는 구조다.
세부 내용 보기
- RLAIF에서는 모델이 사람이나 AI의 실제 선호보다 보상 신호의 허점을 공략할 수 있다는 문제가 있다. 연구는 Generator와 Critic을 서로 대립시키는 2인 Debate Game을 구성해 한 모델의 결과를 다른 모델이 반박하도록 했다.
- 학습 과정에서 Generator가 답을 만들고 Critic이 그 답의 문제를 공격하며, 약한 LLM Judge가 Debate 결과를 판정한다. 이 판정 신호를 RL Fine-Tuning에 사용해 단순한 RLAIF 기준선과 Reward Hacking 발생 정도를 비교하는 흐름이다.
- 포스트가 인용한 연구의 결론은 Debate를 사용한 RL Fine-Tuning이 RLAIF 기준선보다 Reward Hacking을 줄였다는 것이다. 구체적인 감소율은 해당 포스트에 제시되지 않았다.
➖ 로컬 실행과 온디바이스 검색 인프라포스트 1
Qdrant Edge를 사용한 Recall Offline은 휴대전화 내부에서 실행되는 In-Process Vector Engine으로 모델과 질의를 로컬에 유지한다. 모델을 한 번 내려받은 뒤 각 Query가 기기 안에서 처리되는 구조다.
세부 내용 보기
- 네트워크 연결이나 원격 검색 서버에 의존하지 않고 휴대전화에서 검색하려는 요구가 배경이다. Recall Offline은 Qdrant Edge를 앱 내부에 넣어 Vector Engine을 별도 서버가 아닌 기기 프로세스로 실행한다.
- 사용자가 모델을 한 번 다운로드하면 이후 Query가 로컬에서 Embedding과 검색 단계로 이어지고, 외부 서버로 질의를 보내지 않는다. 포스트는 이 구조를 “every query stays local”이라고 설명했다.
- 게시물은 Qdrant Edge 기반 구현과 Recall Offline 앱 링크를 함께 제시했지만, 검색 정확도·지연·저장 용량 수치는 제공하지 않았다. 따라서 이번 사례에서 확인되는 핵심은 온디바이스 실행 경로와 데이터 보관 위치다.
용어 해설
- 사후 학습(Post-Training)
- — 기본 모델을 완성한 뒤 장기 작업 환경과 RL 같은 추가 학습으로 모델의 추론·행동 성능을 조정하는 과정이다. 이번 포스트에서는 GLM-5.3이 같은 모델 규모에서 Post-Training 확장으로 성능을 높인 방식과 연결된다.
- 전문가 혼합 구조(Mixture of Experts)
- — 여러 Expert 중 일부만 활성화해 입력을 처리하는 모델 구조다. 전체 파라미터는 저장 지식의 규모와, 활성 파라미터·유효 깊이는 한 번의 순전파에서 이어갈 추론 단계와 관련된다는 구분이 제시됐다.
- 보상 해킹(Reward Hacking)
- — 모델이 실제 목표를 달성하지 않고 평가 보상만 높이는 현상이다. Debate 방식에서는 생성기와 비평가가 대립하고 약한 LLM Judge가 판정하는 구조로 RLAIF 기준선보다 이 현상을 줄이는 학습 효과를 측정했다.
- Flow Matching
- — 텍스트 조건에 맞춰 이미지 생성 과정을 학습하는 확산 계열 방법이다. Abra 연구는 Flow-Matching Transformer를 여러 Compute 규모에서 학습해 Text-to-Image 모델의 최적 데이터·파라미터 비율을 산출했다.
- 벡터 엔진(Vector Engine)
- — Embedding을 저장하고 유사도 검색을 수행하는 실행 계층이다. Qdrant Edge는 휴대전화 내부 프로세스로 실행되며 모델을 한 번 내려받은 뒤 질의 데이터를 기기 밖으로 보내지 않는 구조를 사용한다.
- 장기 작업 환경(Long-Horizon Environment)
- — 여러 단계의 추론과 행동을 오래 이어가는 학습 환경이다. GLM-5.3 관련 포스트는 한 달 동안 이 환경과 RL을 확장해 같은 기본 모델·구조·파라미터 수에서 성능 변화를 측정했다.
AI 요약 · 북마크 · 개인 피드 설정 — 무료
출처 · 인용 안내
인용 시 "요약 출처: AI Trends (aitrends.kr)"를 표기하고, 사실 확인은 원문 보기 기준으로 진행해 주세요. 자세한 기준은 운영 정책을 참고해 주세요.
