본문으로 건너뛰기
X (Twitter)조회 1

GLM-5.3-Flash 공개, 저비용 컴퓨터 사용 모델과 외부 AI 연구 데이터 개방

GLM-5.3-Flash의 오픈 웨이트·저비용 추론, Qwen3.8-Flash의 로컬 실행, Navigator n2의 데스크톱 자동화

이 요약은 AI가 원문을 분석해 생성했습니다. 정확한 내용은 원문 기준으로 확인하세요.

TL;DR

이번 기간에는 GLM-5.3-Flash가 Ox Alpha의 정식 모델명으로 확인되면서 중국산 AI 칩 기반 서비스 규모, 오픈 웨이트 공개, 가격, 벤치마크 성능이 한꺼번에 주목받았습니다. 320B 전체 파라미터와 18B 활성 파라미터를 사용하는 이 모델은 Artificial Analysis Intelligence Index에서 57점을 기록했고, $0.09 Cost per Task로 GLM-5.3의 $0.68보다 낮은 비용을 보였습니다. Qwen3.8-Flash의 75GB RAM 로컬 실행, Navigator n2의 데스크톱 자동화, Gemini 3.5 Transcribe의 85개 이상 언어 음성 처리도 모델을 실제 작업 환경에 배치하는 흐름으로 묶였습니다. 한편 Anthropic은 개인정보를 보존한 Claude 사용 데이터를 외부 연구자에게 개방했고, Perplexity는 세션·파일·출처를 구조화한 기억 시스템으로 정확성과 회수율을 높였다고 밝혔습니다.

𝕏 실시간 트렌드 토픽

🔥 GLM-5.3-Flash의 정식 공개와 비용 대비 성능포스트 15

Ox Alpha로 알려졌던 모델이 GLM-5.3-Flash로 공개되면서 구조, 중국산 AI 칩 실행, 가격, 라이선스와 여러 평가 결과가 집중됐습니다. 작은 활성 파라미터와 낮은 토큰 가격을 결합해 비용 대비 성능 축에 놓였다는 평가가 이어졌습니다.

세부 내용 보기
  • Ox Alpha로 배포되던 모델의 정체가 GLM-5.3-Flash로 확인됐고, Z.ai는 320B-A18B 구조와 MIT License, 네이티브 멀티모달 기능을 공개했습니다. SemiAnalysis는 하루 100T tokens 처리량이 중국산 칩에서 제공된다는 점을 핵심 변화로 짚었으며, 모델 구조 설명에서는 Kimi Delta Attention, Multi-head Latent Attention, DeepSeek Sparse Attention, sparse Mixture-of-Experts, mHC residual path와 native vision encoder가 함께 언급됐습니다.
  • 전체 320B 파라미터 중 18B만 활성화하는 구조와 낮은 API 가격이 비용을 낮추는 방식입니다. Artificial Analysis Intelligence Index에서 max reasoning effort 기준 57점을 기록했고, Cost per Task는 $0.09로 GLM-5.3의 $0.68보다 약 7.5배 낮았으며, Terminal-Bench v2.1에서는 84.3%로 GLM-5.3의 83.9%와 비슷한 결과를 냈습니다.
  • Z.ai API의 가격은 입력 $0.15 / 1M tokens, 출력 $0.50 / 1M tokens로 제시됐고 cached input은 $0.03/ 1M tokens로 안내됐습니다. GLM-5.3-Flash는 Cline, OpenCode Go, AutoClaw 등 여러 서비스에 빠르게 연결됐으며, Cline에서는 출시 일주일이 안 돼 전체 트래픽의 11% 이상을 차지했다는 수치가 공유됐습니다.
찬성다수

320B 전체 파라미터와 18B 활성 파라미터, MIT License, 낮은 API 가격을 함께 제공해 오픈 웨이트 모델의 비용 대비 성능을 높였다는 평가입니다.

중립소수

Arena의 초기 AutoEval 점수와 여러 벤치마크 결과는 강점을 시사하지만, 일부 평가는 라이브 사용자 투표가 더 쌓여야 수렴 여부를 판단할 수 있다는 입장입니다.

원문 트윗 2개 보기

Artificial Analysis

@ArtificialAnlys

1일 전

GLM-5.3-Flash scores 57 on the Artificial Analysis Intelligence Index. At $0.09 Cost per Task, it sits comfortably on the Intelligence vs. Cost per Task Pareto frontier @Zai_org has released GLM-5.3-Flash, a smaller and cheaper sibling to GLM-5.3 at 320B total parameters and just 18B active parameters. GLM-5.3-Flash supports low/high/max reasoning efforts, and scores 57 evaluation on the Artificial Analysis Intelligence Index with max reasoning effort. This places the model only 3 points behind GLM-5.3 at 60 and in line with GPT-5.6 Terra and Muse Spark 1.2. On Z AI's first-party API, GLM-5.3-Flash is priced at $0.15 / 1M input tokens and $0.50 / 1M output tokens, just over 10% of the price of GLM-5.3. Cached input tokens are priced at $0.026 / 1M tokens, an 80% discount. Its Cost per Task on the Intelligence Index is $0.09, compared to $0.68 for GLM-5.3 (max), and it sits on the Pareto frontier for Intelligence vs. Cost per Task. Key results: ➤ GLM-5.3-Flash is 3 points behind GLM-5.3 (max) on the Artificial Analysis Intelligence Index, at ~7.5x lower Cost per Task. At $0.09 per Intelligence Index task against $0.68 for GLM-5.3, it sits on the Pareto frontier for Intelligence vs. Cost per Task. It ties GPT-5.6 Terra ($0.51) and Muse Spark 1.2 ($0.40) at 57 while costing ~5.7x and ~4.4x less per task. ➤ GLM-5.3-Flash is less token efficient, but its low per-token pricing means this does not translate into a high Cost per Task. The model used 149M output tokens to run the Intelligence Index, ~11% fewer than GLM-5.3 at 168M, but more than Kimi K3 (133M) and Qwen3.8 2.4T A95B (136M) which score the same on the Intelligence Index. Reasoning tokens account for 134M of the 149M total (~90%). ➤ GLM-5.3-Flash matches GLM-5.3 on real-world agentic work on GDPval-AA v2. With an Elo of 1770, the model is tied within the margin of error for GLM-5.3 and Grok 4.6. This places it behind only Claude Opus 5 (xhigh and max). On Terminal-Bench v2.1 it also matches GLM-5.3 (84.3% vs 83.9%), and on τ³-Banking it trails by 3.1 p.p. at 47.2%. ➤ GLM-5.3-Flash demonstrates good real-world knowledge and hallucination rate, scoring +7 on AA-Omniscience. Its AA-Omniscience Accuracy is 28%, 6 p.p. below GLM-5.3 (max) at 34% and well below GPT-5.6 Terra at 47%. However, with a Hallucination Rate of 28%, it is an improvement over GLM-5.3 at 30%. In real-world knowledge, GLM-5.3-Flash knows less than the bigger models and frontier proprietary models in its Intelligence Index tier with an accuracy of 28%. Additional model details: ➤ Pricing: On Z AI's first-party API, $0.15 / 1M input tokens and $0.50 / 1M output tokens . Cached input tokens are priced at $0.03/ 1M tokens, an 80% discount. ➤ Accessibility: Accessible through Z AI's first-party API at launch. ➤ Size: 320B total parameters with 18B active parameters ➤ License: MIT ➤ Context Window: 400k

💬 4 10 81👁 3781

Z.ai

@Zai_org

1일 전

Introducing GLM-5.3-Flash - Leading capabilities at a highly competitive price - Natively multimodal with a 1M-token context window - A 320B-A18B model released under the MIT License - Previously previewed as Ox Alpha, running entirely on Chinese AI chips Blog: http:// z.ai/blog/glm-5.3-f lash … Available now across all official platforms: Weights: http:// huggingface.co/zai-org/GLM-5. 3-Flash … API: http:// docs.z.ai/guides/llm/glm -5.3-flash … Coding Plan: http:// z.ai/subscribe ZCode: http:// zcode.z.ai/en Chat: http:// chat.z.ai AutoClaw: http:// autoclaw.z.ai

💬 13 33 92👁 1385

📈 Qwen3.8-Flash의 로컬 실행과 오픈 웨이트 배포포스트 4

Qwen3.8-Flash가 오픈 웨이트 멀티모달 MoE로 공개되고 QwenCloud API에 연결되면서, 로컬 실행과 저비용 서비스 이용 경로가 함께 부각됐습니다. Unsloth는 125B 모델을 75GB RAM에서 실행하는 GGUF 배포를 제시했습니다.

세부 내용 보기
  • Qwen3.8-Flash는 125B 파라미터와 51B N-gram을 갖춘 멀티모달 MoE로 소개됐고, Qwen4 아키텍처의 초기 미리보기라는 설명이 붙었습니다. QwenCloud API 가격은 입력 $0.16/1M tokens, 출력 $0.47/1M tokens로 안내됐으며, 오픈 웨이트 공개가 API와 직접 실행을 동시에 가능하게 하는 배포 경로가 됐습니다.
  • Unsloth GGUF를 사용하면 125B 모델을 75GB RAM에서 실행할 수 있고, Qwen3.8-Flash-Next는 CPU RAM 또는 unified memory 환경에서 VRAM에 가까운 속도를 목표로 합니다. Alibaba_Qwen은 이 로컬 실행 경로를 공유했고, Unsloth는 Claude-Opus-4.6 (Max)보다 높은 성능을 낸다는 원문 비교를 함께 게시했습니다.
  • 이 흐름은 대형 모델의 사용 조건을 전용 GPU 보유 여부에서 시스템 RAM과 경량화된 파일 배포 여부로 넓혔습니다. 다만 제공된 포스트에는 로컬 실행 속도나 비교 평가의 구체적인 점수가 제시되지 않았으므로, 성능 우위의 범위는 추가 자료 없이 확정하기 어렵습니다.
원문 트윗 2개 보기

📈 Navigator n2의 GUI·CLI 통합 컴퓨터 자동화포스트 4

Navigator n2는 브라우저 전용 자동화에서 전체 데스크톱 환경으로 범위를 넓히고, GUI·CLI·도구·코드를 작업 단계마다 선택하는 컴퓨터 사용 모델로 공개됐습니다. OSWorld 2.0에서 65.2%, 작업당 $1.46이라는 비용과 성능 수치가 제시됐습니다.

세부 내용 보기
  • 기존 브라우저 사용 모델이 스크린샷을 입력받아 브라우저 동작을 출력하는 데 집중했다면, Navigator n2는 브라우저를 포함한 전체 데스크톱에서 그래픽 인터페이스와 명령줄을 오가며 짧은 코드도 작성합니다. 장시간 작업에서 어느 인터페이스로 전환할지 선택하고 수백 단계 동안 진행 상태를 유지하는 것이 핵심 처리 과제로 제시됐습니다.
  • 학습 과정 자체를 컴퓨터 사용 작업으로 구성해 에이전트가 애플리케이션을 탐색하고, 과제를 만들고, 검증기를 시험하고, 보상 해킹과 예외 사례를 찾도록 했습니다. 여기서 발견한 실패가 다음 학습 데이터로 들어가는 순환 구조를 사용해 더 나은 모델이 더 나은 데이터를 만들도록 설계됐습니다.
  • Navigator n2는 27B 파라미터 모델로 소개됐고 OSWorld 2.0에서 65.2%, 작업당 $1.46을 기록했습니다. Yutori는 이 수치를 브라우저 자동화를 넘어 컴퓨터 사용·지식 작업 에이전트를 비용 효율적으로 배포하려는 근거로 제시했으며, Dayton을 통해 격리된 가상 머신을 브라우저에서 실시간으로 조작하는 playground도 공개됐습니다.
찬성다수

GUI와 CLI를 상황에 맞게 교차 선택하고 실패 데이터를 다음 학습에 재투입하는 구조가 장시간 컴퓨터 작업의 정확도와 비용을 함께 개선한다는 입장입니다.

원문 트윗 2개 보기

Rui Wang

@theruiwang

1일 전

Excited to share that Navigator n2 is officially out! The core idea behind n2 is simple: models should use computers the computer way, not the human way. There’s no single best interface. GUI, CLI, tools, code — pick the right one at every step. The hard part is making those choices well over long horizons: knowing when to switch, when not to, and how to keep making progress over hundreds of steps. The training loop ended up following the same idea: training computer-use models is itself a computer-use task. We use computer-use agents to explore applications, create tasks, test verifiers, find reward hacks, edge cases, and where the latest model fails. Those failures shape the next round of training data. It creates a pretty natural recursive loop: better models make better data, which makes better models. The result is a frontier-level model at a fraction of the cost— 65.2% on OSWorld 2.0 at $1.46 per task. Very proud of the team and where n2 landed!

Devi Parikh

Today we’re introducing Navigator n2. It’s a frontier computer-use model, with just 27B parameters.

💬 0 5 13👁 240

Rui Wang

@theruiwang

1일 전

Excited to share that Navigator n2 is officially out! The core idea behind n2 is simple: models should use computers the computer way, not the human way. There’s no single best interface. GUI, CLI, tools, code — pick the right one at every step. The hard part is making those choices well over long horizons: knowing when to switch, when not to, and how to keep making progress over hundreds of steps. The training loop ended up following the same idea: training computer-use models is itself a computer-use task. We use computer-use agents to explore applications, create tasks, test verifiers, find reward hacks, edge cases, and where the latest model fails. Those failures shape the next round of training data. It creates a pretty natural recursive loop: better models make better data, which makes better models. The result is a frontier-level model at a fraction of the cost— 65.2% on OSWorld 2.0 at $1.46 per task. Very proud of the team and where n2 landed!

Devi Parikh

Today we’re introducing Navigator n2. It’s a frontier computer-use model, with just 27B parameters.

트윗에 첨부된 이미지
💬 0 3 6👁 62

📈 Gemini 3.5 Transcribe의 문맥 인식 음성 처리포스트 4

Gemini 3.5 Transcribe가 85개 이상 언어, 다중 화자 식별, 사용자 정의 어휘와 실시간 스트리밍을 지원하는 음성-텍스트 모델로 공개됐습니다. filler word 제거와 화면 문맥 결합을 통해 정리된 문서와 음성 명령까지 처리하는 기능이 강조됐습니다.

세부 내용 보기
  • 기존 음성 입력에서 배경 소음과 반복적인 오탈자 수정이 문제였다면, Gemini 3.5 Transcribe는 음성을 텍스트로 바꾸는 과정에서 filler word를 제거하고 비정형 발화를 서식 있는 문장으로 정리합니다. 화면 문맥과 로컬 파일을 함께 사용해 음성 입력을 이메일 초안으로 변환하는 흐름도 공개됐습니다.
  • 모델은 85개 이상 언어를 자동 감지하고 여러 화자를 구분하며, 전문 용어에 맞춘 custom vocabulary adaptation을 지원합니다. function calling과 realtime streaming도 제공돼 단순 받아쓰기를 넘어 음성에서 의도와 명령을 추출하는 입력 계층으로 확장됐습니다.
  • Google AI Studio와 Gemini에서 API를 사용할 수 있다고 안내됐습니다. 제공된 포스트에는 WER의 구체적인 수치가 없고, 대신 더 정밀한 전사와 낮은 WER을 제품 기능으로 제시했으므로 실제 언어별 성능 비교는 확인된 범위에 포함되지 않습니다.
원문 트윗 2개 보기

Claude 사용 데이터의 외부 연구 개방포스트 2

Anthropic이 개인정보를 보호한 실제 Claude 사용 데이터를 외부 연구자와 조직이 AI 영향 연구에 활용할 수 있도록 도구를 개방했습니다. 플랫폼 telemetry를 외부화하면서 사용자 프라이버시를 유지하는 플랫폼 투명성이 핵심 방향으로 제시됐습니다.

세부 내용 보기
  • AI 시스템의 복잡한 사회기술적 속성과 실제 사용자 상호작용을 모델 개발사 내부만으로 측정하기 어렵다는 문제의식에서 출발했습니다. Anthropic은 외부 연구자가 회사가 독점적으로 보유한 데이터에 접근해 독립 연구를 수행하도록 하되, Claude 사용 데이터는 privacy-preserved 방식으로 제공하는 경로를 만들었습니다.
  • 이번 프로그램은 AI 플랫폼의 속성뿐 아니라 사람들이 플랫폼과 상호작용하는 방식까지 연구 대상으로 확장합니다. Anthropic Institute의 시스템 측정 작업과 함께 추진되며, 외부 연구자에게 다음 프로젝트의 연구 아이디어를 모집하는 단계로 안내됐습니다.
  • 추가 연구로 HIP Lab은 Claude 사용 경험과 사람들의 감정 사이의 관계를, METR은 coding agent의 실제 생산성 향상을 조사하고 있습니다. 두 연구는 진행 중이며, 제공된 포스트에서는 결과 수치가 아직 공개되지 않았습니다.
찬성다수

실제 사용 데이터를 외부 연구자에게 개방하면 AI 플랫폼의 영향과 특성을 개발사 밖에서도 측정할 수 있어 독립적인 연구 생태계를 넓힐 수 있다는 입장입니다.

원문 트윗 2개 보기

Anthropic

@AnthropicAI

1일 전

For the first time, we’ve given external researchers a way to study AI’s impacts using real, privacy-preserved Claude usage data. To date, this work has only been possible within AI labs. We can’t tell the whole story alone, so we opened up our tools.

💬 19 12 106👁 15523

Jack Clark

@jackclarkSF

1일 전

Ever since co-founding Anthropic I've been obsessed with the need for third-party measurement of AI systems. (It's in my tweet announcing Anthropic: https:// x.com/jackclarkSF/st atus/1398304973205630991 …). Now, we're piloting a way for outside researchers and organizations to study what's happening on AI platforms like Claude while protecting user privacy. This matters because AI systems are giant, complex sociotechnical things with innumerable properties. To think AI labs are going to be able to figure out all the appropriate ways to measure and assess these systems is hubristic and just obviously wrong. So we need to figure out ways to externalize not only the properties of the systems, but also data about how these systems are interacting with the people in the world, which is platform telemetry. What we're trying to do here is prototype the next form of transparency that may be necessary, which is platform transparency while preserving user privacy. Along with this, I want to find ways to empower researchers outside Anthropic to perform their own studies on the kinds of data that companies uniquely have access to, as I think this is a good way to create a more vibrant external research ecosystem. This is another step the larger project for The Anthropic Institute about measuring and externalizing more details about the properties of our AI systems and the platforms we deploy them on, and sits alongside existing efforts from our Economics team via the Anthropic Economic Index. We are now soliciting research ideas from others for our next wave of projects - please apply!

Anthropic

For the first time, we’ve given external researchers a way to study AI’s impacts using real, privacy-preserved Claude usage data. To date, this work has only been possible within AI labs. We can’t tell the whole story alone, so we opened up our tools. https:// anthropic.com/research/enabl ing-independent-research …

💬 18 20 148👁 7786

Perplexity Computer의 구조화된 장기 기억포스트 2

Perplexity의 Brain이 세션, 파일, 출처를 구조화된 knowledge wiki로 통합해 Computer의 기억 시스템으로 활용됩니다. 새 평가에서 정확성 9.3점, currentness 8.0점, recall 8.9점이 개선됐고 토큰 사용량은 15% 감소했습니다.

세부 내용 보기
  • 세션과 연결된 파일·출처가 흩어진 상태에서는 컴퓨터 에이전트가 이전 맥락을 반복적으로 찾아야 하는 문제가 생깁니다. Brain은 이 자료를 구조화된 knowledge wiki로 컴파일하고, Dream agent는 파일과 연결 앱의 문맥을 지속적으로 수집해 다중 홉 문맥 그래프를 축적합니다.
  • 축적된 문맥은 이후 작업에서 관련 정보와 출처를 다시 구성하는 입력으로 사용됩니다. 제공된 새 평가에서는 correctness가 9.3점, currentness가 8.0점, recall이 8.9점 향상됐고, 같은 결과를 15% 적은 토큰으로 처리했습니다.
  • 이 구조는 에이전트의 기억을 단순 대화 기록이 아니라 세션·문서·출처 사이의 연결망으로 다루는 방식입니다. 다만 포스트에는 각 점수의 기준값이나 평가 데이터셋이 제시되지 않아 절대적인 정확도 수준까지 판단할 수는 없습니다.
원문 트윗 2개 보기

📈 Late Interaction 검색의 저장 공간 논쟁포스트 1

Late Interaction 기반 임베딩이 단일 벡터 방식보다 높은 검색 품질을 내면서도 저장 공간 부담이 반드시 커지는 것은 아니라는 실험이 공유됐습니다. 307M 파라미터 mLateOn 모델과 PLAID 인덱스가 Qwen3-Embedding-8B의 단일 벡터 표현보다 높은 nDCG를 기록했다는 주장입니다.

세부 내용 보기
  • 단일 벡터 임베딩은 문서 전체를 하나의 벡터로 압축하지만, Late Interaction은 여러 토큰 표현을 보존해 질의 토큰과 문서 토큰을 더 세밀하게 비교합니다. 이번 실험에서는 307M-parameter mLateOn 모델이 zero-shot 상태에서도 단일 벡터 방식보다 두 자릿수 nDCG 퍼센트포인트 높은 결과를 냈다고 보고됐습니다.
  • PLAID는 2022년에 공개된 기술로, 다중 벡터 표현을 색인하고 검색 단계에서 늦은 상호작용을 수행합니다. lightly-finetuned mLateOn-medical은 수억 개 토큰을 약 1 GiB 인덱스로 표현했으며, 게시자는 이 크기가 Qwen3의 single-vector fp16 표현보다 작다고 비교했습니다.
  • 이 결과가 재현되면 다중 벡터 검색의 품질 향상과 저장 공간 증가를 항상 맞바꿔야 한다는 전제가 약해집니다. 다만 포스트에는 전체 실험 설정과 원시 nDCG 수치가 없으므로, OOD 데이터와 다른 압축 조건에서 같은 우위가 유지되는지는 추가 검증이 필요합니다.
찬성소수

Late Interaction이 토큰 단위 표현을 활용해 검색 품질을 높이면서도 PLAID와 소형 모델을 사용하면 단일 벡터보다 작은 인덱스를 만들 수 있다는 입장입니다.

원문 트윗 1개 보기

용어 해설

혼합 전문가 모델(Mixture-of-Experts)
여러 전문가 네트워크 중 입력에 필요한 일부만 활성화하는 구조입니다. 전체 파라미터 수와 실제 계산에 참여하는 활성 파라미터 수를 분리해, 큰 모델의 지식 용량을 유지하면서 추론 비용을 낮추는 데 활용됩니다.
활성 파라미터(Active Parameters)
모델이 한 번의 입력을 처리할 때 실제 계산에 사용되는 파라미터 수입니다. Mixture-of-Experts 모델에서는 전체 파라미터보다 활성 파라미터가 작을 수 있어, 모델 규모와 실행 비용을 별도로 판단하는 기준이 됩니다.
오픈 웨이트(Open Weight)
모델의 학습 가중치를 외부에 공개하는 배포 방식입니다. 사용자는 공개된 가중치를 내려받아 직접 실행하거나 수정할 수 있으며, GLM-5.3-Flash와 Qwen3.8-Flash 관련 포스트에서 접근성과 로컬 실행의 기반으로 언급됐습니다.
컴퓨터 사용 모델(Computer-Use Model)
스크린샷이나 데스크톱 상태를 입력으로 받아 GUI, CLI, 도구, 코드 같은 인터페이스를 선택하고 실제 컴퓨터 작업을 수행하는 모델입니다. Navigator n2는 브라우저를 넘어 전체 데스크톱 환경을 대상으로 학습됐습니다.
플랫폼 투명성(Platform Transparency)
AI 플랫폼이 실제 사용자와 상호작용하는 방식과 그 영향을 외부 연구자가 조사할 수 있도록 관련 데이터와 연구 도구를 개방하는 접근입니다. Anthropic은 개인정보를 보호한 Claude 사용 데이터를 연구에 제공하는 방식으로 이를 시도했습니다.
AI 분석 전체 내용 보기

AI 요약 · 북마크 · 개인 피드 설정 — 무료

출처 · 인용 안내

원문 발행 2026. 08. 27.수집 2026. 08. 27.출처 타입 TWITTER

인용 시 "요약 출처: AI Trends (aitrends.kr)"를 표기하고, 사실 확인은 원문 보기 기준으로 진행해 주세요. 자세한 기준은 운영 정책을 참고해 주세요.