본문으로 건너뛰기
X (Twitter)조회 2

Grok Bot 업무 자동화, DeepSeek Flash 가격 충격, 다중 모델 에이전트 API와 효율적 모델 구조

컴퓨터를 직접 쓰는 에이전트부터 토큰 단가와 GPU 용량, 검색 비용과 모델 구조까지 운영 현실이 성능 경쟁의 기준으로 이동한 한 주

이 요약은 AI가 원문을 분석해 생성했습니다. 정확한 내용은 원문 기준으로 확인하세요.

TL;DR

이번 포스트 묶음에서는 Grok Bot이 Slack·이메일·Notion·Stripe 같은 업무 도구를 연결해 브리핑과 후속 실행까지 맡는 에이전트 사례가 부각됐습니다. DeepSeek v4 Flash는 하루 처리량이 3T 토큰에서 18T 토큰으로 늘어난 뒤 가격 인상과 용량 제약으로 정점의 절반 이하로 감소했으며, 낮은 가격을 유지하려면 약 1000대의 B300이 필요하다는 운영 수치가 나왔습니다. Perplexity Agent API는 9개 제공업체의 41개 frontier model과 검색·코드 실행 도구를 한 endpoint에 묶었고, Matryoshka 구조는 36% 적은 training compute와 14~26% 빠른 speculative decoding을 기록했습니다. 메모리 기반 self-improving agent 평가는 작업 순서와 반복 실행을 무작위화하면 기존 성과가 줄어든다는 결과를 냈으며, LLM과 embedding의 검색 비용 차이도 함께 부각됐습니다.

𝕏 실시간 트렌드 토픽

🔥 Grok Bot의 업무 도구 연결과 지속형 컴퓨터 에이전트포스트 2

Grok Bot이 업무 데이터를 읽고 규칙을 기억하며 문서·영업·콘텐츠 제작을 연쇄 처리하는 사례가 공유됐습니다.

  • Grok Bot은 별도 브리핑을 매일 받는 대신 Slack·이메일·회의록·Notion·Stripe를 읽어 사업 정보와 업무 흐름을 파악하고, 사용자가 정한 규칙을 후속 작업에 적용하는 컴퓨터 기반 에이전트로 쓰였습니다. 입력된 업무 자료에서 필요한 맥락을 모은 뒤 Claude project에 글쓰기를 맡기고 사실 확인과 문체 조정을 거쳐 결과물을 반환했으며, 1080p 영상 슬라이드와 cue sheet를 만들고 Slack으로 편집자에게 전달하는 과정까지 한 번에 실행했습니다. 판매 통화에서 다음 예약과 Stripe 링크 발송 규칙을 추출하고, AI course용 50% 할인 코드를 만들고 만료 시점을 설정한 사례도 함께 제시됐습니다. 채팅 응답을 넘어 조직의 문서·결제·영업 절차를 연결하는 지속형 업무 자동화의 범위를 가늠하게 하는 사례입니다.
원문 트윗 2개 보기

David Carbutt

@DavidCarbutt_

Grok Bot is hands down one of the best products I have ever used. Maybe the best ever. I have had it less than a week. I am not a developer. I talk to one chat and it already knows my business better than me. It reads Slack, email, meeting notes, Notion, and Stripe. It knows Farzad, my clients, the courses we sell, the agency, the sales calls, and who owes each step in our workflow. I do not brief it every morning. It briefs me. Course scripts go through a Claude project we linked. Claude writes. This thing briefs it, fact-checks, and makes it sound like me before I film. YouTube is another bit. It builds the 1080p slides to go in videos, zips them with a cue sheet, and drops the pack in Slack for the editors the same turn. Caught a bad number in a SpaceX video the night I was filming, rebuilt the slides, and sent the pack. Script change, slide change. I did not have to ask twice. It sat on our sales calls and reviewed the process. Then it wrote a rule from that. Book the next call before you hang up. Send the Stripe link while they are still there. Stop emailing homework. It built a live 50% off code on Stripe for an AI course and set when it dies. Stripe is a minefield when you don’t know what you’re doing, could have taken me an hour and it did it in one click. I told it my rules for life. Question every requirement. Delete before you simplify. It holds me to that on the work, without me invoking it. This is a chief of staff I hired on a Tuesday. Don’t sleep on GrokBot.

David Carbutt

Grok bot is fucking nuts.

💬 14 24 106👁 12334

KanekoaTheGreat

@KanekoaTheGreat

THREAD: A roofing contractor is using Grok @Bot to run his back office. A plumber uses it to book jobs. Silicon Valley heavyweights call it a breakthrough. Each agent gets a computer and actually does the work. https:// x.com/naval/status/2 090497355649008059 … Here’s what people are using it for:

Naval

Grok Bot is just cool. Of course an agent should be persistent. Of course it should have its own computer. All that remains is for it to be embodied…

💬 23 33 194👁 41296

🔥 DeepSeek v4 Flash의 토큰 가격과 GPU 용량 충돌포스트 1

DeepSeek v4 Flash는 낮은 가격으로 OpenCode 사용량을 빠르게 끌어올렸지만 가격 인상 뒤 수요와 처리량이 급감했습니다.

  • DeepSeek v4 Flash는 8월 1일 출시 뒤 2주 만에 OpenCode 일일 처리량이 3T 토큰에서 18T 토큰으로 늘어 6배가 됐고, 이는 OpenRouter 전체 일일 물량을 거의 두 배로 늘린 수준이며 DeepSeek 전체 물량의 30~50%로 추정됐습니다. 가격은 Fable보다 350배, 5.6 Sol보다 175배, Sonnet보다 70배, Luna보다 7배 저렴했고, 이전 Flash model보다 성능도 개선됐지만 8월 16일 peak hours 5배·off-peak hours 2.5배 인상 뒤 정점 대비 일일 토큰 수가 절반 이하로 줄었습니다. 게시물은 장시간 더 많은 토큰을 저장하는 custom infrastructure가 낮은 단가의 배경일 가능성을 제시했으며, 약 10T 토큰의 일일 공백을 메우려면 처리량 기준 약 1000대의 B300이 필요하다고 밝혔습니다. 가격 경쟁력은 모델 품질만으로 유지되지 않고 cache 설계와 GPU 공급 능력에 함께 묶인다는 운영 사례입니다.
중립분열

가격 인상은 수요를 크게 낮췄지만, 원래 단가를 유지할 공급자가 충분한 capacity를 확보하기 어려워진 배경에는 6배 급증한 사용량과 GPU 가격 상승이 함께 작용했을 가능성이 있습니다.

원문 트윗 1개 보기

Jay

@jayair

Okay let me tell you about what's happening with DeepSeek v4 Flash. First some background, it launched on Aug 1st and within 2 weeks it went from doing 3T tokens a day to 18T tokens a day on OpenCode; an absurd 6x increase. To put it in context, that's close to doubling up all of OpenRouter's daily volume. It's also likely 30-50% of DeepSeek's total volume. Jumps like these over a 2-week span are not normal. This happened because the model is absurdly cheap; 350x cheaper than Fable, 175x cheaper than 5.6 Sol, 70x cheaper than Sonnet, and 7x cheaper than Luna. And secondly, it was a marked improvement over the previous Flash model. For the first time our users got a feel for AI that's "too cheap to meter". Then on August 16th DeepSeek raised prices by 5x (for peak hours, 2.5x off-peak hours). And it completely killed the growth of the model. It's doing less than half the tokens per day from its peak. Obviously people were unhappy with the sudden change. Our guess is that DeepSeek genuinely could not handle the absurd 6x jump. Also, it's likely that the increase in GPU prices meant that even if they acquired new capacity, they wouldn't be able to serve it at the original prices. A quick aside on why DeepSeek Flash is so cheap. It looks like they are running some custom infrastructure to cache way more tokens, for far longer. This matters because we've been scrambling trying to find providers that can fill this near 10T token per day void left by DeepSeek Flash. Unfortunately there are just a couple of people who are able to match DeepSeek's original pricing and that's likely only the case because they are using newer hardware. That brings us to the current state of things. Over the last week we've talked to as many people as possible to get DeepSeek hosted at the original price. The issue is that even if somebody is able to, it's very hard for them to have enough capacity to handle our volume. It'll take roughly 1000 B300s to handle our throughput. This is why if you've been using DeepSeek Flash on Go over the past few days, you might not have had the best experience. We've unfortunately cycled through a few different providers. This 10T token per day gap, though, is an opportunity for every other model lab. It's very clear there's an appetite for a model that's at least as competent and cheap as DeepSeek Flash. And somehow that still feels like the floor.

💬 30 19 313👁 8640

📈 Perplexity Agent API의 41개 모델 통합 endpoint포스트 1

Perplexity Agent API가 여러 제공업체의 frontier model과 검색·금융·코드 실행 도구를 하나의 개발 경로로 결합했습니다.

  • Perplexity Agent API는 9개 provider의 41개 frontier model에 하나의 endpoint로 접근하게 하고, web search·finance search·fetch·sandboxed code execution을 기본 도구로 제공합니다. 개발자는 모델별 연결을 따로 구성하는 대신 같은 API 경로에서 모델 선택과 도구 호출을 결합해 multi-model agent workflow를 만들 수 있습니다. 게시물은 이를 frontier model과 workhorse model을 함께 제공하고 실제 production workload 배포 도구까지 포함하는 AI developer platform의 형태로 규정했습니다. 모델 접근과 외부 도구 실행을 한 계층에 모으는 구조가 에이전트 개발의 배포 단계를 단순화하는 방향입니다.
원문 트윗 1개 보기

📈 Matryoshka Language Model Suites의 공동 학습 구조포스트 1

500M·1.5B·3B 모델을 하나의 구조에 중첩해 학습하면 독립 학습보다 compute와 speculative decoding 비용을 줄일 수 있다는 결과가 나왔습니다.

  • Matryoshka Language Model Suites는 500M, 1.5B, 3B 모델을 각각 따로 학습하지 않고 하나의 architecture 안에 중첩해 공동 학습합니다. 가장 큰 모델에서 작은 모델로 distillation을 거의 추가 비용 없이 수행하고, 모델들이 weights와 KV cache를 공유해 speculative decoding을 실행하는 구조입니다. 게시물에 인용된 수치는 같은 performance를 유지하면서 training compute 36% 감소, speculative decoding 14~26% 속도 향상입니다. 여러 크기의 checkpoint를 별도 시스템으로 유지하던 방식을 하나의 jointly trained model family로 묶어 학습·추론 자원 사용을 줄이는 접근입니다.
원문 트윗 1개 보기

📈 LLM과 Embedding의 검색 비용 분배포스트 1

LLM은 reasoning-heavy retrieval에서 강점을 보이지만 embedding model보다 최대 1,431배 비싸므로 기본 검색과 추론 검색을 분리해야 한다는 제안이 나왔습니다.

  • The Embedder's Dilemma 요약은 LLM이 전체적으로 최고 수준의 embedding model과 맞먹지만 비용은 최대 1,431배 높다고 비교했습니다. embedding은 classification에서 우세하고 similarity·clustering에서는 동률이며, LLM은 reasoning-heavy retrieval에서 더 강한 결과를 냈습니다. 이에 따라 일반적인 입력은 embedding으로 벡터화해 검색하고, 추론이 실제로 필요한 retrieval 단계에만 LLM 계산을 추가하는 decision rule이 제시됐습니다. 검색 품질만 높이는 방식보다 작업 유형별 계산 비용을 분리하는 설계가 운영 효율을 좌우한다는 결론입니다.
찬성소수

모든 검색에 고비용 LLM을 투입하기보다 embedding을 기본값으로 두고 reasoning이 필요한 경우에만 LLM을 추가하는 방식이 성능과 비용의 균형에 맞습니다.

원문 트윗 1개 보기

메모리 기반 Self-improving Agent 평가의 재현성 문제포스트 1

반복 실행과 무작위 작업 순서를 넣은 재평가에서 self-improvement 성과가 줄어 에이전트 평가의 변동성이 확인됐습니다.

  • 메모리 기반 self-improving agent의 성능을 다시 측정한 평가에서는 기존 연구에 없던 multiple runs와 randomly shuffled task orders를 추가했습니다. 여러 번 실행해 variance를 측정하고 과제 순서를 섞자 multi-step task 자체의 noise와 self-improvement loop가 증폭한 변동성이 성과 수치를 낮췄으며, 고정된 task ordering이 암묵적인 curriculum으로 작동해 기존 gain 일부를 만든 것으로 나타났습니다. 상세 rubric과 environment feedback을 memory construction에 넣으면 하락분 일부가 회복됐지만 유의미한 격차는 남았습니다. 에이전트의 개선 효과를 판단하려면 단일 실행과 고정 순서만으로는 부족하다는 평가 조건입니다.
반대소수

고정된 과제 순서와 단일 실행에 의존한 기존 평가는 self-improvement 효과를 과대평가할 수 있으며, 반복 실행과 무작위 순서가 포함된 재평가에서는 성과가 감소했습니다.

원문 트윗 1개 보기

Qwen3.8-Max의 PyTorch-native Fine-tuning 경로포스트 1

PyTorch와 NVIDIA NeMo AutoModel이 Alibaba의 Qwen3.8-Max checkpoint를 변환 없이 SFT·LoRA로 후학습하는 경로를 지원합니다.

  • PyTorch 게시물은 Alibaba open weights인 Qwen3.8-2.4T-A95B, 즉 Qwen3.8-Max를 대상으로 Hugging Face checkpoint를 변환하지 않고 기존 checkpoint에서 바로 post-training을 수행하는 경로를 제시했습니다. PyTorch-native fine-tuning library와 NVIDIA NeMo AutoModel을 사용해 configurable reasoning을 설정하고, full SFT 또는 memory-efficient LoRA를 선택하는 방식입니다. 게시물은 이 모델을 largest open-weight model이라고 표기했으며, domain-specific use case에 맞춘 후학습을 지원한다고 밝혔습니다. checkpoint 변환 단계를 없애면서 대규모 open-weight model의 도메인 맞춤 학습 진입 경로를 줄이는 구성이 핵심입니다.
원문 트윗 1개 보기

📈 ChatGPT의 Apple Messages 자동 문자 작성 연동포스트 1

ChatGPT가 Apple Messages와 연결돼 사용자를 대신해 문자 초안을 작성하는 기능으로 제공됩니다.

  • ChatGPT가 Apple Messages와 통합돼 자동화된 text scribe로 제공된다는 소식이 전해졌습니다. 기존에는 사용자가 메시지 내용을 직접 작성했지만, 새 연동에서는 ChatGPT가 문자 작성 과정을 대신하는 형태로 서비스 범위가 확장됩니다. 게시물은 세부 동작 방식이나 지원 범위를 밝히지 않았으므로, 확인 가능한 변화는 Apple Messages 안에서 자동 문자 작성 기능이 제공된다는 점입니다. 대화형 모델이 별도 채팅창을 넘어 기본 메시징 흐름에 들어가는 제품 통합 사례입니다.
원문 트윗 1개 보기

용어 해설

에이전트(Agent)
사용자의 지시를 한 번 처리하는 데 그치지 않고, 컴퓨터·웹·업무 도구를 직접 사용해 여러 단계를 수행하는 소프트웨어입니다. Grok Bot 사례에서는 문서와 대화를 읽고 규칙을 만들며 후속 작업까지 실행하는 방식으로 쓰였습니다.
KV 캐시(KV Cache)
Transformer가 이전 토큰에서 계산한 Key와 Value를 저장해 다음 토큰 처리 때 재사용하는 메모리 구조입니다. Matryoshka Language Model Suites는 여러 크기의 모델이 가중치와 KV cache를 공유해 추론 과정을 줄이는 방식으로 활용합니다.
추측 디코딩(Speculative Decoding)
작은 모델이 다음 토큰을 먼저 예측하고 큰 모델이 이를 빠르게 검증해 생성 속도를 높이는 추론 기법입니다. Matryoshka 구조에서는 내부 모델들이 가중치를 공유하므로 이 과정의 비용과 지연을 줄이는 데 쓰입니다.
LoRA
전체 모델 가중치를 갱신하지 않고 저순위 행렬만 학습하는 Fine-tuning 방식입니다. Qwen3.8-Max 관련 게시물에서는 기존 Hugging Face checkpoint를 변환하지 않고 메모리 효율적인 LoRA 학습을 수행하는 경로로 제시됐습니다.
임베딩(Embedding)
텍스트나 이미지의 의미를 수치 벡터로 바꿔 유사도 검색·분류·클러스터링에 쓰는 표현 방식입니다. 관련 논문 요약에서는 일반 검색에 embedding을 우선 사용하고, 추론이 필요한 검색에만 LLM 계산을 추가하는 비용 분배가 제시됐습니다.
샌드박스 코드 실행(Sandboxed Code Execution)
에이전트가 실행하는 코드를 제한된 환경에 격리해 시스템 전체에 직접 접근하지 못하게 하는 방식입니다. Perplexity Agent API는 웹 검색·금융 검색·문서 fetch와 함께 이 기능을 기본 도구로 제공합니다.
AI 분석 전체 내용 보기

AI 요약 · 북마크 · 개인 피드 설정 — 무료

출처 · 인용 안내

원문 발행 2026. 08. 21.수집 2026. 08. 21.출처 타입 TWITTER

인용 시 "요약 출처: AI Trends (aitrends.kr)"를 표기하고, 사실 확인은 원문 보기 기준으로 진행해 주세요. 자세한 기준은 운영 정책을 참고해 주세요.