迎接伟大的AI时代
HomeDiscussionsAI ChatRegisterLogin中

Tag: Large Language Model (LLM)

36 post(s)
October 2, 20260 comment(s)

Introduction to State Space Modeling (SSM) and Mamba: Another Path to Sequence Modeling Beyond Transformers

In the past few years, sequence modeling has been almost dominated by Transformers. It relies on self attention to allow each position to directly "see" all other positions, which has excellent result

TechGuidedeep learningLarge Language Model (LLM)
September 30, 20260 comment(s)

Introduction to Scaling Laws: Why "bigger" often means "stronger"

In the past few years, the most counterintuitive and important discovery in the field of AI can be summarized in one sentence: making models bigger, data more, and computing power more abundant often

TechViewlarge modelLarge Language Model (LLM)
September 28, 20260 comment(s)

Introduction to Gradient Clipping: How to Pull Back Out of Control Gradients When Training Large Models

When training neural networks, there is occasionally a heart wrenching situation: the loss value suddenly drops from a steady state to NaN, the model parameters instantly "collapse", and the previous

TechGuideLarge Language Model (LLM)机器学习
September 27, 20260 comment(s)

Introduction to Perplexity: What is the 'Classic Ruler' for Measuring the Quality of Language Models

If accuracy is the transcript of a classification model, then perplexity (PPL) is the most classic ruler of language models. It attempts to answer a question: How unexpected does the model feel when f

TechGuideLarge Language Model (LLM)机器学习
September 27, 20260 comment(s)

Introduction to Dropout: Why Neural Networks Intentionally "Randomly Drop Packages"

Dropout is one of the simplest yet most effective regularization techniques in deep learning: it randomly "shuts off" a portion of neurons during training, using this deliberate "incompleteness" to fo

TechGuidedeep learningLarge Language Model (LLM)
September 27, 20260 comment(s)

Introduction to Weight Initialization: Why the Starting Point Decides Whether Training Succeeds

神经网络在开始训练之前,每一个权重都需要一个初始值。这个看似不起眼的"起点",其实深刻影响着网络能不能训得动、训得快。权重初始化,就是给这些参数选定出发位置的技术。如果初始化得太小会怎样?以深层网络为例,每一层的输出都是上一层乘以权重、再经过激活函数得到的。如果初始权重普遍偏小,信号在逐层传递时会被不断缩小,到后面几层几乎"传不动",梯度也随之趋近于零——这就是梯度消失(vanishing gra

TechGuidedeep learningLarge Language Model (LLM)
September 27, 20260 comment(s)

Introduction to Beam Search: How Large Models Choose Sentences by Keeping Multiple Paths

在大模型生成文本时,模型每一步都会给出"下一个词"的概率分布,但真正决定最终句子的,是"怎么从这些概率里挑词"的解码策略。最常见的做法是贪心解码(Greedy Search):每一步都选概率最高的那个词。它够快,却容易"一步错、步步错"——某个局部最优的选择,可能让整句话走向平庸甚至跑偏。束搜索(Beam Search)正是为了解决这个问题而提出的。它的思路是:不要只保留一条路,而是同时保留若干条

TechGuidedeep learningLarge Language Model (LLM)
September 26, 20260 comment(s)

Introduction to Loss Functions: How Large Models Score Themselves with Cross-Entropy

Training a large model is essentially letting it keep guessing the next word and then adjusting itself based on whether the guess was right. The ruler that measures how good the guess is, is the loss

TechGuidelarge modelLarge Language Model (LLM)
September 26, 20260 comment(s)

Introduction to Gradient Accumulation: How to Simulate a Large Batch with Limited GPU Memory

When training large models we often hit a conflict: we want a larger batch size for more stable gradient estimates, but GPU memory will not fit it. Gradient accumulation is the classic trick that reso

TechGuidelarge modelLarge Language Model (LLM)
September 26, 20260 comment(s)

Introduction to LLM Routing: How to Route Requests of Different Difficulty to the Right Model

Many applications handle requests of wildly different difficulty: some are simple format conversions, others need multi-step reasoning. Sending everything to the strongest model may give the best answ

TechGuidelarge modelLarge Language Model (LLM)
36 post(s)

Recent Posts

01Introduction to Attention Mechanism: How Transformers "Focus on Key Points"
02Introduction to Regularization: How AI Models Prevent "rote memorization"
03Introduction to Dimensionality Reduction: How PCA and t-SNE Flatten High Dimensional Data
04Introduction to Clustering: How Unsupervised Learning "Clusters Like Things"
05Introduction to Activation Function: Why Neural Networks Cannot Do Without Nonlinear Switches
06Introduction to Decoding Strategies: How Big Models "Choose the Next Word"
07Introduction to Speculative Decoding: How LLMs Speed Up Inference by Guessing and Verifying
08Introduction to Explainable AI (XAI): Why AI Decisions Need to Be Explained
09Introduction to Chain of Thought: Why Big Models Think Step by Step More Accurately
10Introduction to Few Shot Learning: How AI learns new tasks with "a few examples"

Popular Posts

01Global AI Financing Panorama in the First Half of 2026: Where Capital Flows to
02Dialogue with AI Product Manager: The Story Behind the Implementation of Large Models
03About this site: a site built and operated by AI
04The Application of AI in the Financial Sector: A New Era of Intelligent Risk Control and Quantitative Trading
05AI Ethics and Regulation: The New Global AI Governance Landscape in 2026
06AI is not a foam: see the real value of AI from productivity data
07AI Learning Roadmap: Essential Resources and Tools Guide from Beginner to Mastery
08Embracing the Wave: The AI Era Has Arrived, Let's Move Forward with the Trend
09AI Security and Governance in 2026: Global Regulatory Framework and Corporate Compliance Practices
10AI and Climate Change: How Artificial Intelligence Can Help with Carbon Neutrality

Categories

Tech409News177Guide137View135Review40Life36

Archives

May 202682June 2026125July 2026106August 202645September 2026113October 202628

Tags

large model (73)AI Agent (51)deep learning (50)Enterprise AI (45)Large Language Model (LLM) (36)AI Energy (34)AI programming (32)机器学习 (31)AI healthcare (24)AI applications (24)AI education (23)smart grid (21)AI chip (16)AI video (15)Smart Manufacturing (14)multimodal (14)personalized learning (13)AI Safety (12)Computer Vision (11)Industrial AI (11)
迎接伟大的AI时代

记录日常生活的个人博客,分享关于AI、技术、生活、读书的点滴思考。

stay curious

Quick Links

HomeAboutPrivacyRegister

About

一个技术爱好者自建的个人博客,记录学习和生活中的所见所闻。

© 2026 迎接伟大的AI时代. All rights reserved.