DEV Community

Mingxin Technology profile picture

Mingxin Technology

Storage acceleration for LLM inference. KV-cache tiering on NVMe-oF: +29-40% throughput (signed benchmarks, reproducible). mingxinstorage.xyz

Joined Joined on 
Three Clause Types Often Missed in Compute Rental Contracts

Three Clause Types Often Missed in Compute Rental Contracts

Comments
5 min read
Energy Consumption Assessment and Optimization Paths for Data Center-Scale KV Cache Deployment

Energy Consumption Assessment and Optimization Paths for Data Center-Scale KV Cache Deployment

Comments
7 min read
Choosing Compute Rental: Focus on Three SLA Metrics, Not Unit Price

Choosing Compute Rental: Focus on Three SLA Metrics, Not Unit Price

Comments
5 min read
Bandwidth and Latency Optimization Strategies for KV Cache in Edge Computing

Bandwidth and Latency Optimization Strategies for KV Cache in Edge Computing

Comments
5 min read
Four Key Dimensions for Evaluating the Reliability of Domestic AI Storage

Four Key Dimensions for Evaluating the Reliability of Domestic AI Storage

Comments
6 min read
Three-Layer Adaptation of the Ascend Inference Stack: Driver, Operator, and Framework Are All Indispensable

Three-Layer Adaptation of the Ascend Inference Stack: Driver, Operator, and Framework Are All Indispensable

Comments
6 min read
Can Compressed Sensing Be Used for Inference Storage Data Compression?

Can Compressed Sensing Be Used for Inference Storage Data Compression?

Comments
4 min read
NVMe-oF vs. RDMA: Performance Comparison in Inference Storage

NVMe-oF vs. RDMA: Performance Comparison in Inference Storage

Comments
5 min read
Storage Selection Strategy for Inference in Domestic AI Computing Centers

Storage Selection Strategy for Inference in Domestic AI Computing Centers

Comments
5 min read
How KV Cache Prefetch Cuts Storage Latency

How KV Cache Prefetch Cuts Storage Latency

Comments
6 min read
KV Cache Reuse in Multi-Turn Dialogue: 29% Throughput Gain Measured

KV Cache Reuse in Multi-Turn Dialogue: 29% Throughput Gain Measured

Comments
5 min read
Evaluating Long-Context Inference Performance of Domestic Accelerator Cards

Evaluating Long-Context Inference Performance of Domestic Accelerator Cards

Comments
4 min read
Storage Bottlenecks in Real-Time Video Inference and a Tiered Acceleration Approach

Storage Bottlenecks in Real-Time Video Inference and a Tiered Acceleration Approach

Comments
5 min read
KV Cache Reuse in Multi-Turn Dialogue: A Deployment Case Study

KV Cache Reuse in Multi-Turn Dialogue: A Deployment Case Study

Comments
5 min read
How KV Cache Pooling and Sharing Improves Inference Resource Utilization

How KV Cache Pooling and Sharing Improves Inference Resource Utilization

Comments
5 min read
Distributed KV Cache Load Balancing: Strategies and Measured Trade-offs

Distributed KV Cache Load Balancing: Strategies and Measured Trade-offs

Comments
6 min read
Measured Evaluation of KV Cache Acceleration for Inference in Online Education Scenarios

Measured Evaluation of KV Cache Acceleration for Inference in Online Education Scenarios

Comments
4 min read
Deployment Path and Measured Performance of Domestic AI Inference Acceleration Cards in Real-Time Database Queries

Deployment Path and Measured Performance of Domestic AI Inference Acceleration Cards in Real-Time Database Queries

Comments
5 min read
Case Analysis of Huawei OceanStor UCM Inference Acceleration Solution

Case Analysis of Huawei OceanStor UCM Inference Acceleration Solution

Comments
5 min read
Why KV Cache Tiering Is the Key TCO Optimization Point in a Three-Tier Storage Architecture for Compute Centers

Why KV Cache Tiering Is the Key TCO Optimization Point in a Three-Tier Storage Architecture for Compute Centers

Comments
6 min read
The Hidden Cost of Model Switching: A Measured Path from 46.7% to 62.8% Effective Compute Utilization

The Hidden Cost of Model Switching: A Measured Path from 46.7% to 62.8% Effective Compute Utilization

Comments
5 min read
Cross-Platform Methodology Porting: Validating KV Cache Memory Efficiency from Muxi N260 to MI308X

Cross-Platform Methodology Porting: Validating KV Cache Memory Efficiency from Muxi N260 to MI308X

Comments
3 min read
Inference Deployment in Xinchuang Scenarios: Architecture Essentials for Data Not Leaving the Domain

Inference Deployment in Xinchuang Scenarios: Architecture Essentials for Data Not Leaving the Domain

Comments
5 min read
Clos Network Architecture: A Cost-Effectiveness and Selection Framework for Thousand-Card Inference Clusters

Clos Network Architecture: A Cost-Effectiveness and Selection Framework for Thousand-Card Inference Clusters

Comments
7 min read
How to Determine the Storage-to-Compute Ratio for Inference Clusters: Measured Basis for One Array Serving 8 Nodes

How to Determine the Storage-to-Compute Ratio for Inference Clusters: Measured Basis for One Array Serving 8 Nodes

Comments
6 min read
Engineering Practice Essentials of NVMe-oF + RoCEv2 in Inference Storage Scenarios: A Case Study with FX100 Measurements

Engineering Practice Essentials of NVMe-oF + RoCEv2 in Inference Storage Scenarios: A Case Study with FX100 Measurements

Comments
5 min read
What 90% Line-Rate Utilization on a Single 100GbE Port Means: Analyzing Network Bottlenecks in Inference Storage

What 90% Line-Rate Utilization on a Single 100GbE Port Means: Analyzing Network Bottlenecks in Inference Storage

Comments
5 min read
loading...