I am speed

I am speedFast.com-style speed test for LLM APIs

A fast.com-style speed test for LLM APIs that measures how fast models respond in tokens per second and compares throughput across many providers.

free
LLM benchmarktokens per secondAPI speed testmodel comparisonthroughputOpenAIAnthropicdeveloper tool
Indexed
Updated
3.2 (0 Number of reviews)

Log in to rate the project

Try Now

I am speed (the site brands itself as LLMark) is a benchmarking tool for large language model APIs. Its stated goal is to be a fast.com-style speed test for LLMs: instead of measuring internet bandwidth, it measures how quickly different model endpoints generate output.

What it measures

The main metric is throughput in tokens per second (tok/s). The interface shows live throughput timing while a test runs and reports the fastest tok/s result, along with an engine status readout indicating whether the service is active. This makes it useful for comparing the raw response speed of one model or provider against another.

Providers covered

The tool lists support for a broad set of endpoints so speeds can be compared side by side.

  • Hosted providers: OpenAI, Anthropic, Groq, Cerebras, Fireworks AI, Mistral, OpenRouter, Google Gemini, x.ai, z.ai and Kimi.
  • Local models: self-hosted models can also be benchmarked.

Good to know

The public page focuses on the live speed display and provider list. It does not clearly state pricing, account requirements or whether the project is open source, so those details should be confirmed on the site itself. Because it measures speed rather than answer quality, it is best treated as one signal among several when choosing a model.

Pros & Cons

Pros

  • Simple fast.com-style way to compare LLM response speed
  • Covers many major providers plus local models
  • Reports throughput in tokens per second with a live readout
  • Handy for picking the fastest endpoint for a task

Cons

  • Measures speed only, not answer quality or cost
  • Pricing and open-source status are not clearly stated on the page

Frequently Asked Questions

What does I am speed measure?

It measures how fast an LLM endpoint generates output, reported as tokens per second.

Which providers are supported?

It lists OpenAI, Anthropic, Groq, Cerebras, Fireworks AI, Mistral, OpenRouter, Google Gemini, x.ai, z.ai, Kimi and local models.

Does it test answer quality?

No. It focuses on speed, so it is best used alongside other checks for quality and cost.

Is it free?

The page does not clearly state pricing, so this should be confirmed on the official site.

Explore More

Similar Tools

Nadir

Nadir

Nadir introduces a verifier-gated LLM router designed to cut API costs without sacrificing quality. It routes requests to cheaper models first, then uses a calibrated verifier to score responses. If quality falls short, it escalates to a more powerful model. This approach claims up to 60% cost savings while maintaining 98% quality, offering an OpenAI-compatible, two-line integration for high-volume, varied complexity workloads.

StackBuilder

StackBuilder

StackBuilder is a free, AI-driven tool that generates professional cloud architecture diagrams from natural language descriptions. It supports major platforms like AWS, Azure, GCP, and Kubernetes, and allows exports to PNG, SVG, and PDF. No registration is required, making it ideal for system design interviews, architecture documentation, and presentations.

Nest by RAVEN

Nest by RAVEN

Nest by RAVEN (also known as NestMux) is a multi-AI terminal workbench for developers. It allows parallel execution of Claude, Gemini, Codex, Copilot, and Aider within a single window. Features include Git worktrees integration, team terminal sharing, MCP panel, and broadcast prompts. It's local-first, telemetry-free, and available for free download on macOS, Windows, and Linux.

h5i

h5i

h5i is an open-source sandbox tool designed for AI coding agents like Claude Code and Codex. It encapsulates the agent, shell, dependencies, and browser within a single, isolated boundary. This lightweight sandbox launches in under 200ms, supports exporting auditable patches and execution logs, and is local-first, SaaS-free, delivered as a single Rust binary.

Deep Work Plan

Deep Work Plan

Deep Work Plan is an open-source methodology that transforms any code repository into a structured, AI-executable environment using an `init.md` file. It breaks down long-term coding tasks into atomic steps with clear acceptance criteria, validation gates, and recoverable states, preventing AI agents from derailing. It's agent-agnostic, open-source (MIT), and prevents vendor lock-in.

Spanly

Spanly

Spanly offers a specialized observability and monitoring solution for Model Context Protocol (MCP) servers. It helps SaaS teams track error rates, session traces, latency, and client behavior in production environments. With quick CLI/SDK integration, it complements existing monitoring stacks and provides data residency options in the US and EU. A free scanner is available to quickly identify protocol-level vulnerabilities.

Open-source Alternatives

guidellm: Open-Source Tool for Evaluating and Optimizing LLM Inference

guidellm is an open-source tool developed by the vLLM team to evaluate and optimize Large Language Model (LLM) inference performance in production environments. It offers stress testing, latency analysis, and throughput assessment to help developers identify bottlenecks and fine-tune deployment configurations. The project is primarily written in Python and licensed under Apache-2.0. At the time of collection, it had 1214 stars on GitHub.

ai-gateway: Unified AI Gateway Based on Envoy Gateway

ai-gateway is an open-source project built on Envoy Gateway, offering a unified API gateway to manage access to diverse generative AI services. It simplifies AI application integration and operations by providing features like load balancing, caching, and rate limiting for various AI providers. The project is written in Go and licensed under Apache-2.0.

go-micro: Go framework fusing AI agent harness with microservices

go-micro is an open-source Go framework that fuses an AI agent harness with microservices, supporting MCP, A2A, and multi-LLM integration. It is licensed under Apache-2.0 and primarily written in Go. As of the collection time, the project had 22,755 stars on GitHub.

Kun: Local-First AI Agent Workspace

Kun is a local-first AI agent workspace that unifies coding, writing, design, research, and automation through a shared GUI and TUI runtime. The project is primarily developed in TypeScript and has an 'Other' license. As of collection time, it has 4813 GitHub stars.

terax-ai: Lightweight Tauri-based Desktop Dev Environment

terax-ai is a Tauri-based desktop development environment with a size of only 7-8 MB. It integrates a GPU terminal, CodeMirror editor, Git tools, and multi-provider AI agents, offering an all-in-one development experience. The project is primarily written in TypeScript and licensed under Apache-2.0.

jar-analyzer: Open-Source GUI Tool for Java JAR Analysis with AI Assistant

jar-analyzer is an open-source GUI tool for Java JAR package analysis, featuring an integrated AI assistant. It offers robust capabilities like JAR DIFF, method call graph exploration, DFS call chain analysis, taint analysis, and control flow graph (CFG) program analysis. Ideal for Java developers and security researchers, it streamlines code auditing and reverse engineering tasks. The primary language is Java, licensed under GPL-3.0, with 2111 GitHub stars at the time of collection.