Mailgun Review: Put It Through a Controlled Trial
Summarize with AI
Mailgun reviews well on the things a developer evaluates in an afternoon and less well on the thing that decides whether you keep it: how long your own sending data stays searchable. The platform is clean to set up, the API is mature, and the event model is genuinely good. Then a colleague asks what was sent to a named prospect nine days ago, and on three of the four published plans the platform has already forgotten. Run the trial against that question rather than against the integration, because the integration is the part that was never going to fail.
Set up the domain properly before judging anything
Authentication decides what the rest of the trial measures, so it is worth completing properly even in an evaluation. Verify the sending domain instead of individual addresses, publish the DKIM records Mailgun issues, and confirm the SPF record includes the Mailgun sending infrastructure.
Then check DMARC alignment specifically. A message can pass DKIM and still fail DMARC if the signing domain does not align with the domain in the visible From header, and an evaluation that skips this check will produce delivery numbers that flatter the platform relative to how it will behave in production.
Use a subdomain for the sending program instead of the root domain. That isolation keeps a sending problem from reaching the reputation of the domain your invoices and contracts come from, and it makes a future platform change less disruptive.
Our view: a trial run on an unauthenticated or partially aligned domain measures nothing useful. Budget half a day for authentication before the first test send, and treat any delivery figure gathered before that as discarded.
Test the retention window against your real reply cycle
The Mailgun pricing page publishes log retention at 1 day on Free and Basic, 5 days on Foundation and 30 days on Scale, as checked on September 19, 2026.
This is the review finding that matters most for outbound work. A sequenced program produces replies across days and weeks, and a prospect responding on day nine to a message you cannot retrieve is ordinary, not exceptional. On Basic the log is gone before the second step of most sequences has even sent.
Test it deliberately during the trial. Send a real campaign, wait past the retention window, then try to answer a specific question about a specific recipient using only the platform. Whatever you can produce at that moment is what you will be able to produce during a client dispute.
| Plan | Retention | What it can answer |
|---|---|---|
| Free and Basic | 1 day | Yesterday's send only |
| Foundation | 5 days | The current working week |
| Scale | 30 days | A full reply cycle and a monthly review |
If the honest requirement is longer than 30 days, the answer is not a higher tier. It is an event pipeline into storage you control, which is the arrangement that survives a downgrade or a migration.
Evaluate the event webhooks under load, not at rest
The event model deserves a harder test than confirming that a single test event arrives.
Point the webhooks at an endpoint you control and record three things during a real send: latency between the sending event and the webhook arriving, completeness against the accepted count, and behaviour when your endpoint returns an error. The third is the one nobody tests and the one that matters, because a webhook pipeline that silently drops events during a brief outage will leave gaps in exactly the evidence you built it for.
While the pipeline is up, capture events into your own storage from the first day of the trial. Doing so converts the retention question from a plan decision into an engineering one, and it gives the evaluation a dataset the platform cannot age off mid-review.
Price the dedicated IP and the ramp it obliges
Foundation and Scale each include one dedicated IP, and additional IPs are published at $59 per IP per month. Basic lists dedicated IP access from 50,000 messages a month.
An included IP is a real benefit carrying an unpriced obligation. A new IP has no sending history, and receiving systems treat that absence as a risk signal. Amazon's dedicated IP warm-up documentation describes establishing a positive reputation as taking around two weeks with some email providers and up to six weeks with others, a range that reflects how receivers build trust and therefore applies wherever a fresh IP is provisioned.
Plan the trial around that. A dedicated IP provisioned at the start of a two week evaluation will produce delivery numbers from an IP in warm-up, which understates the platform. Either run the evaluation on the shared pool and provision the dedicated IP after the decision, or extend the trial long enough for the ramp to complete.
Test validation accuracy on records you already know
Foundation includes 5,000 email validations, Basic and Foundation price additional validations at $1.20 per 100, and Scale prices them at $0.80 per 100.
Validation is easy to evaluate badly. Running it on an unknown list produces a result with nothing to check it against. Run it instead on a sample where you already know the outcome, including addresses that have previously hard bounced and addresses that have successfully received mail, then compare the verdicts against your own history.
Pay attention to how catch-all and unknown results are reported, since those are the categories that create operating decisions instead of clean answers. A platform that collapses ambiguity into a binary verdict is making a choice on your behalf that you should be making on policy grounds.
For outbound volumes the included credits function as a sample rather than a supply. A prospecting list of 20,000 records consumes four times the Foundation allowance in one pass. Our email verification API guide covers the policy layer that decides which records are permitted to send.
Test support before you need it
Support moves from ticket to phone and chat at the Scale tier. That distinction is invisible until an incident, which is precisely why it belongs in the trial.
Open a substantive ticket during the evaluation. Ask about bounce classification on your own account, or about the behaviour of a specific event type under a specific failure condition. Record the response time and whether the answer was specific to your account or a link to general documentation. Record it as a number, because it becomes the only support evidence you will have when comparing this platform against the next one.
Run the exit test while the account is new
The best moment to confirm you can leave is before you have anything to lose. During the trial, export the suppression list and confirm the format is usable. Retrieve event history through the API and check it against the console. Confirm the sending domain can be repointed without a gap.
That exercise takes an hour during an evaluation and days during a migration under pressure. It is also the check that reveals whether the arrangement you are approving leaves the evidence in your hands or in the vendor's.
LeadHaste builds client sending infrastructure so that domains, mailboxes and warm-up history are registered to the client. Our outbound lead generation services explain how the pieces fit together.
Approve Mailgun on the evidence you gathered
The review turns on retention, on whether the included dedicated IP gets a proper ramp, and on whether the validation allowance matches your list volume. Test the platform against your real reply cycle rather than against a sample send, and capture events into your own storage from day one.
LeadHaste can design and run the controlled trial, then turn the result into a quote-ready operating model with the retention and reputation plan attached. Book your free ICP and campaign-fit discovery call →
Frequently Asked Questions
A modern outbound stack includes: data enrichment (Apollo, Clay, ZoomInfo), email infrastructure (Google Workspace, custom domains), sending tools (Smartlead, Instantly), warm-up services (Warmbox), LinkedIn automation (Expandi, Dripify), CRM integration (HubSpot, Salesforce), and analytics platforms. Most agencies use 15–30 tools orchestrated together.
Building your own stack costs $3K–5K/month in software alone, plus a dedicated person to manage it. With a managed service, you get all the tooling plus the expertise to orchestrate it, often at lower total cost. The key question: can you afford to spend 6–8 weeks setting up instead of generating pipeline?
There's no single 'best' tool. It depends on your volume, budget, and integration needs. Smartlead and Instantly are popular for high-volume sending. Apollo doubles as a data and sequencing platform. The real advantage comes from how tools are orchestrated together, not from any single tool choice.
Look for three things: (1) Do you own the infrastructure they build? (2) Are the engagement terms clear, including what happens after the initial build-and-learn period? (3) Can you see transparent metrics and real case studies with specific numbers? LeadHaste starts with a three-month engagement, then moves month-to-month. Avoid vague reporting and providers that own your domains.
Data enrichment is the process of taking basic company or contact data and adding layers of detail: job titles, direct emails, phone numbers, technographics, intent signals, company size, funding stage, and more. Enrichment tools like Apollo, Clay, and ZoomInfo pull from multiple data sources to build a complete prospect profile before outreach begins.

Jacob Martinez
GTM Engineer, LeadHaste
Builds the machinery behind client campaigns: scraping, enrichment, lead scoring and the automations that keep a list clean before anyone gets emailed.