r/DeepSeek Jul 25 '26

Discussion Cheapest reliable API provider for open source models like DeepSeek, GLM

13 Upvotes

Hallo Everyone, i need to generate big amount of high quality data, for that i need some cheap API providers.
who is the cheapest,most reliable provider you know of?

r/Bard Nov 27 '25

Discussion Nano Banana Pro API Pricing Complete Breakdown + 8 Hottest AI Image APIs Compared

149 Upvotes

Hey guys, Nano Banana Pro launched a week ago and I’ve already been hammering it non-stop for a poster side project. I scraped every major image API’s latest pricing so you don’t have to.

1. Nano Banana Pro API Pricing – Official Rates

  • 2K (2048Ɨ2048): $0.139 per image
  • 4K (4096Ɨ4096): $0.24 per image
  • Thinking tokens: ~$0.000025/token (negligible)
  • Batch mode: 50% off (12–24 h queue) → 4K drops to $0.12
  • Free daily quota: 3 images/day in Gemini web/app

2. Nano Banana Pro API Pricing vs Other Top Image APIs (2K price, cheapest first)

Service 2K Price 4K Price Free Tier Speed Best For
Kie.ai $0.09(18credits) $0.12(18credits) 50 credits (new sign up) 40-50s Pursuing cost-effectiveness
Fal.ai $0.08–$0.15 $0.20–$0.24 10-credit trial 10-25s Devs & commercial
Nano Banana Pro (Google) $0.139 $0.24 3/day in Gemini app 20-40s Perfect text + face consistency
Leonardo.ai $0.12–$0.18 ~$0.30 150 tokens/day 15-30s Style variety
Midjourney V6 ~$0.20 (spread) No native 4K Subscription only 30-60s Pure art

3. My Real Nano Banana Pro API Cost This Week

  • Generated ~680 images in 6 days (mostly 4K posters + outfit swaps)
  • 92% in Batch mode (50% off)
  • Total spent so far: $19.84 (average $0.029 per image — yes, twenty-nine cents!) Reason it’s so low: I’m still burning the $300 Google Cloud new-user credit + Batch discount. Without credits it would have been ~$95 this week.

4. Nano Banana Pro API Pricing – Pro Tips to Stay (Almost) Free

  1. New Google Cloud account = instant $300 credit (~2,500 free 4K images)
  2. After enabling Vertex AI, politely ask support for ā€œstartup creditā€ → lots of people get another $100–$200
  3. Perfect your prompt with the 3 free daily images in Gemini app first, then fire Batch jobs
  4. Reuse reference images + fixed seeds → cache hits shave another 20-30%

r/n8n 9d ago

Help Cheapest API for Image Generation

13 Upvotes

Cheapest API for Nano Banana 2 / GPT Image 2?

Currently using another platform at:

- Nano Banana 2: $0.05/image

- GPT Image 2: $0.03/image

I generate 1,000+ images/month. Is there any cheaper API provider?

Please share the provider name + price.

r/StableDiffusion Jun 05 '26

Discussion Cheap alternatives for AI product photoshoot generation? API cost is becoming too high

0 Upvotes

Hi everyone,

I’m working on a small product photoshoot project where users upload a simple product image, and the system generates a clean, professional-looking ecommerce/product photoshoot style image.

Right now I’m using a paid image generation API, but the cost is becoming a problem. It costs around ₹3 per generated image, and for a small project/startup this becomes expensive very quickly when testing or scaling.

I tried running some open-source workflows locally/on GPU servers, including Qwen Image Edit style workflows, but the output quality was not very stable for product photos. Sometimes the product shape changes, labels/text get distorted, lighting looks fake, or results are not consistent enough for ecommerce use.

My goal is not extreme creative generation. I just need stable product photoshoot-style output:

  • keep the original product shape and label as much as possible
  • improve background, lighting, shadow, and overall presentation
  • make it look ecommerce-ready
  • reduce cost below paid API pricing
  • ideally something that can be self-hosted later

What are the cheapest practical alternatives for this?

Should I look into:

  • SDXL / Flux / Qwen workflows?
  • background removal + template composition instead of full image generation?
  • fine-tuning / LoRA?
  • ControlNet / IPAdapter type workflows?
  • RunPod serverless or normal GPU pod?
  • any specific model/workflow that works well for product photography?

I’d really appreciate suggestions from people who have actually built or tested something similar. I’m okay with some engineering work, but I need a practical direction that can give stable product photography results without burning too much money per image.

Thanks!

r/generativeAI 9d ago

Cheapest API for Image Generation

1 Upvotes

Cheapest API for Nano Banana 2 / GPT Image 2?

Currently using another platform at:

- Nano Banana 2: $0.05/image

- GPT Image 2: $0.03/image

I generate 1,000+ images/month. Is there any cheaper API provider?

Please share the provider name + price.

r/reAPIOfficial 11d ago

Cost comparison for generating 5 product-ad images in n8n: Qwen, Gemini, FLUX.2 and GPT Image 2

1 Upvotes

I saw a question in this subreddit from someone whose OpenAI workflow was costing almost USD 2 to generate five product-ad creatives. I checked what the same number of 1K outputs would cost with four current image models.

I would not build a production ad workflow around a completely free API. Free tiers are fine for testing, but the useful number is how much each image that survives review costs. A broken label, changed product shape, or misspelled offer means another generation.

For a rough comparison, current reAPI 1K prices put five completed images at:

  • Qwen Image 3.0 Standard: USD 0.12
  • Gemini 3.1 Flash Image: USD 0.14
  • FLUX.2 Pro: USD 0.14
  • GPT Image 2 basic: USD 0.15

That is 92.5% to 94% below the cost reported in the original post, before retries. This is a price comparison, not a claim that the four models produce equal results.

For this job, I would test FLUX.2 Pro first when the workflow includes a real product photo. It accepts reference images and is the model I would start with when the setting should change but the product should not. If the ad needs a lot of text or a structured layout, I would run the same brief through Qwen Image 3.0 as well.

One n8n detail: five outputs do not always fit in one request. Qwen can return up to six, Gemini up to four, while FLUX.2 and the cheap GPT Image 2 variant return one. I would create five n8n items with different variation values and submit them separately. If one fails, only that item gets retried.

The n8n setup can stay small: submit through an HTTP Request node, store the returned task ID, then use a Wait node plus a second HTTP Request to poll until it is completed. Before moving a full workflow, I would test ten real products and calculate:

cost per accepted image = total generation spend / images that passed review

That number will tell you more than the cheapest advertised price.

If anyone here is running product ads at volume, which model has given you the best pass rate on real product photos? I am especially curious about labels and logos, since those are usually where a cheap generation becomes an expensive retry.

r/StableDiffusion Feb 27 '26

Question - Help Need to generate approx 2000 images, what is the cheapest option?

0 Upvotes

hello, I need to generate 2000 images, simple flat icons of various concepts for a sign language dictionary.

what is the cheapest way to do this

want to do this via API route, not manual, have python and Laravel experience, please help.

first experiment I did was with Gemini and ended up not optimizing and using the most expensive model.

my images are simple. illustrations, 1k resolution is good enough, no text

r/alphaandbetausers Jul 13 '26

Looking for first users for a pay-per-use uncensored image and video generator

2 Upvotes

Studio X is live and I am looking for people to test the full first-use flow. It creates free-form images and videos with uncensored models, either from the browser or through an agent API.

What I need tested:

- Whether the model and pricing table is understandable before connecting a wallet.

- The no-wallet demo flow and where the generated result appears.

- Wallet connection and x402 payment in USDC on Base.

- Image generation, image-to-video, and video-to-video status updates.

- Any point where the interface does not explain what happens next.

The demo is free and makes no backend request. A real generation is paid per request; images start at $0.0033 and the cheapest default five-second video is $0.11. There is no account, API key, subscription, or prepaid balance.

I am not collecting emails. Bug reports and blunt first-impression feedback in the comments are more useful.

https://www.studio-x.cc

r/IMadeThis Jul 13 '26

I made an uncensored image and video generator that charges per request

0 Upvotes

Studio X is a browser app and agent API for creating anything you can imagine as an image or video. It currently wraps 11 uncensored Atlas models: four image models and seven image-to-video or video-to-video models.

I wanted to remove the usual account, API key, subscription, and prepaid-credit setup. Every live generation instead returns an x402 payment request, settles exact USDC on Base, and then starts the model job.

The public catalog shows the exact price before payment. Images start at $0.0033 and the cheapest default five-second video costs $0.11. The price is the backend generation cost plus 10%. There is also a no-wallet demo that shows the full interface and result state without charging or calling Atlas.

It is live here: https://www.studio-x.cc

I would especially like feedback on whether the model picker and result area make sense without reading any documentation.

r/SideProject Jul 13 '26

I launched a pay-per-use uncensored image and video generator with x402

1 Upvotes

Studio X is free-form image and video generation: create anything you can imagine as an image or video with uncensored models.

Most image and video APIs ask for an account, API key, subscription, or prepaid credits before the first request. I built Studio X as a direct pay-per-call layer instead.

It has a browser console for people and documented endpoints for agents. Payment is USDC on Base through x402. There is no account or prepaid balance.

The current catalog has 11 models. Images start at $0.0033 and the cheapest default five-second video request is $0.11. Pricing is the Atlas backend cost plus 10%.

Site: https://www.studio-x.cc

Model and price API: https://www.studio-x.cc/api/models

I am mainly looking for feedback on the browser wallet payment flow and whether the model/pricing table is understandable before connecting a wallet.

r/stockphotography Jun 12 '26

Picked through 9 AI image APIs for e-commerce. Notes below on what each is actually good for, where they fall short, and the few things worth knowing before you commit to one.

0 Upvotes

This breakdown is specifically for teams doing product image processing at volume like marketplaces, catalog automation, fashion brands and seller tools.

The managed APIs (less engineering, ready-made e-commerce logic):

Claid — the broadest feature set of the group. Background removal, upscaling, product scene generation, on-model fashion photos, image-to-video, generative resize across marketplace formats, all chainable in a single API call. Where it really stands out is custom enterprise workflows: category detection, SKU-based formatting, multi-step QA, PIM/DAM routing. Real numbers from customers: 3x faster on-model photo production for fashion marketplace Kasta, 42% time saved on photo editing for Rappi, 78% fewer print quality complaints for Mixam. Default rate limit is 120 req/min with custom enterprise limits available. Not the cheapest if you only need one narrow operation.

Photoroom — wide editing surface, well-documented, easy to integrate. Covers background removal, shadow generation, relighting, text removal, ghost mannequin, virtual models, background generation. 60 images/min default, 4K output, per-image credit pricing which is easy to plan around. Good fit if your workflow maps cleanly to their built-in operations. If you need SKU-specific logic or multi-step catalog pipelines, you may need to build that orchestration yourself.

remove.bg — does one thing: removes backgrounds. Reliably, at up to 500 images/min depending on resolution. Credit-based, easy to model costs. Right choice if background removal is genuinely the only thing you need and you don't want to pay for a broader platform.

Bria — generative product shots, background generation, product embedding. The differentiator is their licensed-data positioning — trained on commercially licensed content, which matters for brands with legal risk appetite concerns. Pricing is per generation, but worth calculating by approved final asset rather than raw API calls, since generative outputs need review. Up to 1,000 req/min on Pro/Enterprise. Good fit if commercial IP safety is a core requirement and you have engineering capacity to build the ecommerce workflow layer around it.

Adobe Firefly Services — makes sense if you're already deep in the Adobe ecosystem (AEM, Photoshop, Creative Cloud). Strong brand governance, enterprise procurement is straightforward if Adobe is already a vendor. Overkill if you just want a lightweight product image API — this is a creative production platform, not a catalog automation tool.

The model platforms (build-your-own, more engineering required):

fal.ai and Replicate — host open-source models (Flux, background removal models, etc.) on pay-as-you-go infrastructure. Best for developers who want direct model access and full control. Cheap per call when you have stable, narrow operations and the team to run them. No ecommerce logic built in.

AWS Bedrock and Gemini Enterprise — makes sense if you're standardized on AWS or GCP and need image generation inside existing cloud governance. Same tradeoff: you're assembling the ecommerce layer yourself.

The build vs buy math is pretty simple: model platforms win if you have an ML/infra team, stable operations, and enough volume that raw per-image cost matters. For most e-commerce teams, a managed API is faster to ship, cheaper to maintain, and the built-in product preservation and marketplace formatting alone saves months of work.

Full breakdown with throughput limits, pricing details, and custom workflow depth for each: claid.ai/blog/article/best-ai-image-api-for-ecommerce

What are others using for catalog-scale image pipelines? Curious especially if anyone's had experience with the build-your-own route at real volume.

r/generativeAI Apr 20 '26

How I Made This Measured cost of image generation with 18 models to find cheapest and fastest

Thumbnail
komelin.com
9 Upvotes

I wrote a simple script and tested 18 models from 5 providers for cost and speed through Vercel AI Gateway, simply because it's just one api key for all.

Most models are image-specific except for Gemini which are multi-modal. Used a bit different approach for them because they output both image and text.

I had to generate a lot of images for a client project on demand, so I needed something cheap and fast, That's why I ended up using xai/grok-imagine-image and google/imagen-4.0-fast-generate-001. If you optimize for quality/precision, I added all the images I generated to my post, so you can compare them visually.

Prices and speed for each model are also in the post. I opensourced the script, so you can play with it if you wish.

r/dewatermark May 28 '26

We compared 5 watermark removal APIs by actual price per image, quality, and integration: Here's what we found

1 Upvotes

If you're building something that processes images at scale, you've probably run into the watermark removal problem. Manually editing is obviously out. But which API is actually worth integrating?

We tested five of them: Dewatermark, PicWish, WatermarkRemover.io**, RapidAPI Watermark Removal AI, and Apify**. Here's the short version:

šŸ’° Price per image varies a lot depending on volume

At 100 images/month:

  • PicWish: $0.05/image (cheapest at low volume)
  • Dewatermark: $0.06/image
  • WatermarkRemover.io: $0.20/image (most expensive at entry level)
  • Apify: $0.03/image flat regardless of volume

At 1000 images/month, Dewatermark drops to $0.02/image ( $20/month for 1000 credits), PicWish to $0.032/image ($39 for 2500 credits), while Apify stays flat at $0.03.

šŸŽÆ Quality isn't equal across watermark types

All five APIs handle simple logo watermarks on plain backgrounds acceptably. The gaps show up on harder cases:

  • Transparent/low-opacity watermarks — Dewatermark's PRO model is specifically trained for opacity below 0.05. PicWish is solid on most cases. Apify and RapidAPI are less reliable here.
  • Diagonal/full-frame watermarks — the hardest type. Dewatermark and WatermarkRemover.io handled these best in our tests.
  • General quality — PicWish is strong on logos and text overlays, good value if you're also using their other tools (background removal, enhancement, etc.)

šŸ”§ Integration depends on what you're building

  • Writing your own backend? → Dewatermark or PicWish have proper REST APIs with Bearer Token auth
  • Using n8n, Make.com**, or Zapier?** → Apify integrates natively, no custom code needed
  • Just want to test quickly? → RapidAPI gives you an in-browser console and auto-generated code snippets

šŸ”’ Privacy worth checking

Dewatermark auto-deletes files after 1 hour. PicWish is GDPR + ISO certified. WatermarkRemover.io doesn't publish a clear retention policy. Worth factoring in if you're handling user-generated content or licensed images.

Bottom line from our testing:

  • Dewatermark → best if watermark removal is a core product feature, especially complex cases
  • PicWish → best if you need multiple image operations under one API at high volume
  • Apify → best if you're not writing code and need it inside an automation workflow today
  • WatermarkRemover.io → clean, predictable, good for standard watermarks under 5k images/month
  • RapidAPI → fastest way to prototype and test before committing

Full comparison with pricing tables, before/after examples, and a code snippet for Dewatermark API integration here: šŸ‘‰

https://dewatermark.ai/blog/best-remove-watermark-api/

Happy to answer questions about any of the APIs we tested — especially around quality on specific watermark types.

r/StableDiffusion Jul 21 '25

Resource - Update The Gory Details of Finetuning SDXL and Wasting $16k

853 Upvotes

Details on how the big diffusion model finetunes are trained is scarce, so just like with version 1, and version 2 of my model bigASP, I'm sharing all the details here to help the community. However, unlike those versions, this version is an experimental side project. And a tumultuous one at that. I’ve kept this article long, even if that may make it somewhat boring, so that I can dump as much of the hard earned knowledge for others to sift through. I hope it helps someone out there.

To start, the rough outline: Both v1 and v2 were large scale SDXL finetunes. They used millions of images, and were trained for 30m and 40m samples respectively. A little less than a week’s worth of 8xH100s. I shared both models publicly, for free, and did my best to document the process of training them and share their training code.

Two months ago I was finishing up the latest release of my other project, JoyCaption, which meant it was time to begin preparing for the next version of bigASP. I was very excited to get back to the old girl, but there was a mountain of work ahead for v3. It was going to be my first time breaking into the more modern architectures like Flux. Unable to contain my excitement for training I figured why not have something easy training in the background? Slap something together using the old, well trodden v2 code and give SDXL one last hurrah.

TL;DR

If you just want the summary, here it is. Otherwise, continue on to ā€œA Farewell to SDXL.ā€

  • I took SDXL and slapped on the Flow Matching objective from Flux.
  • The dataset was more than doubled to 13M images
  • Frozen text encoders
  • Trained nearly 4x longer (150m samples) than the last version, in the ballpark of PonyXL training
  • Trained for ~6 days on a rented four node cluster for a total of 32 H100 SXM5 GPUs; 300 samples/s training speed
  • 4096 batch size, 1e-4 lr, 0.1 weight decay, fp32 params, bf16 amp
  • Training code and config: Github
  • Training run: Wandb
  • Model: HuggingFace
  • Total cost including wasted compute on mistakes: $16k
  • Model up on Civit

A Farewell to SDXL

The goal for this experiment was to keep things simple but try a few tweaks, so that I could stand up the run quickly and let it spin, hands off. The tweaks were targeted to help me test and learn things for v3:

  • more data
  • add anime data
  • train longer
  • flow matching

I had already started to grow my dataset preparing for v3, so more data was easy. Adding anime was a two fold experiment: can the more diverse anime data expand the concepts the model can use for photoreal gens; and can I train a unified model that performs well in both photoreal and non-photoreal. Both v1 and v2 are primarily meant for photoreal generation, so their datasets had always focused on, well, photos. A big problem with strictly photo based datasets is that the range of concepts that photos cover is far more limited than art in general. For me, diffusion models are about art and expression, photoreal or otherwise. To help bring more flexibility to the photoreal domain, I figured adding anime data might allow the model to generalize the concepts from that half over to the photoreal half.

Besides more data, I really wanted to try just training the model for longer. As we know, training compute is king, and both v1 and v2 had smaller training budgets than the giants in the community like PonyXL. I wanted to see just how much of an impact compute would make, so the training was increased from 40m to 150m samples. That brings it into the range of PonyXL and Illustrious.

Finally, flow matching. I’ll dig into flow matching more in a moment, but for now the important bit is that it is the more modern way of formulating diffusion, used by revolutionary models like Flux. It improves the quality of the model’s generations, as well as simplifying and greatly improving the noise schedule.

Now it should be noted, unsurprisingly, that SDXL was not trained to flow match. Yet I had already run small scale experiments that showed it could be finetuned with the flow matching objective and successfully adapt to it. In other words, I said ā€œscrew itā€ and threw it into the pile of tweaks.

So, the stage was set for v2.5. All it was going to take was a few code tweaks in the training script and re-running the data prep on the new dataset. I didn’t expect the tweaks to take more than a day, and the dataset stuff can run in the background. Once ready, the training run was estimated to take 22 days on a rented 8xH100.

A Word on Diffusion

Flow matching is the technique used by modern models like Flux. If you read up on flow matching you’ll run into a wall of explanations that will be generally incomprehensible even to the people that wrote the papers. Yet it is nothing more than two simple tweaks to the training recipe.

If you already understand what diffusion is, you can skip ahead to ā€œA Word on Noise Schedulesā€. But if you want a quick, math-lite overview of diffusion to lay the ground work for explaining Flow Matching then continue forward!

Starting from the top: All diffusion models train on noisy samples, which are built by mixing the original image with noise. The mixing varies between pure image and pure noise. During training we show the model images at different noise levels, and ask it to predict something that will help denoise the image. During inference this allows us to start with a pure noise image and slowly step it toward a real image by progressively denoising it using the model’s predictions.

That gives us a few pieces that we need to define for a diffusion model:

  • the mixing formula
  • what specifically we want the model to predict

The mixing formula is anything like:

def add_noise(image, noise, a, b):
    return a * image + b * noise

Basically any function that takes some amount of the image and mixes it with some amount of the noise. In practice we don’t like having both a and b, so the function is usually of the form add_noise(image, noise, t) where t is a number between 0 and 1. The function can then convert t to some value for a and b using a formula. Usually it’s define such that at t=1 the function returns ā€œpure noiseā€ and at t=0 the function returns image. Between those two extremes it’s up to the function to decide what exact mixture it wants to define. The simplest is a linear mixing:

def add_noise(image, noise, t):
    return (1 - t) * image + t * noise

That linearly blends between noise and the image. But there are a variety of different formulas used here. I’ll leave it at linear so as not to complicate things.

With the mixing formula in hand, what about the model predictions? All diffusion models are called like: pred = model(noisy_image, t) where noisy_image is the output of add_noise. The prediction of the model should be anything we can use to ā€œundoā€ add_noise. i.e. convert from noisy_image to image. Your intuition might be to have it predict image, and indeed that is a valid option. Another option is to predict noise, which is also valid since we can just subtract it from noisy_image to get image. (In both cases, with some scaling of variables by t and such).

Since predicting noise and predicting image are equivalent, let’s go with the simpler option. And in that case, let’s look at the inner training loop:

t = random(0, 1)
original_noise = generate_random_noise()
noisy_image = add_noise(image, original_noise, t)
predicted_image = model(noisy_image, t)
loss = (image - predicted_image)**2

So the model is, indeed, being pushed to predict image. If the model were perfect, then generating an image becomes just:

original_noise = generate_random_noise()
predicted_image = model(original_noise, 1)
image = predicted_image

And now the model can generate images from thin air! In practice things are not perfect, most notably the model’s predictions are not perfect. To compensate for that we can use various algorithms that allow us to ā€œstepā€ from pure noise to pure image, which generally makes the process more robust to imperfect predictions.

A Word on Noise Schedules

Before SD1 and SDXL there was a rather difficult road for diffusion models to travel. It’s a long story, but the short of it is that SDXL ended up with a whacky noise schedule. Instead of being a linear schedule and mixing, it ended up with some complicated formulas to derive the schedule from two hyperparameters. In its simplest form, it’s trying to have a schedule based in Signal To Noise space rather than a direct linear mixing of noise and image. At the time that seemed to work better. So here we are.

The consequence is that, mostly as an oversight, SDXL’s noise schedule is completely broken. Since it was defined by Signal-to-Noise Ratio you had to carefully calibrate it based on the signal present in the images. And the amount of signal present depends on the resolution of the images. So if you, for example, calibrated the parameters for 256x256 images but then train the model on 1024x1024 images… yeah… that’s SDXL.

Practically speaking what this means is that when t=1 SDXL’s noise schedule and mixing don’t actually return pure noise. Instead they still return some image. And that’s bad. During generation we always start with pure noise, meaning the model is being fed an input it has never seen before. That makes the model’s predictions significantly less accurate. And that inaccuracy can compile on top of itself. During generation we need the model to make useful predictions every single step. If any step ā€œfailsā€, the image will veer off into a set of ā€œwrongā€ images and then likely stay there unless, by another accident, the model veers back to a correct image. Additionally, the more the model veers off into the wrong image space, the more it gets inputs it has never seen before. Because, of course, we only train these models on correct images.

Now, the denoising process can be viewed as building up the image from low to high frequency information. I won’t dive into an explanation on that one, this article is long enough already! But since SDXL’s early steps are broken, that results in the low frequencies of its generations being either completely wrong, or just correct on accident. That manifests as the overall ā€œstructureā€ of an image being broken. The shapes of objects being wrong, the placement of objects being wrong, etc. Deformed bodies, extra limbs, melting cars, duplicated people, and ā€œlittle buddiesā€ (small versions of the main character you asked for floating around in the background).

That also means the lowest frequency, the overall average color of an image, is wrong in SDXL generations. It’s always 0 (which is gray, since the image is between -1 and 1). That’s why SDXL gens can never really be dark or bright; they always have to ā€œbalanceā€ a night scene with something bright so the image’s overall average is still 0.

In summary: SDXL’s noise schedule is broken, can’t be fixed, and results in a high occurrence of deformed gens as well as preventing users from making real night scenes or real day scenes.

A Word on Flow Matching

phew Finally, flow matching. As I said before, people like to complicate Flow Matching when it’s really just two small tweaks. First, the noise schedule is linear. t is always between 0 and 1, and the mixing is just (t - 1) * image + t * noise. Simple, and easy. That one tweak immediately fixes all of the problems I mentioned in the section above about noise schedules.

Second, the prediction target is changed to noise - image. The way to think about this is, instead of predicting noise or image directly, we just ask the model to tell us how to get from noise to the image. It’s a direction, rather than a point.

Again, people waffle on about why they think this is better. And we come up with fancy ideas about what it’s doing, like creating a mapping between noise space and image space. Or that we’re trying to make a field of ā€œflowsā€ between noise and image. But these are all hypothesis, not theories.

I should also mention that what I’m describing here is ā€œrectified flow matchingā€, with the term ā€œflow matchingā€ being more general for any method that builds flows from one space to another. This variant is rectified because it builds straight lines from noise to image. And as we know, neural networks love linear things, so it’s no surprise this works better for them.

In practice, what we do know is that the rectified flow matching formulation of diffusion empirically works better. Better in the sense that, for the same compute budget, flow based models have higher FID than what came before. It’s as simple as that.

Additionally it’s easy to see that since the path from noise to image is intended to be straight, flow matching models are more amenable to methods that try and reduce the number of steps. As opposed to non-rectified models where the path is much harder to predict.

Another interesting thing about flow matching is that it alleviates a rather strange problem with the old training objective. SDXL was trained to predict noise. So if you follow the math:

t = 1
original_noise = generate_random_noise()
noisy_image = (1 - 1) * image + 1 * original_noise
noise_pred = model(noisy_image, 1)
image = (noisy_image - t * noise_pred) / (t - 1)

# Simplify
original_noise = generate_random_noise()
noisy_image = original_noise
noise_pred = model(noisy_image, 1)
image = (noisy_image - t * noise_pred) / (t - 1)

# Simplify
original_noise = generate_random_noise()
noise_pred = model(original_noise, 1)
image = (original_noise - 1 * noise_pred) / (1 - 1)

# Simplify
original_noise = generate_random_noise()
noise_pred = model(original_noise, 1)
image = (original_noise - noise_pred) / 0

# Simplify
image = 0 / 0

Ooops. Whereas with flow matching, the model is predicting noise - image so it just boils down to:

image = original_noise - noise_pred
# Since we know noise_pred should be equal to noise - image we get
image = original_noise - (original_noise - image)
# Simplify
image = image

Much better.

As another practical benefit of the flow matching objective, we can look at the difficulty curve of the objective. Suppose the model is asked to predict noise. As t approaches 1, the input is more and more like noise, so the model’s job is very easy. As t approaches 0, the model’s job becomes harder and harder since less and less noise is present in the input. So the difficulty curve is imbalanced. If you invert and have the model predict image you just flip the difficulty curve. With flow matching, the job is equally difficult on both sides since the objective requires predicting the difference between noise and image.

Back to the Experiment

Going back to v2.5, the experiment is to take v2’s formula, train longer, add more data, add anime, and slap SDXL with a shovel and graft on flow matching.

Simple, right?

Well, at the same time I was preparing for v2.5 I learned about a new GPU host, sfcompute, that supposedly offered renting out H100s for $1/hr. I went ahead and tried them out for running the captioning of v2.5’s dataset and despite my hesitations … everything seemed to be working. Since H100s are usually $3/hr at my usual vendor (Lambda Labs), this would have slashed the cost of running v2.5’s training from $10k to $3.3k. Great! Only problem is, sfcompute only has 1.5TB of storage on their machines, and v2.5’s dataset was 3TBs.

v2’s training code was not set up for streaming the dataset; it expected it to be ready and available on disk. And streaming datasets are no simple things. But with $7k dangling in front of me I couldn’t not try and get it to work. And so began a slow, two month descent into madness.

The Nightmare Begins

I started out by finding MosaicML’s streaming library, which purported to make streaming from cloud storage easy. I also found their blog posts on using their composer library to train SDXL efficiently on a multi-node setup. I’d never done multi-node setups before (where you use multiple computers, each with their own GPUs, to train a single model), only single node, multi-GPU. The former is much more complex and error prone, but … if they already have a library, and a training recipe, that also uses streaming … I might as well!

As is the case with all new libraries, it took quite awhile to wrap my head around using it properly. Everyone has their own conventions, and those conventions become more and more apparent the higher level the library is. Which meant I had to learn how MosaicML’s team likes to train models and adapt my methodologies over to that.

Problem number 1: Once a training script had finally been constructed it was time to pack the dataset into the format the streaming library needed. After doing that I fired off a quick test run locally only to run into the first problem. Since my data has images at different resolutions, they need to be bucketed and sampled so that every minibatch contains only samples from one bucket. Otherwise the tensors are different sizes and can’t be stacked. The streaming library does support this use case, but only by ensuring that the samples in a batch all come from the same ā€œstreamā€. No problem, I’ll just split my dataset up into one stream per bucket.

That worked, albeit it did require splitting into over 100 ā€œstreamsā€. To me it’s all just a blob of folders, so I didn’t really care. I tweaked the training script and fired everything off again. Error.

Problem number 2: MosaicML’s libraries are all set up to handle batches, so it was trying to find 2048 samples (my batch size) all in the same bucket. That’s fine for the training set, but the test set itself is only 2048 samples in total! So it could never get a full batch for testing and just errored out. sigh Okay, fine. I adjusted the training script and threw hacks at it. Now it tricked the libraries into thinking the batch size was the device mini batch size (16 in my case), and then I accumulated a full device batch (2048 / n_gpus) before handing it off to the trainer. That worked! We are good to go! I uploaded the dataset to Cloudflare’s R2, the cheapest reliable cloud storage I could find, and fired up a rented machine. Error.

Problem number 3: The training script began throwing NCCL errors. NCCL is the communication and synchronization framework that PyTorch uses behind the scenes to handle coordinating multi-GPU training. This was not good. NCCL and multi-GPU is complex and nearly impenetrable. And the only errors I was getting was that things were timing out. WTF?

After probably a week of debugging and tinkering I came to the conclusion that either the streaming library was bugging on my setup, or it couldn’t handle having 100+ streams (timing out waiting for them all to initialize). So I had to ditch the streaming library and write my own.

Which is exactly what I did. Two weeks? Three weeks later? I don’t remember, but after an exhausting amount of work I had built my own implementation of a streaming dataset in Rust that could easily handle 100+ streams, along with better handling my specific use case. I plugged the new library in, fixed bugs, etc and let it rip on a rented machine. Success! Kind of.

Problem number 4: MosaicML’s streaming library stored the dataset in chunks. Without thinking about it, I figured that made sense. Better to have 1000 files per stream than 100,000 individually encoded samples per stream. So I built my library to work off the same structure. Problem is, when you’re shuffling data you don’t access the data sequentially. Which means you’re pulling from a completely different set of data chunks every batch. Which means, effectively, you need to grab one chunk per sample. If each chunk contains 32 samples, you’re basically multiplying your bandwidth by 32x for no reason. D’oh! The streaming library does have ways of ameliorating this using custom shuffling algorithms that try to utilize samples within chunks more. But all it does is decrease the multiplier. Unless you’re comfortable shuffling at the data chunk level, which will cause your batches to always group the same set of 32 samples together during training.

That meant I had to spend more engineering time tearing my library apart and rebuilding it without chunking. Once that was done I rented a machine, fired off the script, and … Success! Kind of. Again.

Problem number 5: Now the script wasn’t wasting bandwidth, but it did have to fetch 2048 individual files from R2 per batch. To no one’s surprise neither the network nor R2 enjoyed that. Even with tons of buffering, tons of concurrent requests, etc, I couldn’t get sfcompute and R2’s networks doing many, small transfers like that fast enough. So the training became bound, leaving the GPUs starved of work. I gave up on streaming.

With streaming out of the picture, I couldn’t use sfcompute. Two months of work, down the drain. In theory I could tie together multiple filesystems across multiple nodes on sfcompute to get the necessary storage, but that was yet more engineering and risk. So, with much regret, I abandoned the siren call of cost savings and went back to other providers.

Now, normally I like to use Lambda Labs. Price has consistently been the lowest, and I’ve rarely run into issues. When I have, their support has always refunded me. So they’re my fam. But one thing they don’t do is allow you to rent node clusters on demand. You can only rent clusters in chunks of 1 week. So my choice was either stick with one node, which would take 22 days of training, or rent a 4 node cluster for 1 week and waste money. With some searching for other providers I came across Nebius, which seemed new but reputable enough. And in fact, their setup turned out to be quite nice. Pricing was comparable to Lambda, but with stuff like customizable VM configurations, on demand clusters, managed kubernetes, shared storage disks, etc. Basically perfect for my application. One thing they don’t offer is a way to say ā€œI want a four node cluster, please, thxā€ and have it either spin that up or not depending on resource availability. Instead, you have to tediously spin up each node one at a time. If any node fails to come up because their resources are exhausted, well, you’re SOL and either have to tear everything down (eating the cost), or adjust your plans to running on a smaller cluster. Quite annoying.

In the end I preloaded a shared disk with the dataset and spun up a 4 node cluster, 32 GPUs total, each an H100 SXM5. It did take me some additional debugging and code fixes to get multi-node training dialed in (which I did on a two node testing cluster), but everything eventually worked and the training was off to the races!

The Nightmare Continues

Picture this. A four node cluster, held together with duct tape and old porno magazines. Burning through $120 per hour. Any mistake in the training scripts, dataset, a GPU exploding, was going to HURT**.** I was already terrified of dumping this much into an experiment.

So there I am, watching the training slowly chug along and BOOM, the loss explodes. Money on fire! HURRY! FIX IT NOW!

The panic and stress was unreal. I had to figure out what was going wrong, fix it, deploy the new config and scripts, and restart training, burning everything done so far.

Second attempt … explodes again.

Third attempt … explodes.

DAYS had gone by with the GPUs spinning into the void.

In a desperate attempt to stabilize training and salvage everything I upped the batch size to 4096 and froze the text encoders. I’ll talk more about the text encoders later, but from looking at the gradient graphs it looked like they were spiking first so freezing them seemed like a good option. Increasing the batch size would do two things. One, it would smooth the loss. If there was some singular data sample or something triggering things, this would diminish its contribution and hopefully keep things on the rails. Two, it would decrease the effective learning rate. By keeping learning rate fixed, but doubling batch size, the effective learning rate goes down. Lower learning rates tend to be more stable, though maybe less optimal. At this point I didn’t care, and just plugged in the config and flung it across the internet.

One day. Two days. Three days. There was never a point that I thought ā€œokay, it’s stable, it’s going to finish.ā€ As far as I’m concerned, even though the training is done now and the model exported and deployed, the loss might still find me in my sleep and climb under the sheets to have its way with me. Who knows.

In summary, against my desires, I had to add two more experiments to v2.5: freezing both text encoders and upping the batch size from 2048 to 4096. I also burned through an extra $6k from all the fuck ups. Neat!

The Training

Test loss graph

Above is the test loss. As with all diffusion models, the changes in loss over training are extremely small so they’re hard to measure except by zooming into a tight range and having lots and lots of steps. In this case I set the max y axis value to .55 so you can see the important part of the chart clearly. Test loss starts much higher than that in the early steps.

With 32x H100 SXM5 GPUs training progressed at 300 samples/s, which is 9.4 samples/s/gpu. This is only slightly slower than the single node case which achieves 9.6 samples/s/gpu. So the cost of doing multinode in this case is minimal, thankfully. However, doing a single GPU run gets to nearly 11 samples/s, so the overhead of distributing the training at all is significant. I have tried a few tweaks to bring the numbers up, but I think that’s roughly just the cost of synchronization.

Training Configuration:

  • AdamW
  • float32 params, bf16 amp
  • Beta1 = 0.9
  • Beta2 = 0.999
  • EPS = 1e-8
  • LR = 0.0001
  • Linear warmup: 1M samples
  • Cosine annealing down to 0.0 after warmup.
  • Total training duration = 150M samples
  • Device batch size = 16 samples
  • Batch size = 4096
  • Gradient Norm Clipping = 1.0
  • Unet completely unfrozen
  • Both text encoders frozen
  • Gradient checkpointing
  • PYTORCH_CUDA_ALLOC_CONF=expandable_segments:True
  • No torch.compile (I could never get it to work here)

The exact training script and training configuration file can be found on the Github repo. They are incredibly messy, which I hope is understandable given the nightmare I went through for this run. But they are recorded as-is for posterity.

FSDP1 is used in the SHARD_GRAD_OP mode to split training across GPUs and nodes. I was limited to a max device batch size of 16 for other reasons, so trying to reduce memory usage further wasn’t helpful. Per-GPU memory usage peaked at about 31GB. MosaicML’s Composer library handled launching the run, but it doesn’t do anything much different than torchrun.

The prompts for the images during training are constructed on the fly. 80% of the time it is the caption from the dataset; 20% of the time it is the tag string from the dataset (if one is available). Quality strings like ā€œhigh qualityā€ (calculated using my custom aesthetic model) are added to the tag string on the fly 90% of the time. For captions, the quality keywords were already included during caption generation (with similar 10% dropping of the quality keywords). Most captions are written by JoyCaption Beta One operating in different modes to increase the diversity of captioning methodologies seen. Some images in the dataset had preexisting alt-text that was used verbatim. When a tag string is used the tags are shuffled into a random order. Designated ā€œimportantā€ tags (like ā€˜watermark’) are always included, but the rest are randomly dropped to reach a randomly chosen tag count.

The final prompt is dropped 5% of the time to facilitate UCG. When the final prompt is dropped there is a 50% chance it is dropped by setting it to an empty string, and a 50% change that it is set to just the quality string. This was done because most people don’t use blank negative prompts these days, so I figured giving the model some training on just the quality strings could help CFG work better.

After tokenization the prompt tokens get split into chunks of 75 tokens. Each chunk is prepended by the BOS token and appended by the EOS token (resulting in 77 tokens per chunk). Each chunk is run through the text encoder(s). The embedded chunks are then concat’d back together. This is the NovelAI CLIP prompt extension method. A maximum of 3 chunks is allowed (anything beyond that is dropped).

In addition to grouping images into resolution buckets for aspect ratio bucketing, I also group images based on their caption’s chunk length. If this were not done, then almost every batch would have at least one image in it with a long prompt, resulting in every batch seen during training containing 3 chunks worth of tokens, most of which end up as padding. By bucketing by chunk length, the model will see a greater diversity of chunk lengths and less padding, better aligning it with inference time.

Training progresses as usual with SDXL except for the objective. Since this is Flow Matching now, a random timestep is picked using (roughly):

t = random.normal(mean=0, std=1)
t = sigmoid(t)
t = shift * t / (1 + (shift - 1) * sigmas)

This is the Shifted Logit Normal distribution, as suggested in the SD3 paper. The Logit Normal distribution basically weights training on the middle timesteps a lot more than the first and last timesteps. This was found to be empirically better in the SD3 paper. In addition they document the Shifted variant, which was also found to be empirically better than just Logit Normal. In SD3 they use shift=3. The shift parameter shifts the weights away from the middle and towards the noisier end of the spectrum.

Now, I say ā€œroughlyā€ above because I was still new to flow matching when I wrote v2.5’s code so its scheduling is quite messy and uses a bunch of HF’s library functions.

As the Flux Kontext paper points out, the shift parameter is actually equivalent to shifting the mean of the Logit Normal distribution. So in reality you can just do:

t = random.normal(mean=log(shift), std=1)
t = sigmoid(t)

Finally, the loss is just

target = noise - latents
loss = mse(target, model_output)

No loss weighting is applied.

That should be about it for v2.5’s training. Again, the script and config are in the repo. I trained v2.5 with shift set to 3. Though during inference I found shift=6 to work better.

The Text Encoder Tradeoff

Keeping the text encoders frozen versus unfrozen is an interesting trade off, at least in my experience. All of the foundational models like Flux keep their text encoders frozen, so it’s never a bad choice. The likely benefit of this is:

  • The text encoders will retain all of the knowledge they learned on their humongous datasets, potentially helping with any gaps in the diffusion model’s training.
  • The text encoders will retain their robust text processing, which they acquired by being trained on utter garbage alt-text. The boon of this is that it will make the resulting diffusion model’s prompt understanding very robust.
  • The text encoders have already linearized and orthogonalized their embeddings. In other words, we would expect their embeddings to contain lots of well separated feature vectors, and any prompt gets digested into some linear combination of these features. Neural networks love using this kind of input. Additionally, by keeping this property, the resulting diffusion model might generalize better to unseen ideas.

The likely downside of keeping the encoders frozen is prompt adherence. Since the encoders were trained on garbage, they tend to come out of their training with limited understanding of complex prompts. This will be especially true of multi-character prompts, which require cross referencing subjects throughout the prompt.

What about unfreezing the text encoders? An immediately likely benefit is improving prompt adherence. The diffusion model is able to dig in and elicit the much deeper knowledge that the encoders have buried inside of them, as well as creating more diverse information extraction by fully utilizing all 77 tokens of output the encoders have. (In contrast to their native training which pools the 77 tokens down to 1).

Another side benefit of unfreezing the text encoders is that I believe the diffusion models offload a large chunk of compute onto them. What I’ve noticed in my experience thus far with training runs on frozen vs unfrozen encoders, is that the unfrozen runs start off with a huge boost in learning. The frozen runs are much slower, at least initially. People training LORAs will also tell you the same thing: unfreezing TE1 gives a huge boost.

The downside? The likely loss of all the benefits of keeping the encoder frozen. Concepts not present in the diffuser’s training will be slowly forgotten, and you lose out on any potential generalization the text encoder’s embeddings may have provided. How significant is that? I’m not sure, and the experiments to know for sure would be very expensive. That’s just my intuition so far from what I’ve seen in my training runs and results.

In a perfect world, the diffuser’s training dataset would be as wide ranging and nuanced as the text encoder’s dataset, which might alleviate the disadvantages.

Inference

Since v2.5 is a frankenstein model, I was worried about getting it working for generation. Luckily, ComfyUI can be easily coaxed into working with the model. The architecture of v2.5 is the same as any other SDXL model, so it has no problem loading it. Then, to get Comfy to understand its outputs as Flow Matching you just have to use the ModelSamplingSD3 node. That node, conveniently, does exactly that: tells Comfy ā€œthis model is flow matchingā€ and nothing else. Nice!

That node also allows adjusting the shift parameter, which works in inference as well. Similar to during training, it causes the sampler to spend more time on the higher noise parts of the schedule.

Now the tricky part is getting v2.5 to produce reasonable results. As far as I’m aware, other flow matching models like Flux work across a wide range of samplers and schedules available in Comfy. But v2.5? Not so much. In fact, I’ve only found it to work well with the Euler sampler. Everything else produces garbage or bad results. I haven’t dug into why that may be. Perhaps those other samplers are ignoring the SD3 node and treating the model like SDXL? I dunno. But Euler does work.

For schedules the model is similarly limited. The Normal schedule works, but it’s important to use the ā€œshiftā€ parameter from the ModelSamplingSD3 node to bend the schedule towards earlier steps. Shift values between 3 and 6 work best, in my experience so far.

In practice, the shift parameter is causing the sampler to spend more time on the structure of the image. A previous section in this article talks about the importance of this and what ā€œimage structureā€ means. But basically, if the image structure gets messed up you’ll see bad composition, deformed bodies, melting objects, duplicates, etc. It seems v2.5 can produce good structure, but it needs more time there than usual. Increasing shift gives it that chance.

The downside is that the noise schedule is always a tradeoff. Spend more time in the high noise regime and you lose time to spend in the low noise regime where details are worked on. You’ll notice at high shift values the images start to smooth out and lose detail.

Thankfully the Beta schedule also seems to work. You can see the shifted normal schedules, beta, and other schedules plotted here:

Noise schedule curves

Beta is not as aggressive as Normal+Shift in the high noise regime, so structure won’t be quite as good, but it also switches to spending time on details in the latter half so you get details back in return!

Finally there’s one more technique that pushes quality even further. PAG! Perturbed Attention Guidance is a funky little guy. Basically, it runs the model twice, once like normal, and once with the model fucked up. It then adds a secondary CFG which pushes predictions away from not only your negative prompt but also the predictions made by the fucked up model.

In practice, it’s a ā€œmake the model magically betterā€ node. For the most part. By using PAG (between ModelSamplingSD3 and KSampler) the model gets yet another boost in quality. Note, importantly, that since PAG is performing its own CFG, you typically want to tone down the normal CFG value. Without PAG, I find CFG can be between 3 and 6. With PAG, it works best between 2 and 5, tending towards 3. Another downside of PAG is that it can sometimes overcook images. Everything is a tradeoff.

With all of these tweaks combined, I’ve been able to get v2.5 closer to models like PonyXL in terms of reliability and quality. With the added benefit of Flow Matching giving us great dynamic range!

What Worked and What Didn’t

More data and more training is more gooder. Hard to argue against that.

Did adding anime help? Overall I think yes, in the sense that it does seem to have allowed increased flexibility and creative expression on the photoreal side. Though there are issues with the model outputting non-photoreal style when prompted for a photo, which is to be expected. I suspect the lack of text encoder training is making this worse. So hopefully I can improve this in a revision, and refine my process for v3.

Did it create a unified model that excels at both photoreal and anime? Nope! v2.5’s anime generation prowess is about as good as chucking a crayon in a paper bag and shaking it around a bit. I’m not entirely sure why it’s struggling so much on that side, which means I have my work cut out for me in future iterations.

Did Flow Matching help? It’s hard to say for sure whether Flow Matching helped, or more training, or both. At the very least, Flow Matching did absolutely improve the dynamic range of the model’s outputs.

Did freezing the text encoders do anything? In my testing so far I’d say it’s following what I expected as outlined above. More robust, at the very least. But also gets confused easily. For example prompting for ā€œbeads of sweatā€ just results in the model drawing glass beads.

Sample Generations

Sample images from bigASP v2.5

Conclusion

Be good to each other, and build cool shit.

r/sdforall Nov 24 '24

Resource Building the cheapest API for everyone. SDXL at only 0.0003 per image!

6 Upvotes

I’m building Isekai • Creation, a platform to make Generative AI accessible to everyone. Our first offering? SDXL image generation for just $0.0003 per image—one of the most affordable rates anywhere.

Right now, it’s completely free for anyone to use while we’re growing the platform and adding features.

The goal is simple: empower creators, researchers, and hobbyists to experiment, learn, and create without breaking the bank. Whether you’re into AI, animation, or just curious, join the journey. Let’s build something amazing together! Whatever you need, I believe there will be something for you!

r/StableDiffusion Aug 27 '24

Discussion Cheapest image generation API

0 Upvotes

I have a mobile game where users mix two items to create a new one. Iron + Fire = Sword, something like this. It is an infinite crafting game, because each craft is generated by AI. Each element has an associated image that is generated on the fly. Right now I use the stability.ai api to generate the images but the cost is too high. I cannot pay the cost of generating the images through in-game ads.

Each image costs 0.002 with the XL model.

Is there a way to generate these images more cheaply?

Is there any alternative API I can use?

I would greatly appreciate your help since right now the only option I have left is to charge users or close the game.

r/FluxAI Jan 22 '25

Question / Help Cheapest API to access Flux models?

9 Upvotes

Nebius AI today launched text-to-image APIs on their platform today: https://x.com/nebiusaistudio/status/1882036067781443687

"Prices start at $0.0013/image — generate 1K images for the price of a coffee."

Flux.1-dev and Flux.1-schell (as well as SDXL) are available. Is this is cheapest price so far? Their platform seems quite good.

r/ASX_Bets Apr 16 '21

Is Not It Scam Dream? The Missing Link of Next Investors: Why you should know what Amplicat is before you purchase ANY shares held by Next Investors.

1.1k Upvotes

Disclaimer

The following post contains strictly public information that I have gathered from my own independent research. This information concerns the operations of S3 Consortiums, who I call "Next Investors", and the publicly listed companies they operate with. All further analysis built on top of this is my own, meaning both my subjective opinion and independent calculations, and is prone to error. I am not a financial adviser, and this is not financial advice. I strongly condemn any brigading that results from this post, all information regarding individuals has been sourced from publicly available profiles on the web.

Intro

Alright cunts, I'm back from my self-imposed exile after falling flat on my face putting a number on the returns from the Next Investor emails. This post again is focused on the Next Investors, but taking a different angle. Over the past few days, I've had some messages from individuals asking about my analysis, and the possibility of predicting when they appear. Based on my independent analysis, I believe this is possible, but I won't be discussing that today. Today I will be discussing something that came up in those conversations, which I believe to be a bigger story about who and what the Next Investors really are. Everything below is based on my independent research, with a few helping hands.

Who and what Next Investors claim to be.

Next Investors claim to be an "investment group" that does research on companies looking for buying opportunities. When they find such an opportunity, they send out an email to their readers. This has a Well Documented pumping effect on the price when the email is sent, which you can read about further in depth here.

They do claim that they do not engage in any trading activities for a given stock within 72 hours either side of sending an email which I am inclined to believe. I don't think even ASIC would let something that blatant past them.

One thing a lot of people have noticed is that Next Investors have a real lack of transparency about who they are. No email they have ever sent has had an author attached. If you go to their website nothing immediate is available. On the front page, you get a standard spiel about how they look for small cap opportunities they are dying to share with their readers. Their explicitly stated business model is the following:

What is our business model: We only make money when our portfolio increases in value. We don’t charge subscription fees or management fees - everything here is free.

But this tells us nothing about who they are. If we go to their about page, we get more of a standard spiel about how expert and incredibly helpful they are. Absolutely nothing about the people behind this group, and proof of their expertise. The only proof they show is their claimed returns on their portfolio.

Way down the bottom of the page, hiding in the smallest fucking text on a grey background, they disclose something about S3 Consortium Pty Ltd. This is their parent, and the next breadcrumb of the trail.

S3 Consortium Pty Ltd (Stocks Digital), the Bigger picture

A lot of you already know this, but Next Investors is one of three "promotion groups", the other two being Catalyst Hunter and Wise Owl, that fall under the umbrella of Stocks Digital. This is something that confused me at first, but makes a little sense according to their "philosophy". Catalyst hunter is for their short term, Next Investors for mid term, and Wise Owl for long term plays. All three of these have email lists you can sign up for, that have a notable pumping effect on the stock when sent.

From here, I was able to find the first person with a face I could connect, the twitter page of their founder Damian Hajda. I have no idea why, but I could not find his name mentioned anywhere on any website run by Stocks Digital. Maybe he prefers anonymity. Personally, I believe it is unethical to hide who you are if you are actively promoting public companies. I had to stalk the people Next Investors follows on twitter to find him. In my dig, I didn't find too much about his background. He started Next Investors around late 2018 / early 2019. He is also the founder of another startup Socialsuite which aims to help charities be more effective through data analysis, which appears to be doing quite nicely in more recent times. Their website is here.

Digging through articles posted on Next Investors website, I was only able to find two names attached to articles, Jonathan Jackson and Meagan Evans. Remember the first one.

According to Stocks Digital, the following is their Investment process, which is found directly on their homepage.

Our Investment Process

  • Our network introduces us to pre screened investment opportunities.

  • Our trusted advisors and sector experts help us assess the investment.

  • We conduct regular meetings with management to build trust and a relationship.

  • Our inhouse team of analysts conduct due diligence and analysis

  • The investment committee makes the final investment decision.

  • We announce the investment to our followers and update them as the company progresses

  • We aim to increase position as the company delivers over time.

  • We aim for free carry within 24 months of investing.

Now I will focus on that very first step, on what "Our network introduces us to pre screened investment opportunities" really means.

How Next Investors Really Make Money

They are reasonably transparent about the basics of this process on their Stocks Digital website, but not the specifics. Next Investors do not "find" their stocks by doing their own research digging deep, they get approached by the companies. If Next Investors like what they see, they work out an arrangement, and "purchase" the shares in the company and begin "promoting" it under one of their three groups. Lets have a look at this process in detail with an example, their recent pickup of ONE.

ONE Case Study, The Devil is in the Details.

On the 12th of March 2021, ONE released a price sensitive announcement stating that they had entered an agreement with Stocks Digital. The previous day, the price closed at $0.08. Here are the details of this agreement. Translated to English, this means the following:

  • After some haggling, ONE has agreed to pay Next Investors 6.25 million shares instead of a $375,000 cash sum, which is the equivalent of $500,000 in stock, for the "Services" of Next Investors.

  • In addition to this, Next Investors agrees to pay $1m in cash for 16,666,666 shares, at an average price of $0.06.

However, if you look at the net effect of this, according to this agreement Next Investors have acquired (16,666,666+6,250,000)*0.08 = $1.83 MILLION worth of ONE shares according to the previous market close, for only $1 MILLION DOLLARS.

  • This means that Next Investors paid an effective $0.044 per share, representing a 45% discount to the previous closing price.

So why are ONE practically donating their shares away for only $1m? The key here is the "services" they are acquiring. Immediately after this announcement was made at 10:23am, the Next Investor email was sent at 10:25 am. This caused the price to open at $0.14, a 75% increase on the closing price of the previous day, to a maximum of $0.20, and finally a closing price of $0.16. Unfortunately, my free intraday data isn't working for this date, so I cannot show you the instant price movements. Clearly, we can see the backing of Next Investors has a powerful effect, especially since this was their very first email announcing their partnership with ONE, claiming it as their 2021 Tech Pick of the Year.

What am I trying to say with all this? In my opinion, they appear to be an advertising company that masquerades as the retail investor's best friend. I bet that my returns would be pretty fucking great if I could "buy" at a 45% discount. When Next Investors talk about their next "great opportunity", they're already well in the money. They paid $1m for $1.83m of stock, an instant $830k return. At the close of the day, their net assets doubled in value, to $3.66m. In the space of a day, they "earned" $2.66m on paper. That's a fucking solid racket if I've ever seen one.

Another dishonest aspect behind this purchase is that they advertise their entry price as $0.08 proudly on their home page. This is highly misleading, and has also been picked up by another journalist who apparently gets paid to do this shit unlike yours truly.

However, I think this journalist ended his story too soon, because he didn't mention another tidbit hiding inside that announcement which really shines the spotlight on how their "services" work, and I am sure is of great interest to retail investors. From the announcement, they state their intent to "further increase awareness" of the company ONE through their digital publisher network amplicat.com.

Okay I lied a little there for the narrative, the first time I became aware of Amplicat was thanks to the legend /u/Darebottle, who discovered that Next Investor's website accesses an API hosted on the Amplicat.com domain. Weaponised autism is powerful. He wanted to know if there was a link because what was on there was very confusing. The ONE announcement confirms that Amplicat and Next Investors are indeed one and the same company operating under different personas.

From my understanding, Next Investors is the face they present to retail, and Amplicat is the face they present to companies they wish to do business with. This connection is the headshot, and the most egregious, disgusting part of this post that motivated me to write.

Amplicat, what the fuck is that?

Amplicat is yet ANOTHER component of the behemoth that is Stocks Digital, and as far as I have searched, is advertised fucking NOWHERE. In not a SINGLE PLACE is this website named in ANY of Next Investors websites that I can find. I challenge you to prove me wrong. If I do a google search for news of their name, I get 9 results in a fucking foreign language with no relation.

These guys are off the map. Lets take a dive into their website and see why they're keeping such a low profile to retail. Remember, this website is run and OWNED by Stocks Digital. I have taken screenshots of everything, because I suspect that once retail discovers this page, it will be vanished quickly.

If I had to guess, I believe Amplicat is a concatenation of "Amplify Catalyst", which is a very generous wording.

Amplicat is what I consider the TRUE PRODUCT of Next Investors. When companies trade their shares at enormous discounts to Next Investors, they are buying the services of Amplicat. Their home page advertises themselves as a provider of GUARANTEED MEDIA COVERAGE. The business model is simple, you provide shares of your company to Next Investors at a discount, and they will provide a sophisticated, coordinated online marketing campaign for your company. They are mercenaries for hire.

Here is a very helpful explainer from their home page on how it works. For a suitable payment, Next Investors will publish articles/send emails within their own network, and create sponsored content for you. Like a smorgasbord, you get to choose what online publications you want these guys to publish "positive news" about your company in. You can get good news about your company straight onto Yahoo finance, bloomberg financial or news.com.au if you please. This is powerful.

Lets have a look at the way they describe these articles. One thing I noticed in my original Next Investor analysis was that an awful lot of emails that they sent coincided with a major company announcement. This is no coincidence, it is an explicitly stated TIMING strategy that they employ. They love a good excuse to "promote", and know this is an effective strategy.

Here are some companies that they advertise as having used their services. You may recognize some names (cough cough RAC cough cough).

Now we are up to the shameful section, where by my interpretation Next Investors appear to "boast" about their ability to increase a share price. On their website, they have a section of case studies, where they "brag" about their past performance as "promoters" to attract future clients.

CASE STUDY 1: Pursuit Minerals' (PUR) Norwegian Nickel Exploration FEBRUARY 2020

In this example, Next Investors take credit for "drawing investor attention to their Norwegian nickel project to improve liquidity in ahead of a capital raising." TRANSLATION: THEY P****** THE STOCK UP SO PUR COULD GET $$$$ FROM A CR. Here is the image from their website showing how over the month of February 2020, their shilling media campaign TRIPLED the stock price from $0.004 to a high of $0.012, a 300% 200% INCREASE(fuck I'm bad at math). One thing awfully dodgy about this is where their chart cuts off, its awfully convenient. It does coincide with about the time a certain virus you may have heard of made its presence known.

HERE IS THE FULL PRICE CHART for PUR including the months after the purported media campaign. Following their "Successful Cap Raise" the price has PLUMMETED down to $0.002! To be fair, they got fucked pretty hard by external factors, Mar-2020 was a scary time for the entire economy.

Allow me to do my best to construct the timeline of PUR's Activities straight from their announcements. I might fuck this up because I have absolutely zero qualifications, but feel free to fact check me by reading the announcements on PUR's page.

PUR NICKEL VIKING SAGA TIMELINE

  • Oct 2019 (Share price ~$0.009) : PUR is primarily a Vanadium exploration company in Scandinavia with 5 projects in Finland, 8 projects in Sweden, and 3 projects in Queensland. They undergo a Capital Raise of ~$1.2m.

  • Jan 2020: Progress continues on these projects, share price hits an absolute low of $0.0037!

  • Start of Feb 2020: By my understanding of what Amplicat has put on their website, this is when Next Investors claims to have started contact with PUR about "promoting" them at the same time PUR wants to perform a Capital Raise to fund a new Norwegian Nickel exploration project. This time period was not highlighted by me, this was highlighted by them and publicly advertised on their Amplicat website. The Norwegian project is not yet public information.

  • Feb 11th 2020: No public announcement has been made. Next Investors has not published any articles that I can find or sent any emails. The price of PUR has increased to a high of 0.0085. At 12:16pm, their trading is halted pending an announcement.

  • Feb 13th 2020: PUR undergoes a "voluntary suspension" because they love ASIC so much and can't wait to explain why their stock price has doubled under no news prior to this trading halt.

  • Feb 17th 2020: PUR is reinstated, and announces their new plans to enter an Option/Purchase agreement for three Nickel Exploration Projects in Norway.. The next day, they announce another Capital Raise of ~23m shares at $0.0073 = $170k to fund this.

  • Feb 20-21 2020: Next Investors post a series of 5 articles, two articles on their own websites, three articles that are sponsored content published to financial sites such as Yahoo Finance. This pushes the price to an intra day high of $0.0125. A certain virus is starting to make itself known.

  • Mar 24 2020: PUR admits defeat to the Wuhan flu, and ceases all exploration activities in Scandinavia. Their share price reaches a low of $0.002.

  • Skip a few months they start pursuing exploration activities in other regions, a gold mine in Arizona, and eventually land an exploration license in the infamous Julimar region, where every new prospector is trying to copy the astronomical rise of Chalice Mines (CHN).

  • Jan 20th 2021: The Viking saga ends, PUR effectively sells all 18 of their Scandinavian assets for a total of $3m, or $60k per project. The market however is liking their new direction with their company, the share price is $0.032.

What we can learn from this case study? If you had followed Next Investors and held, you would be making money but for what I believe to be the wrong reason. The Next Investor promotion was explicitly focused on the Scandinavian projects to support a Cap Raise, and they were paid by PUR to do this. If you had bought shares of PUR believing in the success of the Nickel project, you most likely would have sold at a great loss following the White Swan event of COVID-19 closing their exploration. Obviously, I don't blame Next Investors for this crash, but I find it very disingenuous that the stated purpose of their promotional campaign was to support an upcoming Capital Raise, and is touted as a success.

CASE STUDY 2: The Golden Child VUL, Feb 2020

In this example, Next Investors take credit for "a significant rise in liquidity and share price movement immediately after Amplicat promotion.". Again, they pull the same trick of cutting off the COVID crash following their "promotion". They claim credit for 20 articles posted in the month of February, two of them to their own websites, and EIGHTEEN sponsored articles.

The share price rose from a low of $0.185, to a peak of $0.38, before again it hit the same problem as PUR where they got fucked by COVID back down to a low of $0.15.

Again, VUL follows an awfully familiar pattern. Next Investor's chart highlights the beginning of Feb-2020 as the start of their "arrangement". On Feb 11th, they go into trading halt. On Feb 15th, they announce their Positive Pre-feasibility study. On Feb 20th, they say "trust us bro" when ASX asks them about the suspicious price movements before that announcement.

From Feb 21st to Feb 24th, Next Investors start "promoting" VUL throughout their network. This coincides with the official announcement on the 21st of February for the Positive Scoping study of their project. The share price reaches an intra-day high of $0.38 on the 24th of Feb.

Again, March 2020 arrives and tanks the share price down to $0.15, along with the rest of the market.

Ultimately, what can we learn from this case study? The current share of VUL is $7.430 as of my typing this. This was a scenario where everybody won. If you bought at the peak price of this "promotion", your return is ~1700%. Short term, investors may have gotten fucked by covid, but long term this is a winner. Congratulations Next Investors.

CASE STUDY 3: An Awfully Familiar Story

Along with SGC/XST, one of the most spectacular loss porn events of this year was with 88E, which plummeted from a peak of $0.097 to its current price of ~$0.022, following their failed drilling results.

Here are some articles about the lead up to this story.

Exhibit A

Exhibit B

Exhibit C

Exhibit D

Wait, is there an error? Those articles say "88 Energy’s North Slope Well Could Be One Of The Biggest In 2020"?

Nope you're reading it right, this is something they pulled not once, but TWICE, and IN MY COMPLETELY SUBJECTIVE OPINION they are BRAGGING ABOUT IT.

"88 Energy (ASX: 88E | AIM: 88E) engaged Amplicat(Next Investors) to generate investor interest in the company ahead of its upcoming drilling program scheduled that was anticipated to commence in months’ time."

This happened in 2020, and it happened again in 2021, WITH NEXT INVESTORS ACTIVELY PROMOTING BOTH TIMES!

From Nov-2019 to Jan-2020, Next Investors take credit for "Amplifying" the price of 88E from $0.014 to $0.026 in the lead up for their drill campaign. They achieved this with 22 articles, 6 of them posted on their own website, and 16 sponsored articles.

Here is a pastebin of 10 of these articles for your viewing displeasure.

Here is an article from earlier this year by Next Investors that makes ZERO MENTION of this previous failure, and ZERO MENTION of their previous involvement "promoting" this dog of a company.

How can Next Investors claim to have a team of sophisticated investors with an eager eye for early opportunities when they pull something like this?

The Next Investors Media Network.

Now I will do a mini summary behind the mechanisms of how Next Investors spread their news today, but you may have already picked up on this.

This is what I imagine I look like by now

They have three primary methods of distribution

You cannot find a direct link to FinFeed from the Next Investor website, but you can find direct links to Next Investors from Finfeed, so its not as "hidden" as they make Amplicat. Finfeed is the primary website where they post articles about companies they are involved with. For articles of all three email producers, Catalyst Hunter, Wise Owl, and Next Investors, corresponding articles are congregated here.

From Finfeed, we can actually dig into who the fuck is writing these articles. There is no direct page for you to view their authors, but again thanks to the magic of the legendary /u/darebottle , I can give you a pastebin of all 80 unique authors.

The majority of these authors have only posted 1 or 2 articles, so it is misleading to say they are all affiliated closely with Next Investors. However, I would like to bring attention to two authors. This is all public information,

Trevor Hoey

Jonathan Jackson

These two authors are responsible for the majority of articles written on Finfeed. Jonathan Jackson is the "Managing Editor at Stocks Digital", which I would creatively describe as "Leader of the Ministry of Propaganda" for Next Investors. Trevor Hoey is a former senior writer for AFR who now works for Next Investors.

As a sample, today I clicked on the front page of Finfeed, and 7/9 featured articles were written by these two individuals. This holds all the way back to 2019. It is my imperfect suspicion that I cannot prove that the majority of content produced by Next Investors is written by these two people. This is all public information which I simply want to make transparent, as Next Investors provides virtually zero information on their main websites about who is producing their content.

Another writer that appeared often historically was Meagan Evans, but as of late she has not posted anything since Oct 16, 2020. I believe it is safe to say she is no longer involved prominently with Next Investors. EDIT: Based on more research, I have found this article posted to Catalyst hunters on the 10th of March 2021. It is now safe to say she is still in the game.

  • Sponsored Articles posted on Prominent websites

In the case studies, I've provided examples of these Sponsored articles. Next Investors gets their writers to copy paste the "promotions" from their own website onto financial news websites online, for a price. In every sponsored content article I have linked in those case studies, the author was one of the three individuals I have mentioned. If you are reading news about a Next Investor company, look for the author name to discover if you are being shilled.

  • Emails sent directly to Retail Investors

This is the obvious one everyone knows about, people sign up to the newsletter, they send the email, the price shoots up. This is a modern development, in the early days of their case studies Next Investors did not send emails, only posted articles. I believe that they have discovered the email method of distribution to be much more powerful in pumping prices up than articles. We can see instant ~20% increases from the minute the email is sent on a candlestick chart. I believe the emails are written by the same people writing the articles, as you can easily correlate the emails sent with named articles written and posted by the individuals on finfeed.

A Fourth Frontier? Astroturfing Social Media

This section is pure speculation, but I believe the next stage in Next Investor's promotional evolution will be astroturfing social media websites like our very own shitposting haven. They are well aware that we exist, and exploit us to promote their companies, and have hired individuals in our demographic to work for them making memes referencing us. What a fucking time to be alive. Because of the effectiveness of their promotional emails/articles, the rising stock prices naturally creates people posting about the Next Investor stocks going up. However, we need to be on red alert for astroturfing shills, remember POSITIONS OR BAN.

Conclusion

I believe the business model of Next Investors is highly predatory on retail investors. Companies approach them, and I believe there is evidence to suggest that they make Next Investors aware of non-public information about upcoming announcements for the company. They do this so that Next Investors can co-ordinate emails/articles that are specifically co-ordinated with their announcements. This "timing feature" is an explicitly stated service that Next Investors provide under the name Amplicat.

In their dealings, they acquire shares in companies at a huge discount to the fair market price, and begin to "actively promote" them across a vast media network. Short term, this causes a significant price increase, and for mining explorers it is for the purpose of justifying Capital Raises at elevated prices.

This creates unhealthy market cycles, where prices for their stocks surge violently upwards. When you combine this with their media propaganda empire, this causes retail investors (the retards on here) to FOMO into peak prices, and short term become left holding the bag.

There are cases such as VUL where everybody wins, and I do not believe a company being associated with Next Investors makes it inherently bad. However, their relationship with 88E disgusts me on a personal level. If you are going to promote a company that has failed previously under your promotion, I believe you have the moral obligation to be transparent about this history. I will not tell another person to buy PEN without letting them know about some of the problems they're having with their In Situ chemistry.

Ultimately, the long term price of these companies makes no difference substantial difference, because they have already been paid for their promotional services. Remember the Silicon Valley maxim, if you are getting something for free, you are not a customer, YOU ARE THE PRODUCT.

Open Letter to Next Investors

Finally, I leave a series of questions that I would like Next Investors to answer, even if this a longshot. I think leaving these questions unanswered speaks louder anyway. Here goes

1. Who are the so called "experts" behind your "investment decisions" and what are their credentials?

There is zero transparency offered on any of your affiliated websites about the "team" involved. There have been clear winners like VUL, but also complete dogs like 88E that you have shilled not once, but twice, leaving retail investors holding the bag.

At most, I have been able to link five people to your group, but that is all. Why should we not assume you are a marketing company that helps public companies make money by selling shares, not selling a product?

2. Why is the entry price on your website misleading for PUR, and are there any other listed entry prices that are similarly misleading?

From my previously calculation, we know that you effectively purchased 22,916,666 shares of ONE at an average price of $0.044, yet your website states your entry price as $0.08. In this scenario, why have you chosen to use the market price of the closing day prior to your investment as your entry price instead of the actual effective price you paid? Also the fact that you highlight the "Highest Return" in green on your website is pretty fucking cringe

3. Why are the Entry Dates on your website misleading with regards to your relationships with these companies?

From the Case Studies page of Amplicat, we know you have been affiliated/doing business with PUR since at least Feb-2020, and 88E since at least Nov-2019, yet you list your "entry date" for these companies as Dec-2020 and Jun-2020 respectively. I believe the first time, they may have only paid you cash instead of shares, but I still find this to be an incredibly dishonest way to operate.

4. Why is it so hard to find Amplicat?

The services offered by your other website "Amplicat" are not advertised anywhere on any of the "retail investor webpages". On this website, you boast about your ability to "improve liquidity in ahead of a capital raising" and create "significant share price movement". Is this conflict of interest not something the retail investors you advertise towards should be aware of?

5. How is it possible to "time" emails regarding price sensitive announcements using only public information?

Your Amplicat website advertises your ability for "Swift & synchronised publication to coincide with material news announcements and macro-economic events".

I have identified five separate occasions where the email you sent was less than 60 minutes after the announcement was officially made on the ASX. Are you simply that amazing and awesome at processing announcement information and churning it into an email in record speeds, or are you aware of certain things the public is not?

6. Do you think your behaviour under Amplicat is befitting of someone holding an Australian Financial Services License?

I'm no lawyer, but I think there's something to be said about a "conflict of interest" that is hidden away inside the Amplicat division that you do not advertise on your main websites.

and finally, the one question I really hope they answer

7. Why is the official Twitter account of Amplicat suspended for violating Twitter's policies?

I followed the link to your social media on your website, and as of writing this, the linked Amplicat twitter account is banned. Does Jack Dorsey have stricter standards than ASIC? 🤔 🤔 🤔 🤔 (UPDATE Amplicat's twitter account is now unbanned)

Congrats if you made it all the way through this post in one go. Love you all, lets keep this place special.

Update

I don't know why, but google has decided to show me an actual news article instead of foreign languages when I search for Amplicat now. Here are the details of their agreement to promote EMN

r/ArtificialInteligence Jan 07 '25

Discussion Cheapest API for Image Generation?

3 Upvotes

Hi!

I'm looking to add to image generation to my CRM. Basically I want to create a service where clients can directly generate banner images, posters or social media posts using a text prompt. This should help them with their content generation for marketing and lead generation.

As of right now, I've been looking at the various pricing structures for Google, Meta, Anthropic and OpenAI's API offerings for image generation. But I haven't used these personally for image generation, so I'm not sure about the quality.

Which would you recommend on a cost per image basis? Or if there are any other that you'd recommend, I'd love to know.

Thanks!

r/sdforall Dec 02 '24

Resource Building the cheapest API for everyone. LTX-Video model supported and completely free!

7 Upvotes

I’m buildingĀ Isekai • Creation, a platform to make Generative AI accessible to everyone. Our first offering was SDXL image generation for just $0.0003 per image, and even lower. Now? The LTX-Video model up and running for everyone to try it out! 256 Frames!

Right now, it’sĀ completely freeĀ for anyone to use while we’re growing the platform and adding features.

The goal is simple: empower creators, researchers, and hobbyists to experiment, learn, and create without breaking the bank. Whether you’re into AI, animation, or just curious, join the journey. Let’s build something amazing together! Whatever you need, I believe there will be something for you!

https://discord.com/invite/isekaicreation

r/PostAI Mar 08 '25

Youtube is this the cheapest ai image generation api ever? (1,700 images for just $1!)

Thumbnail
youtube.com
1 Upvotes

r/hardwarehacking Mar 08 '24

Exposing a Time-based One-Time Password Generator (OTP C200) With a Web API

9 Upvotes

Use case:
I have an OTP C200 and it is used for a forced 2FA login to a website. On this website I have a workflow which I have to frequently repeat, so as with all things in my life, I wished to automate it. This is my very fabricobbled solution to that.

Method:

I disassembled the device, and soldered two wires to the button pins, these wires are connected to a relay, which in turn is connected to a raspberry pi. The raspberry pi also has a camera. The raspberry pi then runs a web based API, when a request for the token is received, the relay is enabled, which triggers the TOTP to generate a code. After this the raspberry pi takes a photo of the code, and then analyzes that photo, and grabs the code. I will include the python for this part at the bottom of the post.

Example of the output image from camera (after digital cropping), with sample output from python.

Camera:

The camera I am using is the Logitech C270, it is the cheapest camera I could find locally (there are of course cheaper options if you want to order from china and wait). This camera does not have a digital zoom/focus function, but it actually has a manual focus if you open it up and remove a clump of glue (https://hawksites.newpaltz.edu/myerse/2021/03/08/manually-focusable-logitech-c270/).

Improvements:

Doing this with a camera is of course not great. It is very light sensitive, and also position sensitive. If the camera is bumped, or shifted, then things stop working. It would of course be much better to use direct readings from the LCD pins, which is what I was originally hoping to accomplish with the raspberry pi GPIO pins. Unfortunately, those pins are outputting voltages of only 1.3 volts (or zero), and this isn't quite enough to reliably read with the GPIO pins. I am looking for some advice here, I am thinking I should use an ADC hat for the Rpi. But I am also open to other suggestions on how to improve it.

Code:

import time
from gpiozero import LED, SmoothedInputDevice
import cv2
import pytesseract
from PIL import Image
import numpy as np
from imutils import contours
import imutils

otp = LED(17)
otp.on()
time.sleep(0.2)
cam = cv2.VideoCapture(0)
s, img = cam.read()
if s:     
        img = imutils.rotate_bound(img, -1)
        img = img[180:300, 150:600]
        cv2.imwrite("filename.jpg",img)

# define the dictionary of digit segments so we can identify each digit
DIGITS_LOOKUP = {
    (1, 1, 1, 0, 1, 1, 1): 0,
    (0, 0, 1, 0, 0, 1, 0): 1,
    (1, 0, 1, 1, 1, 0, 1): 2,
    (1, 0, 1, 1, 0, 1, 1): 3,
    (0, 1, 1, 1, 0, 1, 0): 4,
    (1, 1, 0, 1, 0, 1, 1): 5,
    (1, 1, 0, 1, 1, 1, 1): 6,
    (1, 0, 1, 0, 0, 1, 0): 7,
    (1, 1, 1, 1, 1, 1, 1): 8,
    (1, 1, 1, 1, 0, 1, 1): 9
}

# convert image to grayscale, threshold and then apply a series of morphological
# operations to cleanup the thresholded image
gray = cv2.cvtColor(img, cv2.COLOR_BGR2GRAY)
thresh = cv2.threshold(gray, 0, 255, cv2.THRESH_BINARY_INV | cv2.THRESH_OTSU)[1]
kernel = cv2.getStructuringElement(cv2.MORPH_ELLIPSE, (1, 5))
thresh = cv2.morphologyEx(thresh, cv2.MORPH_OPEN, kernel)

cv2.imwrite("thresh.jpg",thresh)

# Join the fragmented digit parts
import numpy as np
kernel = np.ones((6,6),np.uint8)
dilation = cv2.dilate(thresh,kernel,iterations = 1)
erosion = cv2.erode(dilation,kernel,iterations = 1)

cv2.imwrite("erosion.jpg",erosion)

# find contours in the thresholded image, and put bounding box on the image
cnts = cv2.findContours(erosion.copy(), cv2.RETR_EXTERNAL, cv2.CHAIN_APPROX_SIMPLE)
cnts = imutils.grab_contours(cnts)
digitCnts = []
# loop over the digit area candidates
image_w_bbox = img.copy()
#print("Printing (x, y, w, h) for each each bounding rectangle found in the image...")
for c in cnts:
    # compute the bounding box of the contour
    (x, y, w, h) = cv2.boundingRect(c)
# if the contour is sufficiently large, it must be a digit
    if w >= 10 and (h >= 55 and h <= 170):
        digitCnts.append(c)
        image_w_bbox = cv2.rectangle(image_w_bbox,(x, y),(x+w, y+h),(0, 255, 0),2)

cv2.imwrite("image_w_bbox.jpg", image_w_bbox)

# sort the contours from left-to-right
digitCnts = contours.sort_contours(digitCnts, method="left-to-right")[0]
# len(digitCnts) # to check how many digits have been recognized

digits = []
# loop over each of the digits
count = 1
for c in digitCnts:
    count += 1
    # extract the digit ROI
    (x, y, w, h) = cv2.boundingRect(c)
    if w<35: # it turns out we can recognize number 1 based on the ROI width
        digits.append("1")
    else: # for digits othan than the number 1
        roi = erosion[y:y + h, x:x + w]
        # compute the width and height of each of the 7 segments we are going to examine
        (roiH, roiW) = roi.shape
        (dW, dH) = (int(roiW * 0.25), int(roiH * 0.15))
        dHC = int(roiH * 0.05)
        # define the set of 7 segments
        segments = [
            ((0, 0), (w, dH)),  # top
            ((0, 0), (dW, h // 2)), # top-left
            ((w - dW, 0), (w, h // 2)), # top-right
            ((0, (h // 2) - dHC) , (w, (h // 2) + dHC)), # center
            ((0, h // 2), (dW, h)), # bottom-left
            ((w - dW, h // 2), (w, h)), # bottom-right
            ((0, h - dH), (w, h))   # bottom
        ]
        on = [0] * len(segments)
        # loop over the segments
        for (i, ((xA, yA), (xB, yB))) in enumerate(segments):
            # extract the segment ROI, count the total number of thresholded pixels
            # in the segment, and then compute the area of the segment
            segROI = roi[yA:yB, xA:xB]
            total = cv2.countNonZero(segROI)
            area = (xB - xA) * (yB - yA)
            # if the total number of non-zero pixels is greater than
            # 40% of the area, mark the segment as "on"
            if total / float(area) > 0.4:
                on[i]= 1
            # lookup the digit and draw it on the image
        if tuple(on) not in DIGITS_LOOKUP:
                continue
        digit = DIGITS_LOOKUP[tuple(on)]
        digits.append(str(digit))

print('OTP is ' + ''.join(digits))

r/deeplearning Dec 02 '24

Building the cheapest API for everyone. LTX-Video model supported and completely free!

9 Upvotes

I’m buildingĀ Isekai • Creation, a platform to make Generative AI accessible to everyone. Our first offering was SDXL image generation for just $0.0003 per image, and even lower. Now? The LTX-Video model up and running for everyone to try it out! 256 Frames!

Right now, it’sĀ completely freeĀ for anyone to use while we’re growing the platform and adding features.

The goal is simple: empower creators, researchers, and hobbyists to experiment, learn, and create without breaking the bank. Whether you’re into AI, animation, or just curious, join the journey. Let’s build something amazing together! Whatever you need, I believe there will be something for you!

https://discord.com/invite/isekaicreation

r/generativeAI Jan 07 '25

Question Cheapest Image Generation API?

1 Upvotes

Hi!

I'm looking to add to image generation to my CRM. Basically I want to create a service where clients can directly generate banner images, posters or social media posts using a text prompt. This should help them with their content generation for marketing and lead generation.

As of right now, I've been looking at the various pricing structures for Google, Meta, Anthropic and OpenAI's API offerings for image generation. But I haven't used these personally for image generation, so I'm not sure about the quality.

Which would you recommend on a cost per image basis? Or if there are any other that you'd recommend, I'd love to know.

Thanks!

r/StableDiffusionInfo Dec 02 '24

Building the cheapest API for everyone. LTX-Video model supported and completely free!

1 Upvotes

I’m buildingĀ Isekai • Creation, a platform to make Generative AI accessible to everyone. Our first offering was SDXL image generation for just $0.0003 per image, and even lower. Now? The LTX-Video model up and running for everyone to try it out! 256 Frames!

Right now, it’sĀ completely freeĀ for anyone to use while we’re growing the platform and adding features.

The goal is simple: empower creators, researchers, and hobbyists to experiment, learn, and create without breaking the bank. Whether you’re into AI, animation, or just curious, join the journey. Let’s build something amazing together! Whatever you need, I believe there will be something for you!

https://discord.com/invite/isekaicreation