Yet-another-efficientdet-pytorch: Anchor's settings for text detection?

Created on 13 May 2020  ·  13Comments  ·  Source: zylo117/Yet-Another-EfficientDet-Pytorch

Hello @zylo117
Thank you for your great work! It's also suprise that you understand the limitation of other Pytorch's implementations :)
I'm working on text detection for invoices. My target is to recognize some important fields in it like picture below

iiiiii

The box's height about 8 -> 25 pixels after resize, while box's width about 1x ->16x bigger
For first step, i only need to detect 3 fields in green box ('form','serial', 'tax_code'). I tried D2 with default anchor's settings and it was not good. After that, i changed settings to:
anchors_scales: '[2 ** 0]'
anchors_ratios: '[(0.25, 0.25), (0.5, 0.25), (0.8, 0.25),(1.2, 0.25),(1.6, 0.25),(2.1, 0.25),(2.7, 0.25),(3.3, 0.25),(4, 0.25)]'
The result was still not good like this (final loss is about 0.3x)
sample_invoice1_visualized

I read some threads about that but i still don't know how to modify them correctly. Could you give me a hint? and do you think efficientDet is good for this type of text detection?

Most helpful comment

And it may not be a good idea to keep only one anchor scale.

thanks. At least i found my mistake for not modify anchor's setting when inference. The result of training must be better like this
sample_invoice1_visualized

All 13 comments

The loss will be faking low if anchors don't match.
Maybe try this repo to re-calculate the anchors.
https://github.com/Cli98/anchor_computation_tool

And it may not be a good idea to keep only one anchor scale.

And it may not be a good idea to keep only one anchor scale.

thanks. At least i found my mistake for not modify anchor's setting when inference. The result of training must be better like this
sample_invoice1_visualized

And it may not be a good idea to keep only one anchor scale.

thanks. At least i found my mistake for not modify anchor's setting when inference. The result of training must be better like this
sample_invoice1_visualized

Hi @titikid Could you share me the final values of anchors_scales and anchors_ratios in your case please?

Thank you very much.

@wenjun90 this one

anchors_scales: '[2 ** 0]'
anchors_ratios: '[(0.25, 0.25), (0.5, 0.25), (0.8, 0.25),(1.2, 0.25),(1.6, 0.25),(2.1, 0.25),(2.7, 0.25),(3.3, 0.25),(4, 0.25)]'

However, they are not good enough and i'm finding another one

@titikid Thank you for your answer!
Could I ask you about your dataset? How many images for training and valid, batch size and learning rate?
Have you try with D1 and D2?
Thank you again!

@titikid Thank you for your answer!
Could I ask you about your dataset? How many images for training and valid, batch size and learning rate?
Have you try with D1 and D2?
Thank you again!

@wenjun90 I tried D0 and D2.
My dataset has 340 training imgs and 20 validation imgs. I don't change batchsize and learning_rate from default settings.
What's your application?

@titikid I am working in the detection of text block too. I tried faster-rcnn well performance for mAP but not good for time of prediction on cpu. On cpu need 4s for prédiction each image. It not difficult to apply in real-time web app.

@titikid I am working in the detection of text block too. I tried faster-rcnn well performance for mAP but not good for time of prediction on cpu. On cpu need 4s for prédiction each image. It not difficult to apply in real-time web app.

If you want to detect all text block in real-time, i recommend another text detection models like Differential Binarization. I tried and it's accurate with high speed

@titikid
I am working on this project. I just wanna tried this algo to know the performance in prediction.

https://miro.medium.com/max/1200/1*gAx3-sIpo09bPDCZ2fI_kw.png

@wenjun90 i'm also interested in Document Layout recognition and i want to apply it in the future. What's your models?

@titikid Cascade rcnn. It give the best metrics but FPS is not good versus other model.

@titikid Could you tell me what's the size of your image before resizing? A followup to that is do you use the author's preprocess function to resize the image?

Was this page helpful?
0 / 5 - 0 ratings

Related issues

lockeregg picture lockeregg  ·  4Comments

HsLOL picture HsLOL  ·  8Comments

lzneu picture lzneu  ·  10Comments

tetsu-kikuchi picture tetsu-kikuchi  ·  3Comments

chunleiml picture chunleiml  ·  5Comments