peter-el-hachem commited on
Commit
325b8a9
ยท
verified ยท
1 Parent(s): f1117ff

Upload folder using huggingface_hub

Browse files
Files changed (3) hide show
  1. .gitattributes +1 -0
  2. README.md +15 -90
  3. architecture.png +3 -0
.gitattributes CHANGED
@@ -33,3 +33,4 @@ saved_model/**/* filter=lfs diff=lfs merge=lfs -text
33
  *.zip filter=lfs diff=lfs merge=lfs -text
34
  *.zst filter=lfs diff=lfs merge=lfs -text
35
  *tfevents* filter=lfs diff=lfs merge=lfs -text
 
 
33
  *.zip filter=lfs diff=lfs merge=lfs -text
34
  *.zst filter=lfs diff=lfs merge=lfs -text
35
  *tfevents* filter=lfs diff=lfs merge=lfs -text
36
+ architecture.png filter=lfs diff=lfs merge=lfs -text
README.md CHANGED
@@ -130,7 +130,7 @@ doc = converter.convert(source=source).document
130
 
131
  print(doc.export_to_markdown())
132
 
133
- ````
134
  </details>
135
 
136
 
@@ -139,7 +139,7 @@ Alternatively, you can use bare **transformers**, **vllm**, **onnx** or **mlx-vl
139
  <details>
140
  <summary>๐Ÿ“„ Single page image inference using plain ๐Ÿค— tranformers ๐Ÿค–</summary>
141
 
142
-
143
  # Run batch inference
144
  start_time = time.time()
145
  outputs = llm.generate(batched_inputs, sampling_params=sampling_params)
@@ -196,7 +196,11 @@ The evaluation can be performed using the [docling-eval](https://github.com/docl
196
  </tr>
197
  <tr>
198
  <td><b>granite-docling-258m</b></td>
199
- <td><b>0.27</b></td><td><b>0.86</b></td><td><b>0.92</b></td><td><b>0.88</b></td>
 
 
 
 
200
  </tr>
201
  </tbody>
202
  </table>
@@ -222,98 +226,18 @@ The evaluation can be performed using the [docling-eval](https://github.com/docl
222
  </tr>
223
  <tr>
224
  <td><b>granite-docling-258m</b></td>
225
- <td><b>0.45</b></td><td><b>0.84</b></td><td><b>0.91</b></td>
226
- <td><b>0.83</b></td><td><b>0.65</b></td><td><b>0.72</b></td>
227
  </tr>
228
- </tbody>
229
- <thead>
230
- <tr><th colspan="7"><b>Code Recognition</b></th></tr>
231
  <tr>
232
- <th></th>
233
- <th>Edit-distance โ†“</th>
234
- <th>F1 โ†‘</th>
235
- <th>Precision โ†‘</th>
236
- <th>Recall โ†‘</th>
237
- <th>BLEU โ†‘</th>
238
- <th>Meteor โ†‘</th>
239
- </tr>
240
- </thead>
241
- <tbody>
242
- <tr>
243
- <td><b>smoldocling-256m-preview</b></td>
244
- <td>0.114</td><td>0.915</td><td>0.94</td><td>0.909</td><td>0.875</td><td>0.889</td>
245
- </tr>
246
- <tr>
247
- <td><b>granite-docling-258m</b></td>
248
- <td><b>0.013</b></td><td><b>0.988</b></td><td><b>0.99</b></td><td><b>0.988</b></td>
249
- <td><b>0.983</b></td><td><b>0.986</b></td>
250
- </tr>
251
- </tbody>
252
- <thead>
253
- <tr><th colspan="7"><b>Equation Recognition</b></th></tr>
254
- <tr>
255
- <th></th>
256
- <th>Edit-distance โ†“</th>
257
- <th>F1 โ†‘</th>
258
- <th>Precision โ†‘</th>
259
- <th>Recall โ†‘</th>
260
- <th>BLEU โ†‘</th>
261
- <th>Meteor โ†‘</th>
262
- </tr>
263
- </thead>
264
- <tbody>
265
- <tr>
266
- <td><b>smoldocling-256m-preview</b></td>
267
- <td>0.119</td><td>0.947</td><td>0.959</td><td>0.941</td><td>0.824</td><td>0.878</td>
268
- </tr>
269
- <tr>
270
- <td><b>granite-docling-258m</b></td>
271
- <td><b>0.073</b></td><td><b>0.968</b></td><td><b>0.968</b></td><td><b>0.969</b></td>
272
- <td><b>0.893</b></td><td><b>0.927</b></td>
273
- </tr>
274
- </tbody>
275
- </table>
276
- <table>
277
- <thead>
278
- <tr><th colspan="3"><b>Table Recognition (FinTabNet 150dpi)</b></th></tr>
279
- <tr>
280
- <th></th>
281
- <th>TEDS (structure) โ†‘</th>
282
- <th>TEDS (w/content) โ†‘</th>
283
- </tr>
284
- </thead>
285
- <tbody>
286
- <tr>
287
- <td><b>smoldocling-256m-preview</b></td>
288
- <td>0.82</td><td>0.76</td>
289
- </tr>
290
- <tr>
291
- <td><b>granite-docling-258m</b></td>
292
- <td><b>0.97</b></td><td><b>0.96</b></td>
293
- </tr>
294
- </tbody>
295
- </table>
296
- <table>
297
- <thead>
298
- <tr><th colspan="3"><b>Other Benchmarks</b></th></tr>
299
- <tr>
300
- <th></th>
301
- <th>MMStar โ†‘</th>
302
- <th>OCRBench โ†‘</th>
303
- </tr>
304
- </thead>
305
- <tbody>
306
- <tr>
307
- <td><b>smoldocling-256m-preview</b></td>
308
- <td>0.17</td><td>338</td>
309
- </tr>
310
- <tr>
311
- <td><b>granite-docling-258m</b></td>
312
- <td><b>0.30</b></td><td><b>500</b></td>
313
  </tr>
314
  </tbody>
315
  </table>
316
 
 
317
  ๐Ÿ’ป Local inference on Apple Silicon with MLX: [see here](https://huggingface.co/ibm-granite/granite-docling-258M-mlx)
318
 
319
  ## Supported Instructions
@@ -370,6 +294,8 @@ The evaluation can be performed using the [docling-eval](https://github.com/docl
370
 
371
  # Model Architecture:
372
 
 
 
373
  The architecture of granite-docling-258m consists of the following components:
374
 
375
  (1) RT-DETRS object detector: [docling-layout-heron](https://huggingface.co/docling-project/docling-layout-heron).
@@ -444,4 +370,3 @@ Its training, which includes both human-annotated and synthetic data informed by
444
 
445
  llm = LLM(model=MODEL_PATH, revision="untied", limit_mm_per_prompt={"image": 1}, dtype="float32")
446
  ```
447
- ````
 
130
 
131
  print(doc.export_to_markdown())
132
 
133
+ ```
134
  </details>
135
 
136
 
 
139
  <details>
140
  <summary>๐Ÿ“„ Single page image inference using plain ๐Ÿค— tranformers ๐Ÿค–</summary>
141
 
142
+ ```python
143
  # Run batch inference
144
  start_time = time.time()
145
  outputs = llm.generate(batched_inputs, sampling_params=sampling_params)
 
196
  </tr>
197
  <tr>
198
  <td><b>granite-docling-258m</b></td>
199
+ <td>0.27</td><td>0.86</td><td>0.92</td><td>0.88</td>
200
+ </tr>
201
+ <tr>
202
+ <td><b>granite-docling-2stage_258m</b></td>
203
+ <td><b>0.31</b></td><td><b>0.90</b></td><td><b>0.93</b></td><td><b>0.92</b></td>
204
  </tr>
205
  </tbody>
206
  </table>
 
226
  </tr>
227
  <tr>
228
  <td><b>granite-docling-258m</b></td>
229
+ <td>0.45</td><td>0.84</td><td>0.91</td>
230
+ <td><b>0.83</b></td><td>0.65</td><td>0.72</td>
231
  </tr>
 
 
 
232
  <tr>
233
+ <td><b>granite-docling-2stage_258m</b></td>
234
+ <td><b>0.27</b></td><td><b>0.85</b></td><td><b>0.92</b></td>
235
+ <td><b>0.83</b></td><td><b>0.70</b></td><td><b>0.79</b></td>
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
236
  </tr>
237
  </tbody>
238
  </table>
239
 
240
+
241
  ๐Ÿ’ป Local inference on Apple Silicon with MLX: [see here](https://huggingface.co/ibm-granite/granite-docling-258M-mlx)
242
 
243
  ## Supported Instructions
 
294
 
295
  # Model Architecture:
296
 
297
+ <img src="https://huggingface.co/docling-project/granite-docling-2stage-258m/resolve/main/architecture.png" alt="2 stage architecutre" style="width: 500px; height: auto; margin-right: 20px;">
298
+
299
  The architecture of granite-docling-258m consists of the following components:
300
 
301
  (1) RT-DETRS object detector: [docling-layout-heron](https://huggingface.co/docling-project/docling-layout-heron).
 
370
 
371
  llm = LLM(model=MODEL_PATH, revision="untied", limit_mm_per_prompt={"image": 1}, dtype="float32")
372
  ```
 
architecture.png ADDED

Git LFS Details

  • SHA256: aff2187b3ebd476756c298f74c53f0185df25c8a9c8041858a399dfc1f2f8d6b
  • Pointer size: 131 Bytes
  • Size of remote file: 289 kB