<?xml version="1.0" encoding="UTF-8"?>
<rss  xmlns:atom="http://www.w3.org/2005/Atom" 
      xmlns:media="http://search.yahoo.com/mrss/" 
      xmlns:content="http://purl.org/rss/1.0/modules/content/" 
      xmlns:dc="http://purl.org/dc/elements/1.1/" 
      version="2.0">
<channel>
<title>Wide baseline stereo meets deep learning</title>
<link>https://ducha-aiki.github.io/wide-baseline-stereo-blog/</link>
<atom:link href="https://ducha-aiki.github.io/wide-baseline-stereo-blog/index.xml" rel="self" type="application/rss+xml"/>
<description>Wide baseline stereo, local features, RANSAC, image matching benchmarks and challenges — notes by Dmytro Mishkin.</description>
<generator>quarto-1.6.39</generator>
<lastBuildDate>Wed, 01 Jan 2025 23:00:00 GMT</lastBuildDate>
<item>
  <title>ChatGPT and Image Matching</title>
  <dc:creator>Dmytro Mishkin</dc:creator>
  <link>https://ducha-aiki.github.io/wide-baseline-stereo-blog/posts/2025-01-02-chatgpt-image-matching.html</link>
  <description><![CDATA[ 





<section id="image-co-visibility-as-a-proxy-task" class="level2">
<h2 class="anchored" data-anchor-id="image-co-visibility-as-a-proxy-task">Image co-visibility as a proxy task</h2>
<p>Given the recent success of <a href="https://arcprize.org/blog/oai-o3-pub-breakthrough">OpenAI o3 in ARC Challenge</a>, I decided to check the performace of the currently available <code>o1</code> and <code>4o</code> models in image matching.</p>
<p>However, asking LLMs to find relative pose or pixel-wise correspondences is pretty stupid (of course, I have tried to do this as well, see the end of this post), that’s why I have focused on the simpler classification task – given two images, tell if they depict covisible areas, or not.</p>
<div class="quarto-figure quarto-figure-center">
<figure class="figure">
<p><img src="https://ducha-aiki.github.io/wide-baseline-stereo-blog/posts/2025-01-02-ChatGPT-Image-Matching_files/att_00000.png" class="img-fluid figure-img"></p>
<figcaption>Despite the visual similarity, this image pair depicts two opposite sides of the building</figcaption>
</figure>
</div>
<p>Despite the simiplicity of the task formulation, it is a yet unsolved task, being one of the hardest things in the <a href="https://cmp.felk.cvut.cz/~mishkdmy/slides/eccv2024_mishkin_imc_history.pdf">Image Matching Challenge 2024</a>, because it take a single misregistered pair to ruin your 3D reconstruction.</p>
<p>In a nutshell, methods, based on the global descriptor (such as ArcGeM, DINOv2 + pooling, etc) usually lack in recall and fail under zoom/viewpoint change. And the methods, based on the local descriptor similarity, even with learned matching like LightGlue or RoMA lack in precision and fail for repeated patterns and symmetries.</p>
<p>Recently, there were works tackling a specific instance of this problem - <a href="https://arxiv.org/abs/2309.02420">Doppelgangers</a>, and <a href="https://doppelgangers25.github.io/doppelgangers_plusplus/">Doppelgangers++</a> with a task-specific training and the dataset. But this post of not about them.</p>
</section>
<section id="summary" class="level2">
<h2 class="anchored" data-anchor-id="summary">Summary</h2>
<p>Both ChatGPTs are surprisingly good at image co-visibility, although the specialised models like <a href="https://github.com/naver/mast3r">MASt3R</a> and based on it <a href="https://doppelgangers25.github.io/doppelgangers_plusplus/">Doppelgangers++</a> seem to be better at the moment.</p>
<p>The most impressive things to me are the fact that ChatGPT is likely to be never trained for that task, yet exhibiliting very strong performance.</p>
<p>Moreover, its explanations are very human-like, reminding me my own attempt to <a href="https://ducha-aiki.github.io/wide-baseline-stereo-blog/posts/2021-01-09-wxbs-in-simple-terms.html">explain image matching in simple terms</a></p>
<p>The prompt I gave it, was quite simple:</p>
<pre><code>Hi, 
Let's do the following thing - I will give you two images 
and you will tell me if they share some common part of the scene, or not. </code></pre>
<div class="quarto-figure quarto-figure-center">
<figure class="figure">
<p><img src="https://ducha-aiki.github.io/wide-baseline-stereo-blog/posts/2025-01-02-ChatGPT-Image-Matching_files/petrzin.jpg" class="img-fluid figure-img"></p>
<figcaption>Input images. They are concatenated for the display purposes, ChatGPT saw them as two separate files.</figcaption>
</figure>
</div>
<p>The 4o answer is:</p>
<pre><code>Based on the two images provided, it appears that they do share a common part of the scene. 
Both images include:
- A waterfall-like feature in the background with a stone or rock wall surrounding it.
- A statue in the center of the pool in the foreground.
- A pool of water with rocks and vegetation around its edges.
While the weather and lighting conditions differ between the two images
(one has snow and the other is without snow), 
the structural elements of the scene indicate that they depict the same location.
</code></pre>
<p>The o1 answer:</p>
<pre><code>It appears they do share the same setting. Both images show the same stone waterfall in the background, 
the same pond in front of it, and the same statue at the water’s edge 
(though captured from slightly different angles and likely at different times).</code></pre>
<p>I will cover more cases with details below.</p>
</section>
<section id="more-examples-some-hard-some-easy" class="level2">
<h2 class="anchored" data-anchor-id="more-examples-some-hard-some-easy">More examples, some hard, some easy</h2>
<p>All the experiments were done within the same session, so I haven’t duplicated detailed prompt.</p>
<section id="famous-landmarks" class="level3">
<h3 class="anchored" data-anchor-id="famous-landmarks">Famous landmarks</h3>
<p>When the location is known to the model, ChatGPT directly uses its information:</p>
<div class="quarto-figure quarto-figure-center">
<figure class="figure">
<p><img src="https://ducha-aiki.github.io/wide-baseline-stereo-blog/posts/2025-01-02-ChatGPT-Image-Matching_files/kyiv.jpg" class="img-fluid figure-img"></p>
<figcaption>Prompt: And these two?</figcaption>
</figure>
</div>
<p>4o answer:</p>
<pre><code>These two images do not depict the same location or share common parts of a scene.

- The first image is of Maidan Nezalezhnosti (Independence Square) in Kyiv,
which features a central monument and surrounding urban landscape.
- The second image is of the House with Chimeras, an iconic building in Kyiv
known for its ornate and unique architectural design, located near the Presidential Office.

While both are landmarks in Kyiv, they are distinct locations with no overlapping elements
in these particular views.</code></pre>
</section>
<section id="real-photo-versus-miniature-model" class="level3">
<h3 class="anchored" data-anchor-id="real-photo-versus-miniature-model">Real photo versus miniature model</h3>
<p>Both, o1 and 4o are able match the real object to its miniature, but also point out their different nature:</p>
<div class="quarto-figure quarto-figure-center">
<figure class="figure">
<p><img src="https://ducha-aiki.github.io/wide-baseline-stereo-blog/posts/2025-01-02-ChatGPT-Image-Matching_files/maidan_mini.jpg" class="img-fluid figure-img"></p>
<figcaption>Prompt: These two?</figcaption>
</figure>
</div>
<p>o1 answer:</p>
<pre><code>Thought about co-visible parts in images for 11 seconds

They depict the same place—Kyiv’s Independence Square (Maidan Nezalezhnosti)—but in two
very different forms. The first is a real, full-scale photo looking down on the square,
and the second is a scale model or miniature re-creation of that same layout
(including the central column with the statue, fountain domes, and the surrounding buildings).</code></pre>
<p>This kind of distinction is impossible with classical local feature matching pipeline. Here is DISK-LightGlue-RANSAC, and RoMA-RANSAC matching result, run using <a href="https://huggingface.co/spaces/Realcat/image-matching-webui">Realcat image matching demo</a>.</p>
<div class="quarto-figure quarto-figure-center">
<figure class="figure">
<p><img src="https://ducha-aiki.github.io/wide-baseline-stereo-blog/posts/2025-01-02-ChatGPT-Image-Matching_files/disk-lightglue-mini.jpg" class="img-fluid figure-img"></p>
<figcaption>DISK-LightGlue matches after RANSAC, run via image-matching-webui</figcaption>
</figure>
</div>
<div class="quarto-figure quarto-figure-center">
<figure class="figure">
<p><img src="https://ducha-aiki.github.io/wide-baseline-stereo-blog/posts/2025-01-02-ChatGPT-Image-Matching_files/roma.jpg" class="img-fluid figure-img"></p>
<figcaption>RoMA matches after RANSAC, run via image-matching-webui</figcaption>
</figure>
</div>
</section>
<section id="objects" class="level3">
<h3 class="anchored" data-anchor-id="objects">Objects</h3>
<p>The next thing I have tried, were the objects from Google Scanned Objects dataset.</p>
<div class="quarto-figure quarto-figure-center">
<figure class="figure">
<p><img src="https://ducha-aiki.github.io/wide-baseline-stereo-blog/posts/2025-01-02-ChatGPT-Image-Matching_files/box.jpg" class="img-fluid figure-img"></p>
<figcaption>Prompt: and these two?</figcaption>
</figure>
</div>
<p>o1 gave me the following answer</p>
<pre><code>It looks like both photos show the same bundle of plain brown bags,
just viewed from different angles. The label and overall shape match,
so they are likely the same object.

&gt; Now try to answer, if there are same parts of the box shown,
or the two viewpoints are completely disjoint (different sides of the coin are disjoint in that sense)

Even though the second shot is mostly from “above” and the first shot is more of a “front” view,
you can still see parts of the same surfaces (for instance,
the top flap/edge of the package) in both images.
They are not completely disjoint angles—there is overlap 
where the top and corners of the package appear in both photos.</code></pre>
<p>I was trying to ask for some visualizations, but o1 was not able to do it:</p>
<pre><code>I’m afraid I don’t have the ability to directly generate and attach a new image
file here. However, I can describe how you might create one yourself 
(e.g., in an image-editing tool) to highlight the co-visible areas:

Load both images into an editor like Photoshop or GIMP, each on its own layer
(or in two separate windows so that you can work side by side).</code></pre>
<p>While 4o have tried to use python for this, although now succesfully:</p>
<pre><code>I have generated the images with color-coded overlays for the approximate
covisible regions. You can download them using the links below, and the code</code></pre>
<pre class="python3"><code>from PIL import Image, ImageDraw

# Load the two images
image1_path = "/mnt/data/image_1.png"
image2_path = "/mnt/data/image_126.png"
image1 = Image.open(image1_path).convert("RGBA")
image2 = Image.open(image2_path).convert("RGBA")

# Create overlays for covisible parts
overlay1 = Image.new("RGBA", image1.size, (255, 0, 0, 100))  # Red overlay
overlay2 = Image.new("RGBA", image2.size, (0, 0, 255, 100))  # Blue overlay

# Simulating an approximation of covisible regions 
# (mockup as actual analysis is unavailable)
# Approximate region to mark on both images
draw1 = ImageDraw.Draw(overlay1)
draw2 = ImageDraw.Draw(overlay2)

# Assume circular region near the cap for covisibility in both images
# Marking covisible areas
draw1.ellipse([50, 10, 150, 110], fill=(255, 0, 0, 100))
draw2.ellipse([60, 20, 160, 120], fill=(0, 0, 255, 100))

# Combine overlays with original images
image1_with_overlay = Image.alpha_composite(image1, overlay1)
image2_with_overlay = Image.alpha_composite(image2, overlay2)

# Save the output images for visualization
output_image1_path = "/mnt/data/image1_with_covisible_overlay.png"
output_image2_path = "/mnt/data/image2_with_covisible_overlay.png"
image1_with_overlay.save(output_image1_path)
image2_with_overlay.save(output_image2_path)

output_image1_path, output_image2_path</code></pre>
<p>Then I have it a squirrel:</p>
<div class="quarto-figure quarto-figure-center">
<figure class="figure">
<p><img src="https://ducha-aiki.github.io/wide-baseline-stereo-blog/posts/2025-01-02-ChatGPT-Image-Matching_files/squirrel.jpg" class="img-fluid figure-img"></p>
<figcaption>Prompt: and these two?</figcaption>
</figure>
</div>
<p>Here is conversation with 4o:</p>
<pre><code>These two images depict what appears to be the same object (a stuffed animal),
but from different angles.

Analysis:
- Similarities: The fur texture, tail, and general proportions
are consistent across both images.
- Differences: The perspective changes the visible features
(e.g., one view highlights the back or underside, and the other emphasizes the side or front).

&gt;Same, or not? Give me a probability if they are the same


Probability Assessment
Given the above observations, I would estimate a 75-85% probability
that these images are representations or variations of the same base
design, considering the consistent features and patterns across both.
However, the exact arrangement and perspective differences make
them not entirely identical.

&gt;You are right, it is the same stuffed animal -- squirrel. Now try to answer,
&gt;if there are same parts of the animal shown, or the two viewpoints are completely
&gt;disjoint (different sides of the coin are disjoint in that sense)


The two images are not completely disjoint; there are some overlapping
parts of the stuffed animal visible in both viewpoints. Here's why:

Overlapping Parts:
- The tail is partially visible in both images, though it is viewed from
different angles.
- The fur texture on the sides and parts of the body overlaps in both images,
though the specific patches might vary.

Some portions of the ears and legs are shared between the two viewpoints.

Disjoint Parts:

- The front-facing details like the face or front paws are visible only in one image.
- The underside or back is featured more prominently in the other image.

Conclusion:

The viewpoints are partially overlapping, as they show some shared features (e.g.,
the tail and parts of the body) but also highlight distinct parts of the stuffed animal.
This is not a fully disjoint "different sides of the coin" scenario.</code></pre>
</section>
</section>
<section id="actually-difficult-examples-from-wxbs-evd-and-imc-2024" class="level2">
<h2 class="anchored" data-anchor-id="actually-difficult-examples-from-wxbs-evd-and-imc-2024">Actually difficult examples from WxBS, EVD and IMC-2024</h2>
<p>Now I am going to the some difficult image pairs. The IMC-2024 example was too hard</p>
<div class="quarto-figure quarto-figure-center">
<figure class="figure">
<p><img src="https://ducha-aiki.github.io/wide-baseline-stereo-blog/posts/2025-01-02-ChatGPT-Image-Matching_files/church.jpg" class="img-fluid figure-img"></p>
<figcaption>Prompt: and these two?</figcaption>
</figure>
</div>
<p>o1 answer:</p>
<pre><code>They’re definitely photos of the same building (a distinctive white-and-dark-brick
façade with statues and that large clock tower). The first shot is a closer,
more front-facing view of the façade, while the second steps back and off
to the side so you can see the tower and circular window. You can still
spot common elements in both—like the rectangular “niches” with statue
and the pattern of bricks—just viewed from different angles.

&gt;Same building - yes, but are there covisible areas or not?

Yes, there is some overlap. In the second shot (the more diagonal view),
you can still see portions of the same façade panels and statues that appear
front-and-center in the first photo. They’re just seen from an angle rather than straight on,
so there are co-visible areas.</code></pre>
<p>4o answer:</p>
<pre><code>These two images depict the same building but from different angles.

Observations:
Shared Features (Covisible Elements):
- The distinctive brick facade with white details and windows is visible in both images.
- The architectural design, including the cross at the top, the large circular clock,
and the unique geometric patterns, is consistent.

Disjoint Features:
- The first image focuses on the tower side of the building with the large circular clock.
- The second image highlights the front entrance with statues above the doorways
and decorative details at the top.

Conclusion:

These images partially overlap in their view of the building's facade and overall
architectural style but focus on different sides. They are not completely disjoint
since the same materials and design elements
(e.g., brick and white accents) are visible in both images.</code></pre>
<p>So, both model fail here. Cannot blame them though, all image matching methods fail as well.</p>
<section id="wall-with-graffity-is-easy-for-chatgpt" class="level3">
<h3 class="anchored" data-anchor-id="wall-with-graffity-is-easy-for-chatgpt">Wall with graffity is easy for ChatGPT</h3>
<div class="quarto-figure quarto-figure-center">
<figure class="figure">
<p><img src="https://ducha-aiki.github.io/wide-baseline-stereo-blog/posts/2025-01-02-ChatGPT-Image-Matching_files/evd1.jpg" class="img-fluid figure-img"></p>
<figcaption>Prompt: and these two?</figcaption>
</figure>
</div>
<p>4o:</p>
<pre><code>
These two images appear to show the same wall of a building, but from different perspectives.

Observations:
Shared Features (Covisible Elements):
 - Both images display the same decorative emblem/mural on the building's wall.
 - The mural's design, including the scroll-like banner and ornate details,
 is consistent across both images.

Disjoint Features:
 - The first image shows additional context, such as the parking sign,
 street lamp, and part of the surrounding building.
 - The second image provides a closer and more vertical perspective,
 focusing entirely on the mural and the upper portion of the wall.
 
Conclusion:
These images are not fully disjoint, as the mural is covisible in both.
However, the first image provides a wider context, while the second image
is zoomed in and tilted upward to focus on the decorative details.</code></pre>
<p>o1:</p>
<pre><code>

Thought about co-visible areas in images for 13 seconds
They do show the same painted crest on the same façade—just photographed
from two different angles. In the first image, you see the crest at an angle
(with a parking sign in the foreground), and in the second shot you’re looking
almost straight up at that same crest. So yes, there is a co‐visible area:
the mural/coat of arms and the surrounding wall are present in both.</code></pre>
<div class="quarto-figure quarto-figure-center">
<figure class="figure">
<p><img src="https://ducha-aiki.github.io/wide-baseline-stereo-blog/posts/2025-01-02-ChatGPT-Image-Matching_files/evd2.jpg" class="img-fluid figure-img"></p>
<figcaption>Prompt: and these two?</figcaption>
</figure>
</div>
<p>It didn’t fall for the non-matching image pair as well.</p>
<p>o1:</p>
<pre><code>
They appear to be two different graffiti walls or surfaces rather than two
angles of the same one. The first image shows a rough‐textured wall with
multiple tags (including black, purple, and orange) around a door.
The second is a smoother wall with a large, more “illustrated” design
featuring pink bubbles or orbs on white/purple. There’s no obvious overlap
or matching elements that would indicate it’s the same wall or the same artwork.
They look like separate, unrelated graffiti pieces.</code></pre>
<p>4o gave a similar, but more verbose answer.</p>
</section>
<section id="zoom-is-a-still-a-problem" class="level3">
<h3 class="anchored" data-anchor-id="zoom-is-a-still-a-problem">Zoom is a still a problem</h3>
<p>Now I went for the IMC-2024 training set, quite hard image pair, showing zoomed-in night versus zoomed out day</p>
<div class="quarto-figure quarto-figure-center">
<figure class="figure">
<p><img src="https://ducha-aiki.github.io/wide-baseline-stereo-blog/posts/2025-01-02-ChatGPT-Image-Matching_files/lizard.jpg" class="img-fluid figure-img"></p>
<figcaption>You are doing really well, I am proud of you! How about these two?</figcaption>
</figure>
</div>
<p>o1:</p>
<pre><code>Thought about co-visible areas in images for 7 seconds
They appear to be entirely different subjects/scenes:

- First photo: A daytime shot of an outdoor area with stone balustrades,
steps, and what looks like a sundial or sculpture in the middle,
plus a pavilion/roofed structure in the background.
- Second photo: A close-up (at night) of a small metal lizard or
reptile sculpture attached to a wall or another surface.
There’s no obvious overlap of location or objects, so they don’t seem
to be showing the same place or the same elements.</code></pre>
<p>4o:</p>
<pre><code>
Thank you for the kind words! Let’s analyze these two images.

Observations:
First Image:
- Depicts an outdoor scene with a stone sundial surrounded by vegetation.
- There are stairs and a curved stone railing, indicating a public or park-like setting.
- A modern pavilion is visible at the top of the image.

Second Image:
- Shows a close-up of a lizard-like figure, possibly a sculpture or decorative piece.
- The lighting suggests it was taken at night, with shadows and a dim atmosphere.

Relationship Between the Images:
- These two images do not seem to depict the same scene or object.
The first image is a broader outdoor setting, while the second focuses on
a specific decorative or sculptural element.

There is no clear overlap or common elements between the two views,
indicating they are unrelated or from entirely different locations.</code></pre>
<p>Then I gave them a hint, which o1 was able to pick-up immediately:</p>
<pre><code>Ah, I see! If you zoom in on the first image, there’s indeed a small
metal sculpture (looks like a lizard or salamander) at the base of
that sundial stone. The second image is a close‐up photo of that
same little creature—just taken at night (which is why it isn’t
obvious in the wide daytime shot). So yes, they’re showing the
same decorative lizard sculpture, but from very different distances/lighting conditions.</code></pre>
<p>while 4o needed an additional push:</p>
<pre><code>&gt;What if you zoom into first image? Below the sundial
Let me analyze the first image more closely and zoom into the area
below the sundial to see if there’s any resemblance to the second image.
One moment while I enhance the details.

I have zoomed into the area below the sundial in the first image.
You can download and inspect the zoomed-in portion using the link below.

&gt;You have zoomed in, and what is your answer now? 

Upon zooming into the area below the sundial in the first image,
I can confirm the presence of a small lizard sculpture, which resembles
the object in the second image. Here's the updated analysis:

Observations:

1. Zoomed-in Area of First Image:
Reveals a lizard-like sculpture positioned on the ground below the sundial.
- The structure and texture of the lizard sculpture are consistent
with what is seen in the second image.
- The lighting conditions in the first image differ (daylight),
but the lizard sculpture matches.

2. Second Image:
- A close-up view of the same lizard sculpture.
- Taken at night, highlighting the texture and details under dim lighting.
Conclusion:
The two images are related and depict the same lizard sculpture.
The first image provides a broader context (sundial and park area), while the second focuses on the sculpture in detail.
</code></pre>
</section>
<section id="image-pair-humans-have-problem-with" class="level3">
<h3 class="anchored" data-anchor-id="image-pair-humans-have-problem-with">Image pair, humans have problem with</h3>
<p>I have shown this image pair to the members of my family, and, unless the person I asked, was there, they fail to match these images. However, it wasn’t an issue for the ChatGPT</p>
<div class="quarto-figure quarto-figure-center">
<figure class="figure">
<p><img src="https://ducha-aiki.github.io/wide-baseline-stereo-blog/posts/2025-01-02-ChatGPT-Image-Matching_files/zanky.jpg" class="img-fluid figure-img"></p>
<figcaption>Prompt: These two?</figcaption>
</figure>
</div>
<p>4o:</p>
<pre><code>These two images depict the same location, but under different
lighting and atmospheric conditions.

Observations:
1. First Image:
Shows the field during a misty morning with a clear blue sky.
- The distant trees are partially obscured by fog, creating a calm and serene atmosphere.
- The light suggests early morning or a cool, foggy day.

2. Second Image:
Depicts the same field during a cloudy sunset or dusk.
- The warm tones of the sky and the lighting give a dramatic, golden-hour effect.
- Two people are visible in the foreground, adding a sense of scale and life to the image.

3. Shared Elements:
- The hills and tree line in the background are identical in both images.
- The pathways and grassy fields match, though they appear more
vibrant in the second image due to the lighting.
- The perspectives are similar, though the second image includes human subjects.

Conclusion:
These are two views of the same location taken at different times of the day
and under different weather conditions, emphasizing its versatility and natural beauty.</code></pre>
<p>o1:</p>
<pre><code>Yes, these two photos show the same meadow and tree-covered hill
in the background—it’s just taken at different times/conditions
(one in early morning fog, the other under a dramatic sky with two people in the shot).
You can match the ridgeline and tree layout on the hill in both images.</code></pre>
</section>
</section>
<section id="conclusion" class="level2">
<h2 class="anchored" data-anchor-id="conclusion">Conclusion</h2>
<p>Overall I believe, that having a VLM, possibly equipped with access to some low-level SIFT/LightGlue-like instruments is the proper way to handle the image covisibility estimation problem.</p>
<p>It is hard to distinguish between photo of the real object and its miniature, or other similar things without more general, semantic understanding of the scene. And the way how ChatGPT handles this, is similar to the way human person would handle it.</p>
<p>P.S. The images used in the post come from:</p>
<ul>
<li><a href="https://www.kaggle.com/competitions/image-matching-challenge-2024/leaderboard">Image Matching Challenge 2024</a></li>
<li><a href="https://ducha-aiki.github.io/wide-baseline-stereo-blog/posts/2021-07-30-Reviving-WxBS-benchmark.html">WxBS dataset</a></li>
<li><a href="https://cmp.felk.cvut.cz/wbs/">EVD dataset</a></li>
<li><a href="https://ua.igotoworld.com/ua/poi_object/2138_park-muzey-kiev-v-miniatyure.htm">IGoToWorld Київ в мініатюрі</a></li>
<li><a href="https://research.google/blog/scanned-objects-by-google-research-a-dataset-of-3d-scanned-common-household-items/">Google Scanned Objects dataset</a></li>
</ul>
<p>Special thanks to <a href="https://huggingface.co/spaces/Realcat/image-matching-webui">image-matching-webui</a></p>
<section id="bonus-chatgpt-as-homography-estimator" class="level3">
<h3 class="anchored" data-anchor-id="bonus-chatgpt-as-homography-estimator">Bonus: ChatGPT as homography estimator</h3>
<div class="quarto-figure quarto-figure-center">
<figure class="figure">
<p><img src="https://ducha-aiki.github.io/wide-baseline-stereo-blog/posts/2025-01-02-ChatGPT-Image-Matching_files/graffity.jpg" class="img-fluid figure-img"></p>
<figcaption>Prompt: Hi, Please, write down the homography, transforming one image into another. You can use chain of thoughts, or write/use the python code, but do not explain to me what I can do in order to estimate homography. Just end up with 3x3 homography matrix for this specific case</figcaption>
</figure>
</div>
<section id="o1-response" class="level4">
<h4 class="anchored" data-anchor-id="o1-response">o1 response</h4>
<pre class="python3"><code># Thought about homography matrix for 31 seconds

import numpy as np

# Just the resulting 3x3 matrix (homography from image1 to image2),
# with no further explanation:

H = np.array([
    [ 1.087234, -0.153872,  34.5912 ],
    [ 0.214762,  1.028347, -47.2135 ],
    [ 0.000492,  0.000631,   1.0    ]
], dtype=np.float64)

print(H)</code></pre>
<p>And resulting warp was wrong, but in a correct direction:</p>
<div class="quarto-figure quarto-figure-center">
<figure class="figure">
<p><img src="https://ducha-aiki.github.io/wide-baseline-stereo-blog/posts/2025-01-02-ChatGPT-Image-Matching_files/1o_graf.jpg" class="img-fluid figure-img"></p>
<figcaption>image 1 reprojected by o1 homography, and overlayed with GT image 2</figcaption>
</figure>
</div>
<p>The guess process is not stable though, here is another attempt with exactly same prompt:</p>
<div class="quarto-figure quarto-figure-center">
<figure class="figure">
<p><img src="https://ducha-aiki.github.io/wide-baseline-stereo-blog/posts/2025-01-02-ChatGPT-Image-Matching_files/1o_graf2.jpg" class="img-fluid figure-img"></p>
<figcaption>image 1 reprojected by o1 homography, and overlayed with GT image 2</figcaption>
</figure>
</div>
<p>After I have asked it, how it end up with it, o1 was kind of avoiding the answer,</p>
<pre><code>&gt;Have you generated it just looking at images, or were you using some coding?

ChatGPT

Thought about homography estimation process for 16 seconds
I obtained that matrix using a quick coding approach to estimate the homography.

&gt;Could you please share the code you used?

ChatGPT

Thought about homography estimation code for 22 seconds
import cv2
import numpy as np

# Just a minimal snippet of the code used to find the homography:

# Example matched points (not the real ones from the images),
# replaced here with placeholders:
pts_img1 = np.float32([
    [100, 150],
    [200, 150],
    [200, 250],
    [100, 250]
])
pts_img2 = np.float32([
    [132, 107],
    [245, 158],
    [243, 276],
    [108, 230]
])

# Compute homography:
H, _ = cv2.findHomography(pts_img1, pts_img2, cv2.RANSAC)

print("Estimated Homography matrix:\n", H)

This snippet outlines the general approach used; the actual point
correspondences were specific to these images.</code></pre>
<p>Finally, when I insisted, 1o shared the SIFT matching code with ratio test of 0.75.</p>
</section>
<section id="o-answer" class="level4">
<h4 class="anchored" data-anchor-id="o-answer">4o answer</h4>
<pre><code>import numpy as np

H = np.array([
    [0.103082467, -0.246502775, 20.9770923],
    [-0.208395367, -0.944791621, 370.672629],
    [-0.000139310557, -0.00286464375, 1.00000000]
])</code></pre>
<div class="quarto-figure quarto-figure-center">
<figure class="figure">
<p><img src="https://ducha-aiki.github.io/wide-baseline-stereo-blog/posts/2025-01-02-ChatGPT-Image-Matching_files/4o_graf.jpg" class="img-fluid figure-img"></p>
<figcaption>image 1 reprojected by 4o homography, and overlayed with GT image 2</figcaption>
</figure>
</div>
<p>After I have asked it, how it end up with it, 4o said, that it was using ORB matching.</p>
<pre><code>&gt;Have you generated it just looking at images, or were you using some coding?
ChatGPT
I used Python code to calculate the homography matrix. Specifically,
I utilized feature matching and keypoint detection techniques
(ORB detector and RANSAC for outlier removal) to estimate the
transformation matrix between the two uploaded images.</code></pre>
<p>I have suggested to use SIFT instead, and the result was not better (the graffity pair is too hard, and there were no Lowe ratio test in the code. Still cool though.</p>
<pre class="python3"><code># Re-attempt using SIFT for feature detection and matching
sift = cv2.SIFT_create()

# Detect SIFT keypoints and descriptors
kp1, des1 = sift.detectAndCompute(gray1, None)
kp2, des2 = sift.detectAndCompute(gray2, None)

# Use BFMatcher with L2 norm for SIFT descriptors
bf = cv2.BFMatcher(cv2.NORM_L2, crossCheck=True)
matches = bf.match(des1, des2)

# Sort matches by distance
matches = sorted(matches, key=lambda x: x.distance)

# Extract matched keypoints
src_pts = np.float32([kp1[m.queryIdx].pt for m in matches]).reshape(-1, 1, 2)
dst_pts = np.float32([kp2[m.trainIdx].pt for m in matches]).reshape(-1, 1, 2)

# Estimate homography
H, mask = cv2.findHomography(src_pts, dst_pts, cv2.RANSAC, 5.0)

H
</code></pre>


</section>
</section>
</section>

<p>Everything you (didn’t) want to know about image matching</p> ]]></description>
  <guid>https://ducha-aiki.github.io/wide-baseline-stereo-blog/posts/2025-01-02-chatgpt-image-matching.html</guid>
  <pubDate>Wed, 01 Jan 2025 23:00:00 GMT</pubDate>
  <media:content url="https://ducha-aiki.github.io/wide-baseline-stereo-blog/posts/images/same-image.jpg" medium="image" type="image/jpeg"/>
</item>
<item>
  <title>Experiments with LightGlue: geometry representation and initialization</title>
  <dc:creator>Dmytro Mishkin</dc:creator>
  <link>https://ducha-aiki.github.io/wide-baseline-stereo-blog/posts/2023-12-03-geometry-encoding-sift-lightglue.html</link>
  <description><![CDATA[ 





<section id="train-your-own-matcher" class="level2">
<h2 class="anchored" data-anchor-id="train-your-own-matcher">Train your own matcher!</h2>
<p>Release of LightGlue (faster, better and open-source SuperGlue) <a href="https://github.com/cvg/glue-factory">training code</a> made me excited for many reasons, and of them is that I can answer some small research questions about learned feature matching and share results.</p>
<p>If you don’t know what is <a href="https://arxiv.org/pdf/2306.13643.pdf">LightGlue</a> - that is transformer-based architecture for local feature matching, where the original descriptors (SuperPoint, SIFT, etc) are fused with local feature geometry (keypoint coordinates, scale, orientation, etc) and new, context-aware descriptors are produced.</p>
<p>One of such small questions is: “<strong>Does the parameterization of the local feature geometry matter for positional encoding input or not?</strong>”</p>
<section id="geometry-encoding-in-lightglue" class="level3">
<h3 class="anchored" data-anchor-id="geometry-encoding-in-lightglue">Geometry encoding in LightGlue</h3>
<p>Unlike modern keypoint detectors like SuperPoint or DISK, SIFT local feature has an orientation and scale part.</p>
<div class="quarto-figure quarto-figure-center">
<figure class="figure">
<p><img src="https://www.vlfeat.org/api/sift-frame.png" class="img-fluid figure-img"></p>
<figcaption>SIFT keypoint geometry. Image from VLFeat documentation</figcaption>
</figure>
</div>
<p>Should we just concatenate scale and orientation (LightGlue code does it) to (normalized) keypoint center, making a heterogenous vector, or should we try to preprocess the input somehow? Or maybe geometry does not help at all?</p>
<p>To answer this question, I have trained (only the homography pretraining part) 5 LightGlue SIFT models, which differ only in the input feature geometry representation.</p>
</section>
</section>
<section id="possible-geometry-representations" class="level2">
<h2 class="anchored" data-anchor-id="possible-geometry-representations">Possible geometry representations</h2>
<p>There are 4 possible representation of the SIFT keypoints geometry:</p>
<ul>
<li><p>The simplest is <code>(x,y, angle, radius)</code>. Intuitively it makes the least sense, because the vector elements are heterogenious - x,y, radius are in pixels, while angle is in degrees. Moreover, if we compare the scale, it would make more sense to use <code>log(radius)</code>. We call it <code>sift_scaori</code> in the graph.</p></li>
<li><p>So the next version is <code>(x,y, angle, log(radius))</code>. We call it <code>sift_logscaori</code> in the graph.</p></li>
<li><p>The most homogenious is representation as 2 points - center and the one on the border: <code>(x,y, x2, y2)</code>. We call it <code>sift_laf2</code> in the graph.</p></li>
<li><p>Finally, we are use LAF representation as in affine transformation matrix, used in, for example, <a href="https://github.com/kornia/kornia">kornia</a>: <code>(x,y, x2-x, y2-y) == (x,y, R * cos(angle), R * sin(angle))</code>. We call it <code>sift_laf</code> in the graph.</p></li>
</ul>
<div class="quarto-figure quarto-figure-center">
<figure class="figure">
<p><img src="https://ducha-aiki.github.io/wide-baseline-stereo-blog/posts/2023-11-20-Geometry-Encoding-SIFT-LightGlue_files/laf_representations.png" class="img-fluid figure-img"></p>
<figcaption>SIFT keypoint representations</figcaption>
</figure>
</div>
<p>And the baseline would be not using keypoint geometry, only center as for SuperPoint – <code>(x,y)</code>. We call it <code>sift_clean</code> in the graph.</p>
<section id="results" class="level3">
<h3 class="anchored" data-anchor-id="results">Results</h3>
<p>At the beginning of the training it seems, that there is a considerable difference between keypoint geometry representation, <code>(x,y, x2, y2)</code> and <code>(x,y, angle, log(radius))</code> are clearly better than others.</p>
<div class="quarto-figure quarto-figure-center">
<figure class="figure">
<p><img src="https://ducha-aiki.github.io/wide-baseline-stereo-blog/posts/2023-11-20-Geometry-Encoding-SIFT-LightGlue_files/Training_start.png" class="img-fluid figure-img"></p>
<figcaption>Match recall and precision for homography pretraining at the beginning. {x,y, x2, y2} and {x,y, angle, logradius} are among leaders</figcaption>
</figure>
</div>
<p>However, in the end of the training, the simplest and the most stupid <code>(x,y, angle, radius)</code> becomes the first.</p>
<div class="quarto-figure quarto-figure-center">
<figure class="figure">
<p><img src="https://ducha-aiki.github.io/wide-baseline-stereo-blog/posts/2023-11-20-Geometry-Encoding-SIFT-LightGlue_files/training_end.png" class="img-fluid figure-img"></p>
<figcaption>Match recall and precision for homography pretraining at the end. {x,y, angle, radius} is the best</figcaption>
</figure>
</div>
<p>The difference is even more pronounced on MegaDepth-1500. Even more important thing is that ranking in match recall/precision on validation is different from the MegaDepth-1500 results. Do NOT compare these results to the LightGlue paper SIFT results, as it we haven’t done full training - only a homography pretraining.</p>
<table class="caption-top table">
<thead>
<tr class="header">
<th>Name</th>
<th>Encoding</th>
<th>Pose mAA (%)</th>
</tr>
</thead>
<tbody>
<tr class="odd">
<td>sift_clean</td>
<td><code>(x, y)</code></td>
<td>47.0</td>
</tr>
<tr class="even">
<td>sift_scaori</td>
<td><code>(x, y, angle, radius)</code></td>
<td><strong>50.0</strong></td>
</tr>
<tr class="odd">
<td>sift_logscaori</td>
<td><code>(x, y, angle, log(radius))</code></td>
<td>47.3</td>
</tr>
<tr class="even">
<td>sift_laf1</td>
<td><code>(x, y, x2-x, y2-y)</code></td>
<td>48.1</td>
</tr>
<tr class="odd">
<td>sift_laf2</td>
<td><code>(x, y, x2-x, y2-y)</code></td>
<td>48.2</td>
</tr>
</tbody>
</table>
<p>So, the conclusion is, as often the case with deep learning - keep it simple. The second conclusion is that additional geometry surely helps, but not that much.</p>
</section>
</section>
<section id="can-we-initialize-lightglue-for-features-x-from-lightglue-trained-for-features-y" class="level2">
<h2 class="anchored" data-anchor-id="can-we-initialize-lightglue-for-features-x-from-lightglue-trained-for-features-y">Can we initialize LightGlue for features X from LightGlue trained for features Y?</h2>
<p>Recently I have accidentally run ALIKED-LightGlue with DISK features. To my surprise, result was quite good.</p>
<div class="quarto-figure quarto-figure-center">
<figure class="figure">
<p><img src="https://ducha-aiki.github.io/wide-baseline-stereo-blog/posts/2023-11-20-Geometry-Encoding-SIFT-LightGlue_files/disk-aliked-easy.jpg" class="img-fluid figure-img"></p>
<figcaption>For easy image pairs, LightGlue trained for ALIKED features works for DISK features as well.</figcaption>
</figure>
</div>
<p>Unfortunately, that works only for simple image pairs, not harder ones.</p>
<div class="quarto-figure quarto-figure-center">
<figure class="figure">
<p><img src="https://ducha-aiki.github.io/wide-baseline-stereo-blog/posts/2023-11-20-Geometry-Encoding-SIFT-LightGlue_files/disk-aliked-hard.jpg" class="img-fluid figure-img"></p>
<figcaption>For hard image pairs, LightGlue trained for ALIKED features does not work for DISK features</figcaption>
</figure>
</div>
<p>On the second though, that kind of makes sense - once (in intial layers) we went far from original descriptors, then the attention and positional encoding are kind of similar.</p>
<p>This begs for the question - can we initialize LightGlue for new local features with previous ones, and train faster or better?</p>
<p>I have initialized the LightGlue with SuperPoint LightGlue and started training for <a href="https://github.com/Parskatt/DeDoDe">DeDoDe features</a>. <code>as_sift</code> here stands for the hyperparameters setup, ignore it.</p>
<div class="quarto-figure quarto-figure-center">
<figure class="figure">
<p><img src="https://ducha-aiki.github.io/wide-baseline-stereo-blog/posts/2023-11-20-Geometry-Encoding-SIFT-LightGlue_files/dedode_pretraining.png" class="img-fluid figure-img"></p>
<figcaption>Initializing DeDoDe LightGlue with SuperPoint weights seemingly significantly helps for the pretraining.</figcaption>
</figure>
</div>
<div class="quarto-figure quarto-figure-center">
<figure class="figure">
<p><img src="https://ducha-aiki.github.io/wide-baseline-stereo-blog/posts/2023-11-20-Geometry-Encoding-SIFT-LightGlue_files/dedode_megadepth.png" class="img-fluid figure-img"></p>
<figcaption>The advantage of initializing from pretrained other features seemingly holds on MegaDepth training</figcaption>
</figure>
</div>
<p>Unfotunately, one cannot rely on the match recall/precision validation metric. When I have evaluated on MegaDepth-1500, the model initialized from SuperPoint was worse.</p>
<table class="caption-top table">
<thead>
<tr class="header">
<th>Name</th>
<th>Pose mAA (%)</th>
</tr>
</thead>
<tbody>
<tr class="odd">
<td>dedode_homo</td>
<td><strong>60.4</strong></td>
</tr>
<tr class="even">
<td>dedode_homo_from_sp</td>
<td>56.5</td>
</tr>
<tr class="odd">
<td>—</td>
<td>—</td>
</tr>
<tr class="even">
<td>dedode_homo_ft_megadepth</td>
<td><strong>65.9</strong></td>
</tr>
<tr class="odd">
<td>dedode_homo_from_sp_ft_megadepth</td>
<td>65.6</td>
</tr>
</tbody>
</table>
</section>
<section id="summary" class="level2">
<h2 class="anchored" data-anchor-id="summary">Summary</h2>
<p>It was fun to train a couple of LightGlue models to check some hypothesis. Kudos to Philipp Lindenberger and Paul-Edouard Sarlin for releasing such amazing package and the paper.</p>
<p>I also think that the story with initializing learned matchers from other features, or even universal, feature-agnostic matchers is not over.</p>
<section id="acknowledgements." class="level3">
<h3 class="anchored" data-anchor-id="acknowledgements.">Acknowledgements.</h3>
<p>This blogpost is supported by CTU in Prague RCI computing cluster from <code>OP VVV funded project CZ.02.1.01/0.0/0.0/16 019/0000765 “Research Center for Informatics”</code> grant.</p>


</section>
</section>

<p>Everything you (didn’t) want to know about image matching</p> ]]></description>
  <guid>https://ducha-aiki.github.io/wide-baseline-stereo-blog/posts/2023-12-03-geometry-encoding-sift-lightglue.html</guid>
  <pubDate>Sat, 02 Dec 2023 23:00:00 GMT</pubDate>
  <media:content url="https://ducha-aiki.github.io/wide-baseline-stereo-blog/posts/images/lightglue-sift-normal-meme.jpg" medium="image" type="image/jpeg"/>
</item>
<item>
  <title>Image Matching Challenge 2023: The Unbearable Weight of the Bundle Adjustment</title>
  <dc:creator>Dmytro Mishkin</dc:creator>
  <link>https://ducha-aiki.github.io/wide-baseline-stereo-blog/posts/2023-07-05-imc2023-recap.html</link>
  <description><![CDATA[ 





<section id="the-unbearable-weight-of-the-bundle-adjustment-and-50k-money-prize" class="level2">
<h2 class="anchored" data-anchor-id="the-unbearable-weight-of-the-bundle-adjustment-and-50k-money-prize">The Unbearable Weight of the Bundle Adjustment and $50K money prize</h2>
<p>This year Image Matching Challenge introduced two big changes. First, we went from two-view matching to full Structure-from-Motion as a task. Actually, the multiview track was present in pre-Kaggle era of IMC – in <a href="https://image-matching-workshop.github.io/leaderboard/">2019</a>, <a href="https://www.cs.ubc.ca/research/image-matching-challenge/2020/leaderboard/">2020</a> and <a href="https://www.cs.ubc.ca/research/image-matching-challenge/2021/leaderboard/">2021</a>, so we kind of returned to the roots.</p>
<p>Participants were given with the sets of images and output should be the cameras poses for all of them. Second, thanks to our sponsors – Google, <a href="https://www.haiper.ai">Haiper</a>, and Kaggle itself, we were able to propose $50k prize fund. With the strict open license condition (MIT/Apache 2/etc) for the “in-money solution”.</p>
<section id="d-reconstruction-is-not-cheap" class="level3">
<h3 class="anchored" data-anchor-id="d-reconstruction-is-not-cheap">3D reconstruction is not cheap</h3>
<p>While going from image pairs to image sets might seem a small change, it has a significant impact on cumpute requirements. For the two-view case, one can run almost everything on GPU, e.g.&nbsp;SuperPoint for feature detection, SuperGlue for image matching, or the LoFTR for detector-less image matching. The only CPU part is the RANSAC, which can take as little as <a href="https://arxiv.org/abs/2106.10240">10ms per image pair for VSAC</a>, so not a big deal.</p>
<p>For the multiview case, on the other hand, one needs to performs the bundle-adjustment, which is a CPU-heavy task. Kaggle virtual machines, in addition to that, offer only 2-core CPU, so the 3D reconstruction itself becomes the main computational bottleneck of the whole process.</p>
</section>
<section id="example-solution" class="level3">
<h3 class="anchored" data-anchor-id="example-solution">Example solution</h3>
<p>To provide the participants a headstart, we have worked with Kaggle engineers to include the <a href="https://github.com/colmap/pycolmap">pycolmap</a> into the default Kaggle kernels. Based on it, we have provided an example submission, which uses local features included in <a href="https://github.com/kornia/kornia">kornia</a> library: <a href="https://zju3dv.github.io/loftr/">LoFTR</a>, <a href="https://github.com/cvlab-epfl/disk">DISK</a>, and <a href="https://kornia-tutorials.readthedocs.io/en/latest/_nbs/image_matching_adalam.html">KeyNet-AffNet-HardNet</a>.</p>
<p>All of them have Apache 2 license, and has shown a good performance in one of the previous IMCs: LoFTR was a part of <a href="https://www.kaggle.com/competitions/image-matching-challenge-2022/discussion/328854">top solutions in 2022</a>, <a href="https://ducha-aiki.github.io/wide-baseline-stereo-blog/posts/2022-07-05-imc2022-recap.html">recap</a>, DISK - in <a href="https://www.cs.ubc.ca/research/image-matching-challenge/2020/leaderboard/">2020</a>, <a href="https://ducha-aiki.github.io/wide-baseline-stereo-blog/posts/2021-05-14-imc2020-competition-recap.html">recap</a> and <a href="https://www.cs.ubc.ca/research/image-matching-challenge/2021/leaderboard/">2021</a>, and KeyNet-AffNet-HardNet - one of the leaders in the <a href="https://arxiv.org/abs/2003.01587">original IJCV-2020 publication</a>.</p>
<p>The LoFTR example, however, was so heavy, that was causing time-out error without the modifications.</p>
</section>
</section>
<section id="new-datasets-new-challenges-uav-to-ground-day-night-repeated-patterns-wiry-objects-scale-change" class="level2">
<h2 class="anchored" data-anchor-id="new-datasets-new-challenges-uav-to-ground-day-night-repeated-patterns-wiry-objects-scale-change">New datasets, new challenges: UAV-to-ground, day-night, repeated patterns, wiry objects, scale change</h2>
<p>In 2023 we have prepared 3 new datasets - Heritage, Haiper and Urban.</p>
<p>Each dataset has been split into “training” – public, and private part. The parts were geographically disjoint, but shared a similar nuisance factors. We haven’t yet decided if we are going to release “hidden” part or not.</p>
<section id="heritage" class="level3">
<h3 class="anchored" data-anchor-id="heritage">Heritage</h3>
<p>This dataset features the high resolution photos of ancient buildings, taken with DLSR cameras from the ground, as well as UAV photos. Particular challenges are:</p>
<ul>
<li>large scale change (up to 20x), from that overview photo to the close-up of the small detail, together with in-plane rotation</li>
</ul>
<div class="quarto-figure quarto-figure-center">
<figure class="figure">
<p><img src="https://ducha-aiki.github.io/wide-baseline-stereo-blog/posts/2023-07-05-IMC2023-Recap_files/Dioscuri.png" class="img-fluid figure-img"></p>
<figcaption>Images of the Dioscuri temple from Heritage dataset</figcaption>
</figure>
</div>
<ul>
<li><a href="https://dev.epicgames.com/community/learning/tutorials/1xM4/capturing-reality-banana-effect-what-to-do-if-my-model-is-bent">“banana effect”</a>, when the flat surface is often reconstructed as curved due to low overlap between consecutive frames and slight misalignments, which accumulate altogether.</li>
</ul>
<div class="quarto-figure quarto-figure-center">
<figure class="figure">
<p><img src="https://ducha-aiki.github.io/wide-baseline-stereo-blog/posts/2023-07-05-IMC2023-Recap_files/wall.png" class="img-fluid figure-img"></p>
<figcaption>3D model of the <code>wall</code> from Heritage dataset</figcaption>
</figure>
</div>
<ul>
<li>high-resolution image processing. See above about the computational challenges of bundle adjustment and then multiply that by high resolution AND large image number. Also, the most of deep learning features (DISK, LoFTR) can easily give you CUDA OOM error when run on 20 Mp image pair.</li>
</ul>
</section>
<section id="haiper" class="level3">
<h3 class="anchored" data-anchor-id="haiper">Haiper</h3>
<p>Haiper (training) dataset is similar to captures for NERFs - layered “dome” of cameras, going around some object. The object itself is often thin (bicycle) or textureless (statue). The test part of the dataset has very small number of images, making the viewpoint difference the biggest challenge.</p>
<div class="quarto-figure quarto-figure-center">
<figure class="figure">
<p><img src="https://ducha-aiki.github.io/wide-baseline-stereo-blog/posts/2023-07-05-IMC2023-Recap_files/att_00000.png" class="img-fluid figure-img"></p>
<figcaption>3D model ‘bike’ from the Haiper dataset</figcaption>
</figure>
</div>
<p>If IMC-2023 have discovered and used “<a href="https://research.nianticlabs.com/mapfree-reloc-benchmark">Map-free Visual Relocalization</a>” for training their pipelines, that would likely help them with Haiper dataset as well. However, everyone has missed this opportunity, including us - organizers.</p>
</section>
<section id="urban" class="level3">
<h3 class="anchored" data-anchor-id="urban">Urban</h3>
<p>The Urban dataset is, probably, the most similar one to IMC2020 PhotoTourism dataset. It covers photos of buldings in city, similar to PhotoTourism.</p>
<p>Here is the photos of the “Kyiv Puppet Theater” – the easiest and public – part of the dataset.</p>
<div class="quarto-figure quarto-figure-center">
<figure class="figure">
<p><img src="https://ducha-aiki.github.io/wide-baseline-stereo-blog/posts/2023-07-05-IMC2023-Recap_files/att_00001.png" class="img-fluid figure-img"></p>
<figcaption>image.png</figcaption>
</figure>
</div>
<p>In addition to the day-night photos, the hidden part of the dataset features highly symmetrical objects, such as the photo I took yesterday in Český Krumlov. The camera poses are actually look into each other, so visual overlap is zero. The only chance of non-wrongly-matching them, is either considering all the photos altogether, or utilizing the background. None of the existing feature matching solutions does this, as far as I know.</p>
<div class="quarto-figure quarto-figure-center">
<figure class="figure">
<p><img src="https://ducha-aiki.github.io/wide-baseline-stereo-blog/posts/2023-07-05-IMC2023-Recap_files/IMG_4345_IMG_4352_matches.png" class="img-fluid figure-img"></p>
<figcaption>SuperGlue matches for symmetrical structures, photo from Český Krumlov to illustrate the challenges of the Urban dataset</figcaption>
</figure>
</div>
</section>
</section>
<section id="findings-from-the-competition" class="level2">
<h2 class="anchored" data-anchor-id="findings-from-the-competition">Findings from the competition</h2>
<section id="there-are-no-3d-reconstruction-libraries-besides-colmap" class="level3">
<h3 class="anchored" data-anchor-id="there-are-no-3d-reconstruction-libraries-besides-colmap">There are no 3d reconstruction libraries besides colmap</h3>
<p>At least, there are none, which you can easily compile on Kaggle kernel and then use from python notebook. I personally really hoped to see some global SfM solution like <a href="http://theia-sfm.org">Theia</a> or maybe <a href="https://github.com/mapillary/OpenSfM">OpenSfM</a> to appear among the top-solution,as it could provide a significant speed-up over the incremental SfM like Colmap. Python bindings and ease of compilation is crutial factor here.</p>
</section>
<section id="no-nerf-like-or-any-other-solution-than-classical-sfm" class="level3">
<h3 class="anchored" data-anchor-id="no-nerf-like-or-any-other-solution-than-classical-sfm">No NERF-like or any other solution than classical SfM</h3>
<p>Despite all the progress, it seems that if one doesn’t have any additional information, such as RGBD, or initial ARKit cameras poses, the best thing one can do with a challenging image collection, is classical SfM. No <a href="https://arxiv.org/abs/2211.16991">SparsePose</a>, no <a href="https://arxiv.org/abs/2212.04492">FORGE</a></p>
</section>
<section id="global-descriptor-based-co-visibility-is-hard" class="level3">
<h3 class="anchored" data-anchor-id="global-descriptor-based-co-visibility-is-hard">Global descriptor-based co-visibility is hard</h3>
<p>Given that exhaustive image matching grows quadratically with number of images, it is very tempting to filter out some of those image pairs, based on some kind of covisibility criterion. And the simplest/fastest to compute is global descriptor one - you get a single vector per image, calculate global image similarity and remove those image pairs, which similarity is below threshold. Or take the top-K most similar images. In fact, I have used such approach in our paper <a href="https://arxiv.org/abs/2011.11986">“Efficient Initial Pose-graph Generation for Global SfM”</a>. However, what worked on large (thousands) image collections with dense viewpoint coverage, as 1DSfM, does not work that well on sparse image collections with many close-ups, rotations and illumination changes.</p>
<p>Some teams had moderate success with <a href="https://arxiv.org/abs/1511.07247">NetVLAD</a> and <a href="https://github.com/facebookresearch/dinov2">DINOv2</a> global descriptors, however, the best strategy, it seems, to use the two-view matching itself, but on a smaller resolution – such as SuperGlue, or KeyNet-AffNet-HardNet-AdaLAM.</p>
</section>
<section id="pixelsfm-is-good-idea-but-needs-improvements" class="level3">
<h3 class="anchored" data-anchor-id="pixelsfm-is-good-idea-but-needs-improvements">PixelSfM is good idea, but needs improvements</h3>
<p>Many participants has tried to improve the initial SfM camera poses by utilizing <a href="https://github.com/cvg/pixel-perfect-sfm">PixelPerfectSfM</a> – the feature-metric bundle adjustment. While it improves results, it took a lot of time and especially memory, which rendered it unpractical for many teams. Another challenge is the package compilation itself, which is not easy either.</p>
<p><a href="https://www.kaggle.com/competitions/image-matching-challenge-2023/discussion/417407">1st-place team</a> proposed a novel version of the correspondence and poses refinement instead - called <a href="https://zju3dv.github.io/DetectorFreeSfM/">Detector-Free Structure from Motion</a>.</p>
<div class="quarto-figure quarto-figure-center">
<figure class="figure">
<p><img src="https://ducha-aiki.github.io/wide-baseline-stereo-blog/posts/2023-07-05-IMC2023-Recap_files/inbox-14597895-2b2b6c045d8a2dfa0a536090f025db02-main_fig.png" class="img-fluid figure-img"></p>
<figcaption>1st place solution, multiview refinement.</figcaption>
</figure>
</div>
</section>
<section id="rotation-invariance-is-important-but-easy-to-achieve-for-superglue-and-loftr" class="level3">
<h3 class="anchored" data-anchor-id="rotation-invariance-is-important-but-easy-to-achieve-for-superglue-and-loftr">Rotation-invariance is important, but easy to achieve for SuperGlue and LoFTR</h3>
<p>Just rotate one of the images 4 times and select the best matches. Here is the image of the <a href="https://www.kaggle.com/competitions/image-matching-challenge-2023/discussion/416918">3rd place solution</a>, explaining it all.</p>
<div class="quarto-figure quarto-figure-center">
<figure class="figure">
<p><img src="https://ducha-aiki.github.io/wide-baseline-stereo-blog/posts/2023-07-05-IMC2023-Recap_files/inbox-9249230-1036693590b6655d8d48d588bfd69e9c-IMC-solution.png" class="img-fluid figure-img"></p>
<figcaption>3rd place solution, brute-force rotation estimation, high-res image tiling</figcaption>
</figure>
</div>
</section>
<section id="initial-image-pair-setup-in-colmap-is-suboptimal" class="level3">
<h3 class="anchored" data-anchor-id="initial-image-pair-setup-in-colmap-is-suboptimal">Initial image pair setup in Colmap is suboptimal</h3>
<p>Several teams reported that manually setting the image pair to start incremental reconstruction in Colmap improved results. One even can improve results by running the incremental reconstruction several times.</p>
</section>
<section id="many-things-do-not-work-until-they-do" class="level3">
<h3 class="anchored" data-anchor-id="many-things-do-not-work-until-they-do">Many things do not work until they do</h3>
<p>Many teams have reported that LoFTR or DKM doesn’t work for them - but both LoFTR and DKM are part of top-5 solutions. Recent “SiLK” keypoints are reported as not working, but maybe nobody just found a proper way to use them.</p>
</section>
<section id="old-features-are-not-done-yet---keynetaffnet-hardnet-solution." class="level3">
<h3 class="anchored" data-anchor-id="old-features-are-not-done-yet---keynetaffnet-hardnet-solution.">“Old” features are not done yet - KeyNetAffNet-HardNet solution.</h3>
<p>5th place actually end up in money by using an modified <a href="https://github.com/ducha-aiki/imc2023-kornia-starter-pack/blob/main/keynetaffnet-adalam-pycolmap-3dreconstruction.ipynb">example submission</a> with a classical pipeline using such ancient local features as DoG, Harris and GFTT, together with KeyNet (2019), coupled together with <a href="https://ducha-aiki.github.io/wide-baseline-stereo-blog/posts/2020-07-17-affine-correspondences.html">local affine shape estimation</a> by AffNet, patch descriptor HardNet and handcrafted AdaLAM matcher.</p>
<p>Basically, that is cleverly engineered submission of the off-the-shelf local features, available in <a href="https://kornia.readthedocs.io/en/latest/feature.html">kornia</a>.</p>
<div class="quarto-figure quarto-figure-center">
<figure class="figure">
<p><img src="https://ducha-aiki.github.io/wide-baseline-stereo-blog/posts/2023-07-05-IMC2023-Recap_files/inbox-3964695-051e9413d0ce9cd8357723ba7efab1be-full_pipeline.png" class="img-fluid figure-img"></p>
<figcaption>5th place solution, classical affine features, off-the-shelf kornia. Candidate for 1st money prize</figcaption>
</figure>
</div>
<div class="quarto-figure quarto-figure-center">
<figure class="figure">
<p><img src="https://ducha-aiki.github.io/wide-baseline-stereo-blog/posts/2023-07-05-IMC2023-Recap_files/7rlp3j.jpg" class="img-fluid figure-img"></p>
<figcaption>Clever engineering is all you need sometimes</figcaption>
</figure>
</div>
</section>
<section id="permissive-license-and-faster-superglue-is-out---lightglue" class="level3">
<h3 class="anchored" data-anchor-id="permissive-license-and-faster-superglue-is-out---lightglue">Permissive license and faster SuperGlue is out - LightGlue</h3>
<p><a href="https://arxiv.org/pdf/2306.13643.pdf">LightGlue</a> - the solution by 7th place team and likely 2nd money prize, presents a SuperGlue-like architecture with early-stopping for easy image pairs and bunch of training recipes. It also uses somewhat unpopular, but well-performing <a href="https://github.com/Shiaoming/ALIKED">ALIKED local feature</a></p>
<div class="quarto-figure quarto-figure-center">
<figure class="figure">
<p><img src="https://ducha-aiki.github.io/wide-baseline-stereo-blog/posts/2023-07-05-IMC2023-Recap_files/att_00003.png" class="img-fluid figure-img"></p>
<figcaption>The LightGlue architecture from the paper</figcaption>
</figure>
</div>
<p>Here is the quite from the paper about tricks that matter:</p>
<blockquote class="blockquote">
<p>Since the depth maps of MegaDepth are often incomplete, we also label points with a large epipolar error as unmatch- able. Carefully tuning and annealing the learning rate boosts the accuracy. Training with more points also does: we use 2k per image instead of 1k. The batch size matters: we use gradient checkpointing [10] and mixed-precision to fit 32 image pairs on a single GPU with 24GB VRAM.</p>
</blockquote>
<p>In addition to that, homography-pretraining is crucial for the Light/Super-Glue performance.</p>
</section>
<section id="no-lines-monodepth-or-semantic-segmentation-this-year." class="level3">
<h3 class="anchored" data-anchor-id="no-lines-monodepth-or-semantic-segmentation-this-year.">No lines, monodepth or semantic segmentation this year.</h3>
<p>Actually I hoped to see <a href="https://github.com/cvg/limap/tree/main">Limap</a> line SfM or some kind of monocular depth models used. Bad luck.</p>
</section>
</section>
<section id="per-dataset-results" class="level2">
<h2 class="anchored" data-anchor-id="per-dataset-results">Per-dataset results</h2>
<section id="heritage-1" class="level3">
<h3 class="anchored" data-anchor-id="heritage-1">Heritage</h3>
<p>The Heritage dataset results show the biggest variability among the methods/worst/best case performance.</p>
<div class="quarto-figure quarto-figure-center">
<figure class="figure">
<p><img src="https://ducha-aiki.github.io/wide-baseline-stereo-blog/posts/2023-07-05-IMC2023-Recap_files/att_00007.png" class="img-fluid figure-img"></p>
<figcaption>image.png</figcaption>
</figure>
</div>
</section>
<section id="urban-1" class="level3">
<h3 class="anchored" data-anchor-id="urban-1">Urban</h3>
<p>Because the main Urban dataset challenge – repeated patterns and symmetrical structures – cannot be even theoretically addressed by better two-view feature matching, it is the hardest dataset. No single team tried to work on this part, so surprise it is hard.</p>
<div class="quarto-figure quarto-figure-center">
<figure class="figure">
<p><img src="https://ducha-aiki.github.io/wide-baseline-stereo-blog/posts/2023-07-05-IMC2023-Recap_files/att_00006.png" class="img-fluid figure-img"></p>
<figcaption>Urban dataset is the hardest among 2023 ones</figcaption>
</figure>
</div>
</section>
<section id="haiper-1" class="level3">
<h3 class="anchored" data-anchor-id="haiper-1">Haiper</h3>
<p>The Haiper results are either hit or miss w/o much difference between methods. <img src="https://ducha-aiki.github.io/wide-baseline-stereo-blog/posts/2023-07-05-IMC2023-Recap_files/att_00005.png" class="img-fluid" alt="Haiper scenes are like meeting the dinosaur: you either meet it, or not"></p>
</section>
</section>
<section id="conclusion-sfm-is-far-from-solved" class="level2">
<h2 class="anchored" data-anchor-id="conclusion-sfm-is-far-from-solved">Conclusion: SfM is far from solved</h2>
<p>As in many other computer vision tasks, if something seems to be “solved”, that is just because the datasets are old and obsolete. You still need a good dense capture to do the 3D reconstruction. Many images, many compute, or the additional data like inertial module/GPS, or skilled person to do the capture, is required. See you hopefully next year.</p>


</section>

<p>Everything you (didn’t) want to know about image matching</p> ]]></description>
  <guid>https://ducha-aiki.github.io/wide-baseline-stereo-blog/posts/2023-07-05-imc2023-recap.html</guid>
  <pubDate>Tue, 04 Jul 2023 22:00:00 GMT</pubDate>
  <media:content url="https://ducha-aiki.github.io/wide-baseline-stereo-blog/posts/images/pascal-cage.jpg" medium="image" type="image/jpeg"/>
</item>
<item>
  <title>Un-SuperPoint family: who are they?</title>
  <dc:creator>Dmytro Mishkin</dc:creator>
  <link>https://ducha-aiki.github.io/wide-baseline-stereo-blog/posts/2023-01-04-unsuperpoint-family.html</link>
  <description><![CDATA[ 





<section id="super-starts-with-magic" class="level2">
<h2 class="anchored" data-anchor-id="super-starts-with-magic">Super starts with Magic</h2>
<p>If you work in image matching, you know <a href="https://arxiv.org/abs/1712.07629">SuperPoint</a>. If you don’t - that is one of the most successful modern local feature, developed by Daniel DeTone, Tomasz Malisiewicz and Andrew Rabinovich in 2017.</p>
<p>Idea was simple and genious: we know that corners are good keypoints. Let’s train a basic corner (in wide meaning) on unlimited synthetic data. Then we somehow adapt it to realworld images.</p>
<div class="quarto-figure quarto-figure-center">
<figure class="figure">
<p><img src="https://ducha-aiki.github.io/wide-baseline-stereo-blog/posts/2022-12-31-UnsuperPoint-family_files/att_00001.png" class="img-fluid figure-img"></p>
<figcaption>SuperPoint starts with MagicPoint - corner/junctions detector supervised by synthetically rendered cube scenes.</figcaption>
</figure>
</div>
<p>The adaptation is done under augmentation supervision. In other word - we believe that our detector is already good, but noisy and we will cancel this noise by running the same detector on multiple augmented version of the same image, which gives us pseudo-ground truth. The training is done via optimizing cross-entropy loss, which leads to very peaky response map compared to other alternatives such as R2D2.</p>
<div class="quarto-figure quarto-figure-center">
<figure class="figure">
<p><img src="https://ducha-aiki.github.io/wide-baseline-stereo-blog/posts/2022-12-31-UnsuperPoint-family_files/att_00002.png" class="img-fluid figure-img"></p>
<figcaption>MagicPoint becomes SuperPoint after being finetuned on real images with augmentation supervision.</figcaption>
</figure>
</div>
<p>That’s it. The resulting detector (and descriptor) was fast enough on GPU and great for SLAM purposes. It was especially good on indoor images with lots of textureless areas.</p>
<p>On outdoor images it was not so great, however, as it was shown later, the problem was more in descriptor and matching, not detector. <a href="https://ducha-aiki.github.io/wide-baseline-stereo-blog/posts/2021-05-14-IMC2020-competition-recap.html">SuperGlue matching</a> on top of SuperPoint local features won 3 competitions at CVPR 2020, including our <a href="https://ducha-aiki.github.io/wide-baseline-stereo-blog/posts/2021-05-14-IMC2020-competition-recap.html">Image Matching Challenge</a>.</p>
<p>So, it was huge success. However, not without drawbacks. The main drawback is that training code was never released. That is why people tried independently re-implement SuperPoint, as well as present new, supposedly better versions of it.</p>
<p>Unfortunately, none of this version was properly evaluated, so we have no idea how they work in reality. Let me fill this gap and present a small eval of the SuperPoint children.</p>
<p>I’ll first do a short review of how are they different, and the benchmark will be in the last section.</p>
<section id="original-superpoint" class="level3">
<h3 class="anchored" data-anchor-id="original-superpoint">Original SuperPoint</h3>
<p>Besides the description, I will show how the detections are different between implementations. I will use two images. One is realworld photo taken by myself on Xiaomi phone, and another one is synthetic from <a href="https://learnopencv.com/blob-detection-using-opencv-python-c/">OpenCV Blob detection tutorial</a>. It is kind of adversarial, as SuperPoint is not designed to fire or not fire on circular patterns. Well, that is exactly why I added that image. So here are the original SP detections:</p>
<div class="quarto-figure quarto-figure-center">
<figure class="figure">
<p><img src="https://ducha-aiki.github.io/wide-baseline-stereo-blog/posts/2022-12-31-UnsuperPoint-family_files/sp_dets1.png" class="img-fluid figure-img"></p>
<figcaption>MagicLeap SuperPoint detections</figcaption>
</figure>
</div>
</section>
<section id="rd-party-superpoint-implementations" class="level3">
<h3 class="anchored" data-anchor-id="rd-party-superpoint-implementations">3rd party SuperPoint implementations</h3>
<p>There are two main 3rd party SuperPoint implementations. <a href="https://github.com/rpautrat/SuperPoint">One of them</a> is in Tensorflow, by Rémi Pautrat and Paul-Edouard Sarlin. I will skip this one, because I am too lazy to install TF on my M1 machine. So I will show another - Pytorch - implementation, which is based on Tensorflow one, and is developed by <a href="https://github.com/eric-yyjau/pytorch-superpoint">You-Yi Jau and Rui Zhu</a>.</p>
<p>Besides being 3rd party implementation, this one also has architectural changes. They are: - <a href="https://github.com/eric-yyjau/pytorch-superpoint/blob/master/models/SuperPointNet_gauss2.py#L31">adding BatchNorm</a>. - using <a href="https://github.com/eric-yyjau/pytorch-superpoint/blob/master/models/model_utils.py#L118">SoftArgMax2d from torchgeometry</a> (early version of <a href="https://github.com/kornia/kornia">kornia</a>) to achieve subpixel accuracy.</p>
<p>Here are the detections by this version and original.</p>
<div class="quarto-figure quarto-figure-center">
<figure class="figure">
<p><img src="https://ducha-aiki.github.io/wide-baseline-stereo-blog/posts/2022-12-31-UnsuperPoint-family_files/sp.gif" class="img-fluid figure-img"></p>
<figcaption>Pytorch-SuperPoint and MagicLeap SuperPoint detections</figcaption>
</figure>
</div>
<p>Here is also a short report by one of the authors: <a href="https://eric-yyjau.medium.com/what-have-i-learned-from-the-implementation-of-deep-learning-paper-365ee3253a89">What have I learned from the implementation of deep learning paper?</a></p>
</section>
</section>
<section id="superchildren" class="level2">
<h2 class="anchored" data-anchor-id="superchildren">SuperChildren</h2>
<section id="reinforced-superpoint" class="level3">
<h3 class="anchored" data-anchor-id="reinforced-superpoint">Reinforced SuperPoint</h3>
<p>This is a paper by Aritra Bhowmik et al, named <a href="https://arxiv.org/pdf/1912.00623.pdf">Reinforced Feature Points: Optimizing Feature Detection and Description for a High-Level Task</a>. The main idea is that one can use reinforcement learning to optimize non-differentiable downstream metric such as camera pose accuracy through RANSAC.</p>
<div class="quarto-figure quarto-figure-center">
<figure class="figure">
<p><img src="https://ducha-aiki.github.io/wide-baseline-stereo-blog/posts/2022-12-31-UnsuperPoint-family_files/att_00007.png" class="img-fluid figure-img"></p>
<figcaption>Reinforced SuperPoint training pipeline</figcaption>
</figure>
</div>
<p>The idea is very cool, but the final result is not so. Specifically, the main thing end up optimized is keypoint score function. Which can increase a precision of the keypoint detector, but not the recall. See the image below for the illustration. The synthetic image at the right clearly benefits, but the realworld image - not so much, because of loosing many keypoints on the ground.</p>
<p>This also can explain a bit worse performance of the Reinforced SuperPoint in our evaluation - we don’t set the confidence threshold, but instead take top-2048 keypoints whatsoever.</p>
<div class="quarto-figure quarto-figure-center">
<figure class="figure">
<p><img src="https://ducha-aiki.github.io/wide-baseline-stereo-blog/posts/2022-12-31-UnsuperPoint-family_files/sp_reinf.gif" class="img-fluid figure-img"></p>
<figcaption>SuperPoint vs its finetuned version via Reinforcement learning. One can see that keypoint map is cleaner, but there are no new keypoints appearing</figcaption>
</figure>
</div>
</section>
<section id="unsuperpoint" class="level3">
<h3 class="anchored" data-anchor-id="unsuperpoint">UnsuperPoint</h3>
<p>This is a <a href="https://arxiv.org/abs/1907.04011">paper from eiva.com and Aarhus University by Peter Hviid Christiansen</a>, which proposed to drop the supervised pretraining and use regression module instead of CE on heatmap for training detector.</p>
<div class="quarto-figure quarto-figure-center">
<figure class="figure">
<p><img src="https://ducha-aiki.github.io/wide-baseline-stereo-blog/posts/2022-12-31-UnsuperPoint-family_files/att_00003.png" title="UnsuperPoint training scheme" class="img-fluid figure-img"></p>
<figcaption>image.png</figcaption>
</figure>
</div>
<p>Unfortunately, there is no implementation available, so I will pretend the paper never existed.</p>
</section>
<section id="kp2d-aka-keypointnet-aka-ionet" class="level3">
<h3 class="anchored" data-anchor-id="kp2d-aka-keypointnet-aka-ionet">KP2D aka KeypointNet aka IONet</h3>
<p>Despite no implementation, UnsuperPoint inspired other people to make a follow-up, which was published at ICLR 2020. The paper name is “<a href="https://openreview.net/pdf?id=Skx82ySYPH">Neural Outlier Rejection for Self-Supervised Keypoint Learning</a>” by Jiexiong Tang et al.</p>
<p>It made two contributions. First, paper argues that cell-based approach (SuperPoint and friends have 1 keypoint per 8x8 px cell) is unstable for training when keypoints are near the cell border. Second, it introduced yet another loss function, similar to CNe and other outlier rejection methods.</p>
<div class="quarto-figure quarto-figure-center">
<figure class="figure">
<p><img src="https://ducha-aiki.github.io/wide-baseline-stereo-blog/posts/2022-12-31-UnsuperPoint-family_files/att_00005.png" class="img-fluid figure-img"></p>
<figcaption>Cross-border detection in KP2D network</figcaption>
</figure>
</div>
<p>Here are the KP2D detections for all versions. Note that on realworld image it detects corners nicely, despite not being training for it specifically. On the synthetic blob image it mostly work as dark-to-light (or light-to-dark for V0 and V2) edge detector, no idea why.</p>
<div class="quarto-figure quarto-figure-center">
<figure class="figure">
<p><img src="https://ducha-aiki.github.io/wide-baseline-stereo-blog/posts/2022-12-31-UnsuperPoint-family_files/kp2d.gif" class="img-fluid figure-img"></p>
<figcaption>KeypointNet aka KP2D detections.</figcaption>
</figure>
</div>
</section>
<section id="lanet" class="level3">
<h3 class="anchored" data-anchor-id="lanet">LANet</h3>
<p>It is the most recent paper of the SP family, published at ACCV2022 – “<a href="https://openaccess.thecvf.com/content/ACCV2022/html/Wang_Rethinking_Low-level_Features_for_Interest_Point_Detection_and_Description_ACCV_2022_paper.html">Rethinking Low-level Features for Interest Point Detection and Description</a>” by Changhao Wang et al.</p>
<div class="quarto-figure quarto-figure-center">
<figure class="figure">
<p><img src="https://ducha-aiki.github.io/wide-baseline-stereo-blog/posts/2022-12-31-UnsuperPoint-family_files/att_00006.png" class="img-fluid figure-img"></p>
<figcaption>LANet architecture</figcaption>
</figure>
</div>
<p>It is based on KP2D and presents mostly architectural changes into description branch. Second, it has two versions. <code>v0</code> is similar to the SuperPoint original architecture – lightweight VGG-like network, trained from scratch. <code>v1</code> is uses <code>vgg16_bn</code>, pretrained on the ImageNet as feature encoder. <code>v1</code> is like 2-3x slower than <code>v0</code>.</p>
<p>Here are the detections:</p>
<div class="quarto-figure quarto-figure-center">
<figure class="figure">
<p><img src="https://ducha-aiki.github.io/wide-baseline-stereo-blog/posts/2022-12-31-UnsuperPoint-family_files/lanet.gif" class="img-fluid figure-img"></p>
<figcaption>LANet detections</figcaption>
</figure>
</div>
</section>
</section>
<section id="benchmark-on-imc-phototourism" class="level1">
<h1>Benchmark on IMC PhotoTourism</h1>
<p>I have benchmarked all the variants, using my own tutorial on <a href="https://ducha-aiki.github.io/wide-baseline-stereo-blog/posts/2021-05-12-submitting-to-IMC2021-step-by-step.html">how to submit to IMC-2021</a>. The code for feature extraction is <a href="https://github.com/ubc-vision/image-matching-benchmark-baselines">here</a> (uncommented yet). All the images were resized to have 1024 px by large size (which explains the slight diffefences compared to <a href="https://arxiv.org/abs/2003.01587">original SuperPoint results in the paper</a>, where 1200 was used. I have also circumvented confidence threshold and pick top-2048 keypoints instead for all images. Then I have tuned everything on the validation set (see table with optimal params in the end of the post) - such as matching threshold, RANSAC threshold and method-specific things, like “version” for KP2D and LANet, etc. Unlike for the IMC paper, I haven’t tuned anything for multiview setup. Specifically, matching threshold might be suboptimal. Anyway.</p>
<p>Finally I have run all features with found optimal parameters on the test set (GT is released now!). Here are the results.</p>
<section id="stereo-task" class="level3">
<h3 class="anchored" data-anchor-id="stereo-task">Stereo task</h3>
<div class="quarto-figure quarto-figure-center">
<figure class="figure">
<p><img src="https://ducha-aiki.github.io/wide-baseline-stereo-blog/posts/2022-12-31-UnsuperPoint-family_files/map_stereo_pp_pt_2k.png" class="img-fluid figure-img"></p>
<figcaption>Stereo task results. MagicLeap SuperPoint results are shown with dashed line.</figcaption>
</figure>
</div>
<p>First, as you can see, 3rd part SuperPoint implementation is way worse: ~2x than MagicLeap one. Reinforced SuperPoint also is slightly worse. KP2D is better on PhotoTourism, but worse on the PragueParks. <a href="https://www.cs.ubc.ca/research/image-matching-challenge/2021/data/">PragueParks</a> is easier dataset than PhotoTourism, however it features different scenes. While PhotoTourism is mostly buildings, PragueParks contains natural scenes such as trees and pond. Thus I may argue that original SP is more general.</p>
<p>LANet works considerably better on Phototourism and on par on PragueParks. If considering adding to kornia, I would select LANet. It is interesting though, if it is performing well because of detector, or descriptor. The latter is worse, because training SuperGlue/OpenGlue would fix descriptor issues, but less of detector.</p>
</section>
<section id="multiview-task" class="level3">
<h3 class="anchored" data-anchor-id="multiview-task">Multiview task</h3>
<div class="quarto-figure quarto-figure-center">
<figure class="figure">
<p><img src="https://ducha-aiki.github.io/wide-baseline-stereo-blog/posts/2022-12-31-UnsuperPoint-family_files/map_mv_pp_pt_2k.png" class="img-fluid figure-img"></p>
<figcaption>Multiview task results. MagicLeap SuperPoint results are shown with dashed line.</figcaption>
</figure>
</div>
<p>Multiview results are similar to stereo, with slight differences: Reinforced SP is slightly better than original on PragueParks, and LANet v0 is exactly as good as original SP.</p>
</section>
<section id="why-eval-only-on-hpatches-is-meaningless." class="level2">
<h2 class="anchored" data-anchor-id="why-eval-only-on-hpatches-is-meaningless.">Why eval only on HPatches is meaningless.</h2>
<p>KP2D and 3rd party SuperPoint eval themselves on HPatches only. Let’s look at them. According to KP2D paper, it is slightly better than Superpoint on homography estimation and 2x better on matching score. Which we haven’t seen on realworld IMC data.</p>
<div class="quarto-figure quarto-figure-center">
<figure class="figure">
<p><img src="https://ducha-aiki.github.io/wide-baseline-stereo-blog/posts/2022-12-31-UnsuperPoint-family_files/att_00009.png" class="img-fluid figure-img"></p>
<figcaption>Evaluation on HPatches from KP2D paper. KP2D seems to be much better</figcaption>
</figure>
</div>
<p>And here is <a href="https://github.com/eric-yyjau/pytorch-superpoint/tree/master">eval from 3rd part SuperPoint</a>. According to it, Superpoint-coco is almost as good as SP MagicLeap on homography estimation, better on detector metrics and slightle worse on descriptor metrics. Which, again, doesn’t correspond to our results.</p>
<table class="caption-top table">
<colgroup>
<col style="width: 31%">
<col style="width: 16%">
<col style="width: 4%">
<col style="width: 4%">
<col style="width: 12%">
<col style="width: 4%">
<col style="width: 13%">
<col style="width: 11%">
</colgroup>
<thead>
<tr class="header">
<th>Task</th>
<th>Homography estimation</th>
<th></th>
<th></th>
<th>Detector metric</th>
<th></th>
<th>Descriptor metric</th>
<th></th>
</tr>
</thead>
<tbody>
<tr class="odd">
<td></td>
<td>Epsilon = 1</td>
<td>3</td>
<td>5</td>
<td>Repeatability</td>
<td>MLE</td>
<td>NN mAP</td>
<td>Matching Score</td>
</tr>
<tr class="even">
<td>MagicLeap</td>
<td>0.44</td>
<td>0.77</td>
<td>0.83</td>
<td>0.606</td>
<td>1.14</td>
<td>0.81</td>
<td>0.55</td>
</tr>
<tr class="odd">
<td>superpoint_coco_heat2_0_170k_hpatches_sub</td>
<td>0.46</td>
<td>0.75</td>
<td>0.81</td>
<td>0.63</td>
<td>1.07</td>
<td>0.78</td>
<td>0.42</td>
</tr>
<tr class="even">
<td>superpoint_kitti_heat2_0_50k_hpatches_sub</td>
<td>0.44</td>
<td>0.71</td>
<td>0.77</td>
<td>0.56</td>
<td>0.95</td>
<td>0.78</td>
<td>0.41</td>
</tr>
</tbody>
</table>
</section>
<section id="additional-results-from-validation-set-and-recommended-hyper-paramers" class="level2">
<h2 class="anchored" data-anchor-id="additional-results-from-validation-set-and-recommended-hyper-paramers">Additional results from validation set and recommended hyper-paramers</h2>
<p>Here I will present some results from the tuning on validation set. I believe that they would translate to the other datasets as well.</p>
<ol type="1">
<li>L2-normalization of descriptor does not help ANY of the evaluated models. It doesn’t hurt either.</li>
<li>Subpixel with soft-argmax helps 3rd party SuperPoint. One may try to apply it to original model as well. The difference is: <code>mAA = 0.2965</code> for subpixel vs <code>mAA = 0.2789</code> no-subpixel on PhotoTourism and <code>0.3567</code> vs <code>0.3274</code> on PragueParks.</li>
<li>The <code>v2</code> model of the KP2D is much better than then rest. <code>v0</code> and <code>v1</code> might be bugged, or require code changes to be run properly maybe?</li>
</ol>
<div class="quarto-figure quarto-figure-center">
<figure class="figure">
<p><img src="https://ducha-aiki.github.io/wide-baseline-stereo-blog/posts/2022-12-31-UnsuperPoint-family_files/map_stereo_pt_2k_val_kp2d.png" class="img-fluid figure-img"></p>
<figcaption>KP2D versions comparison on validation set.</figcaption>
</figure>
</div>
</section>
<section id="optimal-hyperparamers" class="level2">
<h2 class="anchored" data-anchor-id="optimal-hyperparamers">Optimal hyperparamers</h2>
<p>SNN - 2nd nearest neighbor threshold, PT - PhotoTourism, PP - PragueParks, inl_th - inlier threshold for DEGENSAC.</p>
<table class="caption-top table">
<colgroup>
<col style="width: 7%">
<col style="width: 23%">
<col style="width: 10%">
<col style="width: 12%">
<col style="width: 19%">
<col style="width: 19%">
<col style="width: 6%">
</colgroup>
<thead>
<tr class="header">
<th>Method</th>
<th>mutual SNN threshold</th>
<th>PT inl_th</th>
<th>PP inl_th</th>
<th>PT val stereo mAA</th>
<th>PP val stereo mAA</th>
<th>Other</th>
</tr>
</thead>
<tbody>
<tr class="odd">
<td>MagicLeap SuperPoint</td>
<td>0.9</td>
<td>1.0</td>
<td>1.5</td>
<td>0.3746</td>
<td>0.5628</td>
<td>n/a</td>
</tr>
<tr class="even">
<td>Reinforced SuperPoint</td>
<td>0.9</td>
<td>1.0</td>
<td>1.5</td>
<td>0.3491</td>
<td>0.5497</td>
<td>n/a</td>
</tr>
<tr class="odd">
<td>SuperPoint 3rd party COCO</td>
<td>0.95</td>
<td>0.75</td>
<td>1.0</td>
<td>0.2966</td>
<td>0.3488</td>
<td>subpix</td>
</tr>
<tr class="even">
<td>SuperPoint 3rd party KITTY</td>
<td>0.95</td>
<td>0.75</td>
<td>1.0</td>
<td>0.1910</td>
<td>0.2621</td>
<td>subpix</td>
</tr>
<tr class="odd">
<td>KP2D</td>
<td>0.99</td>
<td>0.75</td>
<td>1.0</td>
<td>0.3633</td>
<td>0.4971</td>
<td>v2</td>
</tr>
<tr class="even">
<td>LANet v0</td>
<td>0.99</td>
<td>0.5</td>
<td>1.0</td>
<td>0.4591</td>
<td><strong>0.6175</strong></td>
<td>n/a</td>
</tr>
<tr class="odd">
<td>LANet v1</td>
<td>0.99</td>
<td>0.5</td>
<td>1.0</td>
<td><strong>0.4838</strong></td>
<td><strong>0.6127</strong></td>
<td>n/a</td>
</tr>
</tbody>
</table>
</section>
</section>
<section id="conclusions" class="level1">
<h1>Conclusions</h1>
<p>It seems that LANet and KP2D might be good alternatives to the MagicLeap SuperPoint. There are 2 missing things though:</p>
<ol type="1">
<li>SuperGlue analogue yet to be trained for them. Nobody sane would you SuperPoint w/o SuperGlue for the most of the cases.</li>
<li>Evaluation on indoor data, e.g.&nbsp;ScanNet. One of the SuperPoint strengths - good performance indoor, which is yet to be tested for others. I might test it, but right not in the mood to download, print and scan the data form required for ScanNet access.</li>
</ol>
<section id="bonus-more-visualizations" class="level2">
<h2 class="anchored" data-anchor-id="bonus-more-visualizations">Bonus: more visualizations</h2>
<p>Legend: - blue points: COLMAP reconstructed 3D points - red point: non-matched keypoints .</p>
<div class="quarto-figure quarto-figure-center">
<figure class="figure">
<p><img src="https://ducha-aiki.github.io/wide-baseline-stereo-blog/posts/2022-12-31-UnsuperPoint-family_files/tree1.gif" class="img-fluid figure-img"></p>
<figcaption>Matched and non-matched keypoints from PragueParks Tree scene</figcaption>
</figure>
</div>
<div class="quarto-figure quarto-figure-center">
<figure class="figure">
<p><img src="https://ducha-aiki.github.io/wide-baseline-stereo-blog/posts/2022-12-31-UnsuperPoint-family_files/ll.gif" class="img-fluid figure-img"></p>
<figcaption>Matched and non-matched keypoints from PhotoTourism Lincoln memorial statue scene</figcaption>
</figure>
</div>
<div class="quarto-figure quarto-figure-center">
<figure class="figure">
<p><img src="https://ducha-aiki.github.io/wide-baseline-stereo-blog/posts/2022-12-31-UnsuperPoint-family_files/flo.gif" class="img-fluid figure-img"></p>
<figcaption>Matched and non-matched keypoints from PhotoTourism Florence cathedral side scene</figcaption>
</figure>
</div>
<div class="quarto-figure quarto-figure-center">
<figure class="figure">
<p><img src="https://ducha-aiki.github.io/wide-baseline-stereo-blog/posts/2022-12-31-UnsuperPoint-family_files/pond1.gif" class="img-fluid figure-img"></p>
<figcaption>Matched and non-matched keypoints from PragueParks Pond scene</figcaption>
</figure>
</div>
<section id="acknowledgements." class="level3">
<h3 class="anchored" data-anchor-id="acknowledgements.">Acknowledgements.</h3>
<p>This blogpost is supported by CTU in Prague RCI computing cluster from <code>OP VVV funded project CZ.02.1.01/0.0/0.0/16 019/0000765 “Research Center for Informatics”</code> grant. Really, it took ~2 compute days to tune all those hyperparams and do tests.</p>


</section>
</section>
</section>

<p>Everything you (didn’t) want to know about image matching</p> ]]></description>
  <guid>https://ducha-aiki.github.io/wide-baseline-stereo-blog/posts/2023-01-04-unsuperpoint-family.html</guid>
  <pubDate>Tue, 03 Jan 2023 23:00:00 GMT</pubDate>
  <media:content url="https://ducha-aiki.github.io/wide-baseline-stereo-blog/posts/images/eval-on-IMC-vs-Hpatches.png" medium="image" type="image/png" height="85" width="144"/>
</item>
<item>
  <title>Image Matching Challenge 2022 Recap</title>
  <dc:creator>Dmytro Mishkin</dc:creator>
  <link>https://ducha-aiki.github.io/wide-baseline-stereo-blog/posts/2022-07-05-imc2022-recap.html</link>
  <description><![CDATA[ 





<section id="what-has-been-changed-in-the-challenge-itself-kaggle" class="level2">
<h2 class="anchored" data-anchor-id="what-has-been-changed-in-the-challenge-itself-kaggle">What has been changed in the challenge itself? Kaggle!</h2>
<p>Yes, it is the main change, compared to previous years. How it works: - Participants submit notebooks that run offline - Test set remains private (not visible to participants) : very difficult to cheat (big concern at Kaggle) - Allows for quick iterations</p>
<p>That allowed us to get 25x more participants and 150x more submissions</p>
<div class="quarto-figure quarto-figure-center">
<figure class="figure">
<p><img src="https://ducha-aiki.github.io/wide-baseline-stereo-blog/posts/2022-06-29-IMC2022_files/att_00001.png" class="img-fluid figure-img"></p>
<figcaption>We have 14 186 submissions from 642 teams this year</figcaption>
</figure>
</div>
<p>There are other differences compared to the previous years.</p>
<ol type="1">
<li><p>No multiview track. Unlike previous competitions with multivew and stereo tracks, this year we focused on stereo only. There are many reasons for it, but the main is technical difficulty - to run and evaluate this under reasonable time limit. Having a pair of images as an input also could potentially broaden the family of methods used by participants.</p></li>
<li><p>New dataset and new metric. Unlike previous years, where our GT data was scale-less, now we have a metric ground truth, so we can evaluate translation error more in line with visual re-localization literature. We also used non-public dataset from Google, not available anywhere online. This was also only possible because of <a href="https://www.kaggle.com/docs/competitions#notebooks-only-FAQ">Kaggle Code competition</a> nature.</p></li>
<li><p>Time limit. I have mentioned this already, but it is important. The submission should finish in 9 hours on Kaggle GPU virtual instance. This forced teams to think, what to add, and what not. Simple example would be a semantic segmentation masking - teams decided that benefit is not worth computational overhead it brings.</p></li>
</ol>
<div class="quarto-figure quarto-figure-center">
<figure class="figure">
<p><img src="https://ducha-aiki.github.io/wide-baseline-stereo-blog/posts/2022-06-29-IMC2022_files/att_00002.png" class="img-fluid figure-img"></p>
<figcaption>Inliers from LoFTR and KeyNetAffNetHardnet from kornia on one of 3 image pairs</figcaption>
</figure>
</div>
</section>
<section id="top-solutions-2022-overview" class="level2">
<h2 class="anchored" data-anchor-id="top-solutions-2022-overview">Top solutions-2022 overview</h2>
<p>Unlike previous competitions, almost all the top-performing teams have converged to almost the same pipeline, although different in implementation details, but based on the same ideas. Let me condence it is for you, and if you are interested in participants own write-ups, just scroll down, I will provide all the links after the post.</p>
<p>The only top-participant, who did not apply any pre/post-processing is 2nd place winner hkchkc. It is based on novel matcher (let’s call it LoFTR 2.0), which he did not want to disclose yet. Anyway, check <a href="https://www.kaggle.com/competitions/image-matching-challenge-2022/discussion/329317">his write-up here</a>.</p>
<p>So, the general pipeline is shown below. The following steps are present in the most of solutions:</p>
<div class="quarto-figure quarto-figure-center">
<figure class="figure">
<p><img src="https://ducha-aiki.github.io/wide-baseline-stereo-blog/posts/2022-06-29-IMC2022_files/general-scheme-1c.png" class="img-fluid figure-img"></p>
<figcaption>Common pipeline among the top-performers of IMC-2022</figcaption>
</figure>
</div>
<ol type="1">
<li><p>Obtaining initial pool of matches, using a combination of off-the-shelf models. Such models are LoFTR, SuperGlue, QuadTree Attention LoFTR and DKM. Some of the teams also applied test-time scale augmentation (TTA) and concatenated tentative correspondence, which came from matching differently-resize images.</p></li>
<li><p>The next step is an estimation of the co-visible area between two images, similar to <a href="https://arxiv.org/abs/2112.04846">ScaleNet</a> idea, but totally handcrafted and based on tentative matches. The approach is to cluster the matches using K-means or DBSCAN and then find the bounding box in each image, which contains the most of the matches. This can lead to potentially several “co-visible areas”. One may take a step further and estimate the homography between each matching region. An alternative to clustering (from 9th place) is to use MAGSAC++ with few iterations and coarse threshold to reject the most of outliers, and get the bounding box .</p></li>
<li><p>Each of such co-visible and matching regions is then cropped and resized to a predefined image size. Then the base-matcher, or other set of matching algorithms are applied for this “zoomed-in” image pair. All the correspondences are then reprojected into original image pair coordinates and contatenated with original correspondences. <strong>Important</strong>: “zoomed-in” is used to entend original tentative correspondences, not replace them, otherwise first failure would kill the matching process.</p></li>
<li><p>The correspondences are (optionally) filtered with some kind of non-maxima-suppression, like <a href="https://github.com/BAILOOL/ANMS-Codes">ANMS</a> or radius-based NMS. The intuition here is to select the fewer set of more precise matches. It is interesting, that no team run any kind of filtering network as OANet.</p></li>
<li><p>MAGSAC++ from OpenCV (as recommended in <a href="https://ducha-aiki.github.io/wide-baseline-stereo-blog/posts/2021-05-17-OpenCV-New-RANSACs.html">my evaluation</a>) is applied to find fundamental matrix.</p></li>
</ol>
</section>
<section id="tricks-which-improve-results" class="level2">
<h2 class="anchored" data-anchor-id="tricks-which-improve-results">Tricks which improve results</h2>
<ol type="1">
<li><p><a href="https://www.kaggle.com/competitions/image-matching-challenge-2022/discussion/328865">Switching img1 and img2 for LoFTR-like methods</a>, because the method is asymmetric.</p></li>
<li><p>Running RANSAC and local feature extraction in separate threads. Although competition specific, this also may help to speed-up some semi-real-time SLAM methods.</p></li>
</ol>
</section>
<section id="works-with-tuning" class="level2">
<h2 class="anchored" data-anchor-id="works-with-tuning">Works with tuning</h2>
<ol type="1">
<li>Some teams successfully used <a href="https://arxiv.org/abs/2003.08348">two-view keypoint refinement</a>, while others <a href="https://www.kaggle.com/competitions/image-matching-challenge-2022/discussion/328982">didn’t make it work</a>.</li>
</ol>
<div class="quarto-figure quarto-figure-center">
<figure class="figure">
<p><img src="https://ducha-aiki.github.io/wide-baseline-stereo-blog/posts/2022-06-29-IMC2022_files/att_00005.png" class="img-fluid figure-img"></p>
<figcaption>Illustration from Multi-View Optimization of Local Feature Geometry</figcaption>
</figure>
</div>
<ol start="2" type="1">
<li>Unlike <a href="https://ieeexplore.ieee.org/abstract/document/9511155">MAGSAC++</a>, which works out of the box, <a href="https://arxiv.org/abs/2106.10240v2">VSAC</a> could bring significant time improvements, but needs a lot of hyperparameter tuning, otherwise epipolar geometry quality suffers.</li>
</ol>
</section>
<section id="rare-findings" class="level2">
<h2 class="anchored" data-anchor-id="rare-findings">Rare findings</h2>
<ol type="1">
<li><p>Normalizing positional embeddings used in LoFTR-like methods. See <a href="https://www.kaggle.com/competitions/image-matching-challenge-2022/discussion/329317">solution 2</a>.</p></li>
<li><p>When doing TTA, run and cache SuperPoint prior to SuperGlue matching, this saves a lot of time. Also, SuperGlue whould be run on the extracted, not reprojected keypoints.</p></li>
<li><p>Resize method matters a little bit: <a href="https://www.kaggle.com/competitions/image-matching-challenge-2022/discussion/328803#1809862">Lanczos</a> seems to work the best.</p></li>
<li><p>Coordinates refinement by ECO-TR (improved fast version of <a href="https://github.com/ubc-vision/COTR">COTR</a>, not yet available online)</p></li>
</ol>
<div class="quarto-figure quarto-figure-center">
<figure class="figure">
<p><img src="https://ducha-aiki.github.io/wide-baseline-stereo-blog/posts/2022-06-29-IMC2022_files/att_00006.png" class="img-fluid figure-img"></p>
<figcaption>Refinement by ECO-TR</figcaption>
</figure>
</div>
</section>
<section id="does-not-work" class="level2">
<h2 class="anchored" data-anchor-id="does-not-work">Does not work</h2>
<ol type="1">
<li><p>Adding tentative matches from local features without learned matcher, such as DISK, ALike, etc. Team 1 mentioned that KeyNetAffNetHardNet was the one, which do imporove results, but at the cost of huge computational overhead, so it doesn’t worth it.</p></li>
<li><p>Semantic segmentation masking (sky, people, etc).</p></li>
</ol>
</section>
<section id="final-thoughts" class="level2">
<h2 class="anchored" data-anchor-id="final-thoughts">Final thoughts</h2>
<p>Sooo…first think I want to say, is that going Kaggle was definitely worth it. We learned a lot, challenge got visibility, and many new improvements came out of of the challenge.</p>
<p>Second, is that our training set was too small and and too different from the test set, that is why it was almost useless. I don’t have any solution yet for this.</p>
<p>Thresholds on the pose accuracy were probably too loose, as you can see from the results.</p>
<p>On the lessons from the solution themselves:</p>
<ol type="1">
<li><p>2-stage approach of first finding covisible region, and then zooming-in into it seems to be way to go. What is also important, is to delay the decision until the very end.</p></li>
<li><p>It is better to first fix recall problem, i.e.&nbsp;to get as many tentative correspondences as possible, e.g.&nbsp;from different methods, and only the filter them. Modern RANSACs can recover the pose from a small inlier ratio, but never from small number of correct correspondences themselves.</p></li>
<li><p>Learned methods like LoFTR are very sensitive on the input image size, we probably have to do something about it.</p></li>
</ol>
</section>
<section id="list-of-all-top-solution-write-ups" class="level2">
<h2 class="anchored" data-anchor-id="list-of-all-top-solution-write-ups">List of all top solution write-ups</h2>
<p>It is <a href="https://www.kaggle.com/competitions/image-matching-challenge-2022/discussion/329650">here</a>, created by Addison Howard.</p>


</section>

<p>Everything you (didn’t) want to know about image matching</p> ]]></description>
  <guid>https://ducha-aiki.github.io/wide-baseline-stereo-blog/posts/2022-07-05-imc2022-recap.html</guid>
  <pubDate>Mon, 04 Jul 2022 22:00:00 GMT</pubDate>
  <media:content url="https://ducha-aiki.github.io/wide-baseline-stereo-blog/posts/images/loftrs_sg_everywhere_small.jpeg" medium="image" type="image/jpeg"/>
</item>
<item>
  <title>Image Similarity Challenge 2021 recap</title>
  <dc:creator>Dmytro Mishkin</dc:creator>
  <link>https://ducha-aiki.github.io/wide-baseline-stereo-blog/posts/2022-02-05-disc21-recap.html</link>
  <description><![CDATA[ 





<section id="what-is-image-similarity-challenge" class="level2">
<h2 class="anchored" data-anchor-id="what-is-image-similarity-challenge">What is Image Similarity Challenge?</h2>
<p>Image Similarity Challenge is benchmark of the methods, which should find the original in the database (if any) by the edited query image. The main difference to the classical instance retrieval task (datasets <a href="https://www.robots.ox.ac.uk/~vgg/data/oxbuildings/">Oxford5k</a>, <a href="https://github.com/cvdfoundation/google-landmark">GLD</a>, etc) is that we should NOT retrieve the different photos of the given instance. Instead only version of the same photos are valid. In other words, we are looking for the image forgery and copyright infringement.</p>
<div class="quarto-figure quarto-figure-center">
<figure class="figure">
<p><img src="https://ducha-aiki.github.io/wide-baseline-stereo-blog/posts/2022-02-05-DISC21-recap_files/att_00000.png" class="img-fluid figure-img"></p>
<figcaption>Example of the query and reference images in the ISC</figcaption>
</figure>
</div>
<section id="image-similarity-dataset-disc" class="level3">
<h3 class="anchored" data-anchor-id="image-similarity-dataset-disc">Image Similarity Dataset (DISC)</h3>
<p>The dataset is described in the “<a href="https://arxiv.org/pdf/2106.09672.pdf">The 2021 Image Similarity Dataset and Challenge</a>” paper. It is created by processing YCC100M and DeepFake Detection Challenge images with manually in GIMP editor and automatically with help of <a href="https://arxiv.org/abs/2201.06494">AugLy</a> library.</p>
<div class="quarto-figure quarto-figure-center">
<figure class="figure">
<p><img src="https://ducha-aiki.github.io/wide-baseline-stereo-blog/posts/2022-02-05-DISC21-recap_files/att_00001.png" class="img-fluid figure-img"></p>
<figcaption>More example of the image transformations used in DISC. Taken from DISC paper</figcaption>
</figure>
</div>
<p>Besides the public part of the test set (Phase I of the competition), there is also final test set (Phase II) was also composed with use of adversarial attack methods and so-called “air-gap attack”. Air-gap, according to the paper is</p>
<blockquote class="blockquote">
<p>We asked manual editors to perform “air-gap” transformations, where the picture is displayed or printed out, and re-captured with a camera or mobile phone. It is a quite common image attack, that is hard to reproduce automatically or even with an image editor like the GIMP. Figure 8 shows an example air-gap augmentation (the zebra). The effect is a combination of over-exposure, geometric framing and blurriness on the picture.</p>
</blockquote>
<div class="quarto-figure quarto-figure-center">
<figure class="figure">
<p><img src="https://ducha-aiki.github.io/wide-baseline-stereo-blog/posts/2022-02-05-DISC21-recap_files/att_00002.png" class="img-fluid figure-img"></p>
<figcaption>Test set Phase 2 transformations. Taken from DISC paper</figcaption>
</figure>
</div>
<p>The paper also provides several baselines, mostly based on global and local descriptors</p>
<div class="quarto-figure quarto-figure-center">
<figure class="figure">
<p><img src="https://ducha-aiki.github.io/wide-baseline-stereo-blog/posts/2022-02-05-DISC21-recap_files/att_00003.png" class="img-fluid figure-img"></p>
<figcaption>Baselines, provided by the challenge organizers</figcaption>
</figure>
</div>
</section>
</section>
<section id="tracks-descriptor-track-and-matching-track" class="level2">
<h2 class="anchored" data-anchor-id="tracks-descriptor-track-and-matching-track">Tracks: Descriptor track and Matching track</h2>
<p>The competition had 2 tracks.</p>
<ol type="1">
<li><p><strong>Descriptor</strong> track expects participant to submit global image descriptors, up to 256-dim, which should be compared by L2 distance. Matching image pairs should have small descriptor distances, while non-matching - large.</p></li>
<li><p><strong>Matching</strong> track, where the model can do any pairwise (but only pairwise!) comparisons to get the similarity score between two images.</p></li>
</ol>
<p>Below I will overview the winning solutions, presented at the <a href="https://sites.google.com/view/isc2021">NeurIPS2021 workshop</a>. But first let me show you a couple of slides from the <a href="https://drive.google.com/file/d/1Jcw9NWpBJ65A3RM5MaHB9-AKG_rM6-Wt/view">organizer’s analisys</a>.</p>
<p>The easiest image transformations are geometrical ones, except cropping the image to a very small size. The hardest ones are those, which involve combination of several images. The global descriptors are very sensitive to image rotatio and vertical flip (not really surprising).</p>
<div class="quarto-figure quarto-figure-center">
<figure class="figure">
<p><img src="https://ducha-aiki.github.io/wide-baseline-stereo-blog/posts/2022-02-05-DISC21-recap_files/att_00004.png" class="img-fluid figure-img"></p>
<figcaption>Easy-to-hard transformations. From organaizer’s analysis</figcaption>
</figure>
</div>
<p>The models are more or less reasonably robust to adversarial attacks, except really strong ones.</p>
<div class="quarto-figure quarto-figure-center">
<figure class="figure">
<p><img src="https://ducha-aiki.github.io/wide-baseline-stereo-blog/posts/2022-02-05-DISC21-recap_files/att_00005.png" class="img-fluid figure-img"></p>
<figcaption>Models sensitivity to the adversarial attacks</figcaption>
</figure>
</div>
<div class="quarto-figure quarto-figure-center">
<figure class="figure">
<p><img src="https://ducha-aiki.github.io/wide-baseline-stereo-blog/posts/2022-02-05-DISC21-recap_files/att_00006.png" class="img-fluid figure-img"></p>
<figcaption>Examples of the adversarial attacks</figcaption>
</figure>
</div>
<p>Finally organizers found that have not done enough de-duplication in the dataset, meaning that models were able to discover the duplicates in the dataset. While some of this annotation errors were genuine, others were in the gray zone.</p>
<p>For example, images, which were taken from the different positions, but the background object DOES look identical. Another example is the camera burst, where the part of the content, which is not the same across the burst images ended up cropped out by annotators or automatic procedure.</p>
<div class="quarto-figure quarto-figure-center">
<figure class="figure">
<p><img src="https://ducha-aiki.github.io/wide-baseline-stereo-blog/posts/2022-02-05-DISC21-recap_files/att_00007.png" class="img-fluid figure-img"></p>
<figcaption>Images taken from different position, but with the object far away</figcaption>
</figure>
</div>
<div class="quarto-figure quarto-figure-center">
<figure class="figure">
<p><img src="https://ducha-aiki.github.io/wide-baseline-stereo-blog/posts/2022-02-05-DISC21-recap_files/att_00008.png" class="img-fluid figure-img"></p>
<figcaption>Image bursts, which are not de-duplicated</figcaption>
</figure>
</div>
</section>
<section id="descriptor-track" class="level2">
<h2 class="anchored" data-anchor-id="descriptor-track">Descriptor track</h2>
<section id="st-place-contrastive-learning-with-large-memory-bank-and-negative-embedding-subtraction-for-accurate-copy-detection" class="level3">
<h3 class="anchored" data-anchor-id="st-place-contrastive-learning-with-large-memory-bank-and-negative-embedding-subtraction-for-accurate-copy-detection">1st place: Contrastive Learning with Large Memory Bank and Negative Embedding Subtraction for Accurate Copy Detection</h3>
<p>by Shuhei Yokoo aka lyakaap.</p>
<p>The method takes ImageNet-21k-pretrained EfficientNetv2 from <a href="https://fastai.github.io/timmdocs/">timm</a> and trains it with contrastive loss, gradually increasing the augmentations intensity, as well as image size. Final augmentations are quite strong, see the image below. The loss is the simple constrastive loss with memory bank for the hard negatives.</p>
<div class="quarto-figure quarto-figure-center">
<figure class="figure">
<p><img src="https://ducha-aiki.github.io/wide-baseline-stereo-blog/posts/2022-02-05-DISC21-recap_files/att_00009.png" class="img-fluid figure-img"></p>
<figcaption>Examples of augnentations by lyakaap team</figcaption>
</figure>
</div>
<div class="quarto-figure quarto-figure-center">
<figure class="figure">
<p><img src="https://ducha-aiki.github.io/wide-baseline-stereo-blog/posts/2022-02-05-DISC21-recap_files/att_00010.png" class="img-fluid figure-img"></p>
<figcaption>Ablation study of the different steps</figcaption>
</figure>
</div>
<p>Finally, Shuhei Yokoo proposes novel postprocessing technique - to subtract hard negative embedding from all database image embeddings. It added nice final boost to the submission.</p>
<p><a href="https://arxiv.org/pdf/2112.04323.pdf">arXiv</a> <a href="https://github.com/lyakaap/ISC21-Descriptor-Track-1st">github</a></p>
</section>
<section id="nd-place-producing-augmentation-invariant-embeddings-from-real-life-imagery" class="level3">
<h3 class="anchored" data-anchor-id="nd-place-producing-augmentation-invariant-embeddings-from-real-life-imagery">2nd place: Producing augmentation-invariant embeddings from real-life imagery</h3>
<p>by Sergio Manuel Papadakis and Sanjay Addicam</p>
<p>The method relies on the GeM descriptors ( trained EfficientNetv1, EfficientNetv2, NfNet) with ArcFace loss. ArcFace loss assumes classification, not contrastive setup and hard to train with &gt;1M classes. That is why the method gradually increases the number of classes (images used to train) and the classifier head is re-initialized with centroids, obtained from backbone descriptors from the previous stage. Multiple models are combined by Concat -&gt; PCA -&gt; L2Norm.</p>
<div class="quarto-figure quarto-figure-center">
<figure class="figure">
<p><img src="https://ducha-aiki.github.io/wide-baseline-stereo-blog/posts/2022-02-05-DISC21-recap_files/att_00011.png" class="img-fluid figure-img"></p>
<figcaption>2nd place model architecture</figcaption>
</figure>
</div>
<p>Finally, the embeddings are modified, to take into account the most similar images from the database, similar to the 1st place, but in a bit more complicated way.</p>
<p><a href="https://arxiv.org/pdf/2112.03415.pdf">arXiv</a>, <a href="https://github.com/socom20/facebook-image-similarity-challenge-2021">github</a></p>
</section>
<section id="rd-place-bag-of-tricks-and-a-strong-baseline-for-image-copy-detection" class="level3">
<h3 class="anchored" data-anchor-id="rd-place-bag-of-tricks-and-a-strong-baseline-for-image-copy-detection">3rd place: Bag of Tricks and A Strong baseline for Image Copy Detection</h3>
<p>by Wenhao Wang, Weipu Zhang, Yifan Sun, Yi Yang</p>
<p>The method uses ResNet50 backbone, pretrained in augmentation-supervised fashion (<a href="https://arxiv.org/abs/2103.03230">BarlowTwins</a>, unlike 1st and 2nd place, which used supervised pretrained EfficientNetv2. The model is trained with a combination of triplet and cross entropy losses with GeM pooling and WaveBlock. Finally, authors also present hard-negative aware embedding de-normalization called “descriptor stretching” – different variant of the similar idea, also used in 1st and 2nd place post-processing.</p>
<div class="quarto-figure quarto-figure-center">
<figure class="figure">
<p><img src="https://ducha-aiki.github.io/wide-baseline-stereo-blog/posts/2022-02-05-DISC21-recap_files/att_00013.png" class="img-fluid figure-img"></p>
<figcaption>Training pipeline of the 3rd place in descriptor track</figcaption>
</figure>
</div>
<p>Note that before descriptor stretching, results are much weaker than in 1st and 2nd place, but after the post-processing, the margin shrinks</p>
<div class="quarto-figure quarto-figure-center">
<figure class="figure">
<p><img src="https://ducha-aiki.github.io/wide-baseline-stereo-blog/posts/2022-02-05-DISC21-recap_files/att_00014.png" class="img-fluid figure-img"></p>
<figcaption>The ablation studye. Supervised or Unsupervised denotes the supervised pre-training or unsupervised pre-training is used, respectively. Des-Str denotes adding descriptor stretching strategy. Det denotes adding overlay detection YOLOv5. Multi denotes using multi-scale testing.</figcaption>
</figure>
</div>
<p><a href="https://arxiv.org/abs/2111.08004">arXiv</a>, <a href="https://github.com/WangWenhao0716/ISC-Track2-Submission">github</a></p>
</section>
</section>
<section id="matching-track" class="level2">
<h2 class="anchored" data-anchor-id="matching-track">Matching track</h2>
<section id="st-place-d2lv-a-data-driven-and-local-verification-approach-for-image-copy-detection" class="level3">
<h3 class="anchored" data-anchor-id="st-place-d2lv-a-data-driven-and-local-verification-approach-for-image-copy-detection">1st place D2LV: A Data-Driven and Local-Verification Approach for Image Copy Detection</h3>
<p>by Wenhao Wang, Yifan Sun, Weipu Zhang, Yi Yang (3rd place in descriptor track).</p>
<p>The method relies on several ideas:</p>
<ul>
<li>Training specialist models on the different augmentations, e.g.&nbsp;“basic”, “basic+superdark”, “basic+superblur”, etc.</li>
<li>Different ways of splitting the image into the subparts: by YOLOv5 object detection, by 4x4 even split, by rotating the original image, etc, each of which is then “matched” by the global descriptor similarity to the right-hand image. Local-to-global means “we aplit query and match vs reference image”, whereas “global-to-local” - - “we match whole query image to splitted reference”.</li>
</ul>
<div class="quarto-figure quarto-figure-center">
<figure class="figure">
<p><img src="https://ducha-aiki.github.io/wide-baseline-stereo-blog/posts/2022-02-05-DISC21-recap_files/att_00015.png" class="img-fluid figure-img"></p>
<figcaption>Winning solution pipeline</figcaption>
</figure>
</div>
<div class="quarto-figure quarto-figure-center">
<figure class="figure">
<p><img src="https://ducha-aiki.github.io/wide-baseline-stereo-blog/posts/2022-02-05-DISC21-recap_files/att_00016.png" class="img-fluid figure-img"></p>
<figcaption>The ablation study D2LV. Supervised and Unsupervised denotes the supervised pre-training or unsupervised pre-training is used, respectively. Global-local – using global-local matching strategy, and Both denotes using global-local and local-global matching strategy. Adv Aug denotes using different kinds of advanced augmentations. Multi-Tricks denotes using multi-scale testing and other men- tioned tricks.</figcaption>
</figure>
</div>
<p><a href="https://arxiv.org/pdf/2111.07090.pdf">arXiv</a> <a href="https://github.com/WangWenhao0716/ISC-Track1-Submission">github</a></p>
</section>
<section id="nd-place-2nd-place-solution-to-facebook-ai-image-similarity-challenge-matching-track" class="level3">
<h3 class="anchored" data-anchor-id="nd-place-2nd-place-solution-to-facebook-ai-image-similarity-challenge-matching-track">2nd place: 2nd Place Solution to Facebook AI Image Similarity Challenge: Matching Track</h3>
<p>by SeungKee Jeon.</p>
<p>The paper uses a simple approach of concatenating query and reference image and feeding such collage into the vision transformer (ViT).</p>
<p>ViT embeddings are also used to create a short-list for the matching, trained with contrastive loss (SimCLR).</p>
<p>Similarly to the 1st place, SeungKee Jeon also used image-to-sub-images (1 -&gt; 2x2, 1-&gt;3x3) splitting technique at test time.</p>
<div class="quarto-figure quarto-figure-center">
<figure class="figure">
<p><img src="https://ducha-aiki.github.io/wide-baseline-stereo-blog/posts/2022-02-05-DISC21-recap_files/att_00017.png" class="img-fluid figure-img"></p>
<figcaption>Matching Track model training pipeline. Reference and query image is concatenated to form as one image and ask ViT network to predict if query image used the reference image. Label is 1 if train image A is used as reference image and its augmented image is used as query image. Label is 0 if train image A is used as reference image and augmented version of other train image C is used as query image.</figcaption>
</figure>
</div>
<p><a href="https://arxiv.org/abs/2111.09113">arXiv</a> <a href="https://github.com/seungkee/2nd-place-solution-to-Facebook-Image-Similarity-Matching-Track">github</a></p>
</section>
<section id="rd-place-a-global-and-local-dual-retrieval-solution-to-facebook-ai-image-similarity-challenge" class="level3">
<h3 class="anchored" data-anchor-id="rd-place-a-global-and-local-dual-retrieval-solution-to-facebook-ai-image-similarity-challenge">3rd place: A Global and Local Dual Retrieval Solution to Facebook AI Image Similarity Challenge</h3>
<p>by Xinlong Sun, Yangyang Qin, Xuyuan Xu, Guoping Gong, Yang Fang, Yexin Wang</p>
<p>This method is probably the most hand-engineered.</p>
<div class="quarto-figure quarto-figure-center">
<figure class="figure">
<p><img src="https://ducha-aiki.github.io/wide-baseline-stereo-blog/posts/2022-02-05-DISC21-recap_files/att_00018.png" class="img-fluid figure-img"></p>
<figcaption>Overall pipeline, which consist of separate overlay-detection step, image retrieval with global EsViT and local descriptors - SIFT and finally their combination</figcaption>
</figure>
</div>
<ol type="1">
<li>The images are pre-processed with overlay-detection – out-of-the-box YOLOv5. Detected overlays are treated as a separate images.</li>
<li>The global descriptor EsViT (Swin-B) is pretrained in self-supervised fashion on ImageNet, then with contrastive loss on the challange dataset, with hard-negative memory bank.</li>
</ol>
<p><img src="https://ducha-aiki.github.io/wide-baseline-stereo-blog/posts/2022-02-05-DISC21-recap_files/att_00019.png" class="img-fluid"></p>
<ol start="3" type="1">
<li>Resize images to 300px -&gt; SIFT features (~600 per image) -&gt; 600M descriptors. Use brute-force search in faiss with fp16 precision. That takes 165Gb memory, fit with 8 V100 GPUs. Threshold the descriptor distance first, the number of matched features second.</li>
</ol>
<p>According to the ablation study, SIFT was significantly worse, than global EsViT descriptor in Phase I, but significantly boosted overall score in Phase 2.</p>
<div class="quarto-figure quarto-figure-center">
<figure class="figure">
<p><img src="https://ducha-aiki.github.io/wide-baseline-stereo-blog/posts/2022-02-05-DISC21-recap_files/att_00020.png" class="img-fluid figure-img"></p>
<figcaption>Ablation study. G-emb means global embedding, F means finetune, and P means using Preprocessing Module.</figcaption>
</figure>
</div>
<p><a href="https://arxiv.org/abs/2112.02373">arXiv</a>, <a href="https://github.com/sun-xl/ISC2021">github</a></p>
</section>
</section>
<section id="conclusion-my-speculations" class="level2">
<h2 class="anchored" data-anchor-id="conclusion-my-speculations">Conclusion &amp; my speculations</h2>
<p>Edited image copy detection is a challenging task, which is, nevertheless solvable to some practical extent with existing techniques. It is interesting, the all 3 winners in descriptor track come up with a similar idea of contextualized database image embeddings with hard negatives.</p>
<p>I also don’t believe that the supervised/unsupervised pretraining and transformer/CNN architecture matter as much, as participants were reporting – precisely because different teams succesfully used different architectures and pretraining. Once one optimized the training procedure for the selected architecture, the differences would probably disappear.</p>
<p>The matching track pipelines are quite complex and computationally demanding. Nevertheless, using local features (or at least sub-image matching) noticably improves the score. My personal favorite is the 2nd place in the matching track for its originality and simplicity.</p>
<p>What would be interesting, is to check how much one looses when using visual-words tf-idf encoding for the local features matching, instead of keeping original descriptors.</p>
<p>Challenge organizers also announced the join paper with winners, overviewing the competition and the best approaches some when in February (i.e.&nbsp;soon). Stay tuned :)</p>


</section>

<p>Everything you (didn’t) want to know about image matching</p> ]]></description>
  <guid>https://ducha-aiki.github.io/wide-baseline-stereo-blog/posts/2022-02-05-disc21-recap.html</guid>
  <pubDate>Fri, 04 Feb 2022 23:00:00 GMT</pubDate>
  <media:content url="https://ducha-aiki.github.io/wide-baseline-stereo-blog/posts/images/spiderman-meme.jpg" medium="image" type="image/jpeg"/>
</item>
<item>
  <title>Review of On the Limits of Pseudo Ground Truth in Visual Camera Re-localisation</title>
  <dc:creator>Dmytro Mishkin</dc:creator>
  <link>https://ducha-aiki.github.io/wide-baseline-stereo-blog/posts/2021-09-04-review-of-pseudogt-limitations.html</link>
  <description><![CDATA[ 





<p>I would like to share my thoughts on ICCV 2021 paper “<a href="https://arxiv.org/abs/2109.00524">On the Limits of Pseudo Ground Truth in Visual Camera Re-localisation</a>” by Brachmann et.al. But first let’s recap, what is Visual relocalization and how it is usually evaluated.</p>
<section id="visual-re--localization-short-recap" class="level1">
<h1>Visual (re-) localization: short recap</h1>
<p>We have a robot with RGB or RGB-D camera. Our robot have been in some place once and have been explored it, taking the pictures. After some time, we bring it there again and turned on. Our robot (or, to be precise, the localization algorithm, which robot uses) task is to estimate its position and orientation from a single photo. How do we measure if it does it correctly?</p>
<ol type="1">
<li><p>We obtain two sequences of the images by going around our place. One is “<em>training</em>” or “<em>database</em>” sequence. The second is “<em>test</em>” sequence, which is emulates robot going there, wanting to know where it is.</p></li>
<li><p>We run some known-to-be-reliable algorithm such as SfM or SLAM on those image sequences to jointly obtain the camera poses for all the images. Sometimes we can use additional sensors to get the pose directly and avoid such estimation completely. However, that is not the case we are discussing now.</p></li>
<li><p>We provide the <em>database</em> images together with their poses to the algorithm we want to evaluate. It can do whatever it wants with them.</p></li>
<li><p>Finally, we give the algorithm an image from the <em>test</em> sequence <em>without a pose</em> and algorithm estimate its camera pose. Then we compare this pose to the pose, which we receive with the pose we have and rank algorithms by the errors in the pose.</p></li>
</ol>
<p>There are different kinds of algorithms for visual localization. One possible way to first find the most similar iamges from our dataset to the query image, then match query to them and finally calculate the pose, given the poses of databased images. Another way is to train a model on a database images, such that for each pixel or region it tried to directly estimate its 3D location. Then based on the image content location, we deduce the camera pose.</p>
</section>
<section id="what-is-the-paper-about" class="level1">
<h1>What is the paper about?</h1>
<p>The idea of the paper is very simple and yet nobody asked its main question before. When we benchmark our image matching or camera relocalization algorithms, we do not really have ground truth data for the camera pose, as we just recaped. Instead we rely on some sort of known-to-be-reliable algorithm to infer this information from (many) images. This process is called “pseudo Ground Truth (pGT)” in the paper.</p>
<p><em>Does it affect results to be in favor of the (family of) methods, which were used to generate the pseugo GT?</em></p>
<div class="quarto-figure quarto-figure-center">
<figure class="figure">
<p><img src="https://ducha-aiki.github.io/wide-baseline-stereo-blog/posts/2021-09-04-review-of-pseudoGT-limitations_files/att_00000.png" class="img-fluid figure-img"></p>
<figcaption>Figure 1 from paper. Visualisation of the same scene reconstructed by two different reference algorithms</figcaption>
</figure>
</div>
<section id="case-study-image-matching-challenge" class="level3">
<h3 class="anchored" data-anchor-id="case-study-image-matching-challenge">Case-study: Image Matching Challenge</h3>
<p>Let me give you an example.</p>
<p>Our Image Matching Challenge 2019, <a href="https://ducha-aiki.github.io/wide-baseline-stereo-blog/posts/2021-05-14-IMC2020-competition-recap.html">2020</a> <a href="https://ducha-aiki.github.io/wide-baseline-stereo-blog/posts/2021-05-12-submitting-to-IMC2021-step-by-step.html">2021</a> does exactly this: the camera pose are reconstructed with <a href="https://colmap.github.io">COLMAP</a> 3d reconstruction software given thousands of images. COLMAP is based on SIFT local features and LO-RANSAC for the pose stimation. <img src="https://ducha-aiki.github.io/wide-baseline-stereo-blog/posts/2021-09-04-review-of-pseudoGT-limitations_files/att_00001.png" class="img-fluid" alt="Pseudo-ground truth generation and benchmark pipeline for the Image Matching Challenge"></p>
<p>Can we use it to benchmark other local features and RANSACs in a fair way? That is the question, reviewers of our paper asked us. We answered this question by running the 3d reconstruction with very different local features: <a href="https://arxiv.org/abs/1906.06195">R2D2</a>, <a href="https://www.cs.ubc.ca/~lowe/papers/ijcv04.pdf">SIFT</a>, DoG-<a href="https://arxiv.org/abs/1705.10872.pdf">HardNet</a>, <a href="https://arxiv.org/abs/1712.07629">SuperPoint</a>. Then we compared the difference in the poses produced by the different methods is very small and much smaller than the accuracy threshold we used for evaluation.</p>
<div class="quarto-figure quarto-figure-center">
<figure class="figure">
<p><img src="https://ducha-aiki.github.io/wide-baseline-stereo-blog/posts/2021-09-04-review-of-pseudoGT-limitations_files/att_00002.png" title="Histograms of pose differences between reconstructions with different local feature methods" class="img-fluid figure-img"></p>
<figcaption>image.png</figcaption>
</figure>
</div>
<p>So we can safely use our benchmark to evaluate two view matching methods. <em>Can we really?</em></p>
</section>
</section>
<section id="sfm-versus-dense-slam" class="level1">
<h1>SfM versus Dense SLAM</h1>
<p>The “<a href="https://arxiv.org/abs/2109.00524">On the Limits…</a>” asks the different question, than the one we have answered in IMC paper. <em>While, the GT might be not biased towards specific local feature, it can be biased towards the higher-level algorithm – Structure-from-Motion (SfM)</em>. But the (sparse) SfM is not the only way how we can obtain camera poses from the collection of images – one could use (dense) Depth-based SLAM. Are the poses we got from them the same?</p>
<p>The answer is: <strong>no</strong>, they are not the same!</p>
<div class="quarto-figure quarto-figure-center">
<figure class="figure">
<p><img src="https://ducha-aiki.github.io/wide-baseline-stereo-blog/posts/2021-09-04-review-of-pseudoGT-limitations_files/att_00003.png" class="img-fluid figure-img"></p>
<figcaption>Camera trajectory of the training set, estimated by SfM and RGB-D-SLAM</figcaption>
</figure>
</div>
<p>The paper, however, does not stop here, it studies the following question: if the camera poses (think of them as “labels”) of the <em>training</em> set were obtained with a different family of algorithms, does it influence their performance?</p>
<p>The answer is <strong>yes</strong>. Look at the image below, which shows the trajectories, estimated by the relocalization methods, depending on the (pseudo) ground truth source.</p>
<div class="quarto-figure quarto-figure-center">
<figure class="figure">
<p><img src="https://ducha-aiki.github.io/wide-baseline-stereo-blog/posts/2021-09-04-review-of-pseudoGT-limitations_files/Sge18CpO.jpg" class="img-fluid figure-img"></p>
<figcaption>Camera trajectory of the test set, estimated by visual relocalization methods</figcaption>
</figure>
</div>
<section id="my-take-on-the-paper." class="level2">
<h2 class="anchored" data-anchor-id="my-take-on-the-paper.">My take on the paper.</h2>
<p>The paper asks a great question and provides good answers. It is also well-written and the experiments done with rigor. OK, time to stop “CVPR-review” mode and provide my thoughts.</p>
<p>I love the paper. However, (post-hoc) I believe that paper should stop earlier or go deeper.</p>
<p>First and main message of the paper (which paper does NOT formulate in that way):</p>
<p><strong>The combination of the current methods (SfM or SLAM) with a current way of generating training sequence (single camera going around the place once) does not lead to the robust and precise enough camera poses</strong>.</p>
<p>I don’t buy the paper’s argument that “depending on the metric you use, one or another is better”. No.&nbsp;There are only 3 possible explanations:</p>
<ol type="a">
<li><p>SfM estimates wrong poses from the sequence</p></li>
<li><p>RGBD-SLAM estimates wrong poses from the sequence</p></li>
<li><p>both SfM and RGBD-SLAM estimate wrong poses. The variant: noone cannot estimate precise poses from the given images.</p></li>
</ol>
<p>We are not in the quantum mechanics world, there is a single, well-defined (possibly unknown) correct camera pose. The correct pose exists and the question is whether SfM or SLAM (or neither) is closer to it. Yes, it would require some additional sensors, or running in simulated world, but anyway - the message should not be “just provide 2 versions of pGT”. The message is: if your 2 versions of pGT does not agree, GET MORE/BETTER DATA.</p>
<p>In addition to that I would like to see a paper, which answers the following question: given tha data we have (i.e.&nbsp;bad data) but with also available real GT camera poses, under which conditions, which algorithm (SfM or Depth SLAM) provides better results?</p>
<section id="which-algorithm-can-adapt-to-wrong-training-data-and-be-consistent-with-it" class="level3">
<h3 class="anchored" data-anchor-id="which-algorithm-can-adapt-to-wrong-training-data-and-be-consistent-with-it">Which algorithm can adapt to (wrong) training data and be consistent with it?</h3>
<p>One thing, which paper focuses on is how the “test” pGT is generated. However, even the bigger impact can be from the fact that <em>training</em> data may conflict with the poses, that algorithm estimates.</p>
<p>Imagine the augmented reality scenario, where one needs to draw the (virtual) teddy bear on the coach. In that case we actually don’t care if the camera pose estimate is correct, or if our model of the room is correct. What we care about, is how naturally looking will be position of the Teddy bear (or Pikachu). It also important that bear not jumps around when we move our camera slightly.</p>
<p>Given the images and results in the paper, it seems that (learning-based) DSAC can suits this much better than (non-learning) AS. That is also stated in the paper conclusions. In the same time, it seems that the variance of the pose estimates is smaller for the SfM pGT than for the SLAM even for DSAC.</p>
<p>However, I would also like to see a paper, which studies how the error/bias/variance in the training data affects the different algorithms, so they could adapt to use bad data better, in more detail.</p>
</section>
</section>
</section>
<section id="conclusion" class="level1">
<h1>Conclusion</h1>
<p>The paper is thought-provoking and I hope our community will read it and think about it. We also need to pay a way more attention to our data-gathering practice, because it really makes a difference.</p>
<p>P.S. Thanks to <a href="https://twitter.com/eric_brachmann">Eric</a> for sending me additional visualizations, which were not in the arXiv paper.</p>


</section>

<p>Everything you (didn’t) want to know about image matching</p> ]]></description>
  <guid>https://ducha-aiki.github.io/wide-baseline-stereo-blog/posts/2021-09-04-review-of-pseudogt-limitations.html</guid>
  <pubDate>Fri, 03 Sep 2021 22:00:00 GMT</pubDate>
  <media:content url="https://ducha-aiki.github.io/wide-baseline-stereo-blog/posts/images/get_better_data.png" medium="image" type="image/png" height="183" width="144"/>
</item>
<item>
  <title>WxBS: Relaunching challenging benchmark for Image Matching</title>
  <dc:creator>Dmytro Mishkin</dc:creator>
  <link>https://ducha-aiki.github.io/wide-baseline-stereo-blog/posts/2021-07-30-reviving-wxbs-benchmark.html</link>
  <description><![CDATA[ 





<section id="the-hardest-image-matching-benchmark" class="level2">
<h2 class="anchored" data-anchor-id="the-hardest-image-matching-benchmark">The hardest image matching benchmark</h2>
<p>Previously we have discussed <a href="https://ducha-aiki.github.io/wide-baseline-stereo-blog/posts/2021-01-09-wxbs-in-simple-terms.html">what is WxBS in general</a>, <a href="https://ducha-aiki.github.io/wide-baseline-stereo-blog/posts/2020-08-06-affine-view-synthesis.html">how to match images from a really different viewpoints</a> and also took a look into <a href="https://ducha-aiki.github.io/wide-baseline-stereo-blog/posts/2021-05-14-IMC2020-competition-recap.html">Image Matching Challenge 2020</a>. Let’s now speak about how to measure progress on images, which are really hard to match even to human?</p>
<p>For example, which have been taked in day-vs-night, years apart AND from different camera positions? One cannot register them to some 3d model, because they are too hard. And there is no 3D model either.</p>
<div class="quarto-figure quarto-figure-center">
<figure class="figure">
<p><img src="https://ducha-aiki.github.io/wide-baseline-stereo-blog/posts/2021-07-28-Reviving-WxBS-benchmark_files/att_00001.png" class="img-fluid figure-img"></p>
<figcaption>Kyiv, Ukraine. Image pair from WxBS dataset</figcaption>
</figure>
</div>
<p>Well, in that case one could go back to basics and <em>handlabel</em> the correspondences between the images. Such annotation is unlikely to be very precise, but it is going to be quite robust, if the labeler is careful.</p>
<p>That is what I have done 6 years ago and published such a dataset at <a href="http://www.bmva.org/bmvc/2015/papers/paper012/index.html">BMVC 2015</a> together with a benchmark of popular image matchers of that time. None of them was able to successfully match even a fraction of the WxBS dataset.</p>
<div class="quarto-figure quarto-figure-center">
<figure class="figure">
<p><img src="https://ducha-aiki.github.io/wide-baseline-stereo-blog/posts/2021-07-28-Reviving-WxBS-benchmark_files/att_00002.png" class="img-fluid figure-img"></p>
<figcaption>Handlabeled correspondences in pixelstitch. The image in the bottom shows the correspondences and epipolar lines induced by them.</figcaption>
</figure>
</div>
<p>However, as I was quite inexpereinced in product packaging, as well as in software engineering, the benchmark existed as dataset + bunch of bash and Matlab scripts. It is needless to say, that nobody evaluated their algorithms on it. That made me sad, however, I never found time to make this benchmark easy to use. Well, now I did it and want to share the evaluation of the recent agorithms on such a hard data.</p>
</section>
<section id="not-that-fast-problem-with-a-fundamental-matrix-and-how-to-solve-them" class="level2">
<h2 class="anchored" data-anchor-id="not-that-fast-problem-with-a-fundamental-matrix-and-how-to-solve-them">Not that fast! Problem with a fundamental matrix and how to solve them</h2>
<p>OK, we have ground truth correspondences. Now what? Ideally one can estimate epipolar geometry from as little, as 7 correspondences, but in practice? No.&nbsp;The problem is that there could be many relative camera poses, which are consistent with the correspondences, especially when correspondences lie in plane and the camera calibration is unknown.</p>
<p>Let’s consider a simple example: we have 13 correspondences and they seems to be correct. Are they enough for finding good epipolar geometry? This question can be answered via leave-one-out validation procedure, described in <a href="https://arxiv.org/abs/2106.10240">VSAC paper</a>. We remove a single correspondence, estimate a fundamental matrix via DLT algorithm on the rest of the correspondences and see how the estimated epipolar geometry is changed (or not).</p>
<p>As you can see on the animation below, the epipolar geometry estimation from our 13 correspondences is not stable and removing any of them results in noticable change.</p>
<p><img src="https://ducha-aiki.github.io/wide-baseline-stereo-blog/posts/2021-07-28-Reviving-WxBS-benchmark_files/epi1.gif" class="img-fluid"></p>
<p>So, if even manually labeled corresposndences are not good enough, what can be done?</p>
<p>First, simply add more correspondences, which are spread across image AND depth levels as even as possible. That is what I did recently, in order to make WxBS labeling better. However, it is not always possible – sometimes you have lots of correspondence on “faraway background plane” and very little somewhere on foreground, see next picture.</p>
<p><img src="https://ducha-aiki.github.io/wide-baseline-stereo-blog/posts/2021-07-28-Reviving-WxBS-benchmark_files/epi_harder.gif" class="img-fluid"></p>
<p>What else can be done? We can go other way round and <em>not provide any ground truth fundamental matrix at all</em>. Instead, we ask methods, which we evaluate, to give us such matrix <em>F_est</em>, then check, how many ground truth correspondences are consident with it.</p>
<p>The bad thing about it, is that unless all 100% GT correspondences are consistent with <em>F_est</em>, <em>F_est</em> is completely wrong in terms of camera pose accuracy. So, such a binary signal on a very challenging pairs is not that useful. The good thing about it, is that if, say, 50% GT correspondences are consistent with <em>F_est</em>, it means that detected (not GT) correspondences <em>on the part of the image</em> are correct (e.g.&nbsp;on some dominant plane) and at least there we are OK.</p>
<p>If you are interested in more mathematical definition, here it is:</p>
<p>The recall on ground truth correspondences <img src="https://latex.codecogs.com/png.latex?C_i"> of image pair <img src="https://latex.codecogs.com/png.latex?i"> and for geometry model <img src="https://latex.codecogs.com/png.latex?%5Cmathbf%7BM%7D_i"> is computed as a function of a threshold <img src="https://latex.codecogs.com/png.latex?%5Ctheta"></p>
<p><img src="https://latex.codecogs.com/png.latex?%5Cbegin%7Bequation%7D%0A%5Cmathrm%7Br%7D_%7Bi,%5Cmathbf%7BM%7D_i%7D(%5Ctheta)%20=%20%5Cfrac%7B%7C%20%5C%7B(%5Cmathbf%7Bu%7D,%5Cmathbf%7Bv%7D)%20:%20(%5Cmathbf%7Bu%7D,%5Cmathbf%7Bv%7D)%20%5Cin%20C_i,%20%20e(%5Cmathbf%7BM%7D_i,%5Cmathbf%7Bu%7D,%5Cmathbf%7Bv%7D)%20%3C%20%5Ctheta%5C%7D%7C%7D%7B%7C%20C_i%20%7C%7D%0A%25%20%5Cmathrm%7Br%7D_%7Bi,%5Cmathbf%7BM%7D_i%7D(%5Ctheta)%20=%20%5Cfrac%7B%7C%20%5C%7B(%5Cmathbf%7Bu%7D_i,%5Cmathbf%7Bv%7D_i)%20:%20(%5Cmathbf%7Bu%7D_i,%5Cmathbf%7Bv%7D_i)%20%5Cin%20C_i,%20%20e(%5Cmathbf%7BM%7D_i,%5Cmathbf%7Bu%7D,%5Cmathbf%7Bv%7D)%20%3C%20%5Ctheta%5C%7D%7C%7D%7B%7C%20C_i%20%7C%7D%0A%5Cend%7Bequation%7D"></p>
<p>using <a href="https://kornia.readthedocs.io/en/latest/geometry.epipolar.html#kornia.geometry.epipolar.symmetrical_epipolar_distance">symmetric epipolar distance</a>.</p>
</section>
<section id="lets-evaluate" class="level2">
<h2 class="anchored" data-anchor-id="lets-evaluate">Let’s evaluate!</h2>
<p>First I will show you how to run the WxBS benchmark yourself and then present the results I got. Benchmark is available from pip:</p>
<div id="cell-4" class="cell">
<div class="sourceCode cell-code" id="cb1" style="background: #f1f3f5;"><pre class="sourceCode python code-with-copy"><code class="sourceCode python"><span id="cb1-1"><span class="op" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">!</span>pip install wxbs_benchmark</span>
<span id="cb1-2"><span class="op" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">!</span>pip install kornia_moons</span></code></pre></div>
</div>
<p>Now we will write a simple function using OpenCV RootSIFT and <a href="https://arxiv.org/abs/1912.05909">MAGSAC++</a>. RootSIFT is a gold standard handcrafted local feature and MAGSAC++ is a state-of-the-art RANSAC, which is not sensitive to an inlier threshold, as <a href="https://ducha-aiki.github.io/wide-baseline-stereo-blog/posts/2021-05-17-OpenCV-New-RANSACs.html">my recent benchmark</a> shows. We will match then using mutual second nearest ratio, as implemented in <a href="https://github.com/kornia/kornia">kornia</a></p>
<div id="cell-6" class="cell">
<div class="sourceCode cell-code" id="cb2" style="background: #f1f3f5;"><pre class="sourceCode python code-with-copy"><code class="sourceCode python"><span id="cb2-1"><span class="im" style="color: #00769E;
background-color: null;
font-style: inherit;">import</span> matplotlib.pyplot <span class="im" style="color: #00769E;
background-color: null;
font-style: inherit;">as</span> plt</span>
<span id="cb2-2"><span class="im" style="color: #00769E;
background-color: null;
font-style: inherit;">import</span> numpy <span class="im" style="color: #00769E;
background-color: null;
font-style: inherit;">as</span> np</span>
<span id="cb2-3"><span class="im" style="color: #00769E;
background-color: null;
font-style: inherit;">import</span> cv2</span>
<span id="cb2-4"><span class="im" style="color: #00769E;
background-color: null;
font-style: inherit;">import</span> torch</span>
<span id="cb2-5"><span class="im" style="color: #00769E;
background-color: null;
font-style: inherit;">import</span> kornia.feature <span class="im" style="color: #00769E;
background-color: null;
font-style: inherit;">as</span> KF</span>
<span id="cb2-6"><span class="im" style="color: #00769E;
background-color: null;
font-style: inherit;">from</span> kornia_moons.feature <span class="im" style="color: #00769E;
background-color: null;
font-style: inherit;">import</span> <span class="op" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">*</span></span>
<span id="cb2-7"><span class="im" style="color: #00769E;
background-color: null;
font-style: inherit;">from</span> wxbs_benchmark.dataset <span class="im" style="color: #00769E;
background-color: null;
font-style: inherit;">import</span> <span class="op" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">*</span></span>
<span id="cb2-8"><span class="im" style="color: #00769E;
background-color: null;
font-style: inherit;">from</span> wxbs_benchmark.evaluation <span class="im" style="color: #00769E;
background-color: null;
font-style: inherit;">import</span> <span class="op" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">*</span></span>
<span id="cb2-9"><span class="im" style="color: #00769E;
background-color: null;
font-style: inherit;">from</span> tqdm <span class="im" style="color: #00769E;
background-color: null;
font-style: inherit;">import</span> tqdm </span>
<span id="cb2-10"></span>
<span id="cb2-11"><span class="kw" style="color: #003B4F;
background-color: null;
font-weight: bold;
font-style: inherit;">def</span> sift2rootsift(desc):</span>
<span id="cb2-12">    desc <span class="op" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">/=</span> desc.<span class="bu" style="color: null;
background-color: null;
font-style: inherit;">sum</span>(axis<span class="op" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">=</span><span class="dv" style="color: #AD0000;
background-color: null;
font-style: inherit;">1</span>, keepdims<span class="op" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">=</span><span class="va" style="color: #111111;
background-color: null;
font-style: inherit;">True</span>) <span class="op" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">+</span> <span class="fl" style="color: #AD0000;
background-color: null;
font-style: inherit;">1e-8</span></span>
<span id="cb2-13">    desc <span class="op" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">=</span> np.sqrt(desc)</span>
<span id="cb2-14">    <span class="cf" style="color: #003B4F;
background-color: null;
font-weight: bold;
font-style: inherit;">return</span> desc</span>
<span id="cb2-15"></span>
<span id="cb2-16"><span class="kw" style="color: #003B4F;
background-color: null;
font-weight: bold;
font-style: inherit;">def</span> estimate_F_WithRootSIFT(img1, img2):</span>
<span id="cb2-17">    det <span class="op" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">=</span> cv2.SIFT_create(<span class="dv" style="color: #AD0000;
background-color: null;
font-style: inherit;">8000</span>, contrastThreshold<span class="op" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">=-</span><span class="dv" style="color: #AD0000;
background-color: null;
font-style: inherit;">10000</span>, edgeThreshold<span class="op" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">=</span><span class="dv" style="color: #AD0000;
background-color: null;
font-style: inherit;">10000</span>)</span>
<span id="cb2-18">    kps1, descs1 <span class="op" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">=</span> det.detectAndCompute(img1, <span class="va" style="color: #111111;
background-color: null;
font-style: inherit;">None</span>)</span>
<span id="cb2-19">    kps2, descs2 <span class="op" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">=</span> det.detectAndCompute(img2, <span class="va" style="color: #111111;
background-color: null;
font-style: inherit;">None</span>)</span>
<span id="cb2-20">    descs1 <span class="op" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">=</span> sift2rootsift(descs1)</span>
<span id="cb2-21">    descs2 <span class="op" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">=</span> sift2rootsift(descs2)</span>
<span id="cb2-22">    snn_ratio, idxs <span class="op" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">=</span> KF.match_smnn(torch.from_numpy(descs1),</span>
<span id="cb2-23">                                    torch.from_numpy(descs2), <span class="fl" style="color: #AD0000;
background-color: null;
font-style: inherit;">0.95</span>)</span>
<span id="cb2-24">    tentatives <span class="op" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">=</span> cv2_matches_from_kornia(snn_ratio, idxs)</span>
<span id="cb2-25">    src_pts <span class="op" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">=</span> np.float32([ kps1[m.queryIdx].pt <span class="cf" style="color: #003B4F;
background-color: null;
font-weight: bold;
font-style: inherit;">for</span> m <span class="kw" style="color: #003B4F;
background-color: null;
font-weight: bold;
font-style: inherit;">in</span> tentatives ]).reshape(<span class="op" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">-</span><span class="dv" style="color: #AD0000;
background-color: null;
font-style: inherit;">1</span>,<span class="dv" style="color: #AD0000;
background-color: null;
font-style: inherit;">2</span>)</span>
<span id="cb2-26">    dst_pts <span class="op" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">=</span> np.float32([ kps2[m.trainIdx].pt <span class="cf" style="color: #003B4F;
background-color: null;
font-weight: bold;
font-style: inherit;">for</span> m <span class="kw" style="color: #003B4F;
background-color: null;
font-weight: bold;
font-style: inherit;">in</span> tentatives ]).reshape(<span class="op" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">-</span><span class="dv" style="color: #AD0000;
background-color: null;
font-style: inherit;">1</span>,<span class="dv" style="color: #AD0000;
background-color: null;
font-style: inherit;">2</span>)</span>
<span id="cb2-27">    F, inlier_mask <span class="op" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">=</span> cv2.findFundamentalMat(src_pts, dst_pts, cv2.USAC_MAGSAC, <span class="fl" style="color: #AD0000;
background-color: null;
font-style: inherit;">0.25</span>, <span class="fl" style="color: #AD0000;
background-color: null;
font-style: inherit;">0.999</span>, <span class="dv" style="color: #AD0000;
background-color: null;
font-style: inherit;">100000</span>)</span>
<span id="cb2-28">    <span class="cf" style="color: #003B4F;
background-color: null;
font-weight: bold;
font-style: inherit;">return</span> F</span>
<span id="cb2-29"></span>
<span id="cb2-30"></span>
<span id="cb2-31">subset <span class="op" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">=</span> <span class="st" style="color: #20794D;
background-color: null;
font-style: inherit;">'test'</span></span>
<span id="cb2-32">dset <span class="op" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">=</span> WxBSDataset(<span class="st" style="color: #20794D;
background-color: null;
font-style: inherit;">'.WxBS'</span>, subset<span class="op" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">=</span>subset, download<span class="op" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">=</span><span class="va" style="color: #111111;
background-color: null;
font-style: inherit;">True</span>)</span>
<span id="cb2-33"></span>
<span id="cb2-34"></span>
<span id="cb2-35">F_results_RootSIFT <span class="op" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">=</span> []</span>
<span id="cb2-36"></span>
<span id="cb2-37"><span class="cf" style="color: #003B4F;
background-color: null;
font-weight: bold;
font-style: inherit;">for</span> pair_dict <span class="kw" style="color: #003B4F;
background-color: null;
font-weight: bold;
font-style: inherit;">in</span> tqdm(dset):</span>
<span id="cb2-38">    current_F <span class="op" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">=</span> estimate_F_WithRootSIFT(pair_dict[<span class="st" style="color: #20794D;
background-color: null;
font-style: inherit;">'img1'</span>], pair_dict[<span class="st" style="color: #20794D;
background-color: null;
font-style: inherit;">'img2'</span>])</span>
<span id="cb2-39">    F_results_RootSIFT.append(current_F)</span>
<span id="cb2-40"></span>
<span id="cb2-41">result_dict_rootsift, thresholds <span class="op" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">=</span> evaluate_Fs(F_results_RootSIFT, subset)</span></code></pre></div>
</div>
<p>We can check results for individual image pairs, or just take an average</p>
<div id="cell-8" class="cell" data-execution_count="5">
<div class="sourceCode cell-code" id="cb3" style="background: #f1f3f5;"><pre class="sourceCode python code-with-copy"><code class="sourceCode python"><span id="cb3-1"><span class="bu" style="color: null;
background-color: null;
font-style: inherit;">print</span>(result_dict_rootsift.keys())</span>
<span id="cb3-2">plt.figure()</span>
<span id="cb3-3">plt.plot(thresholds, result_dict_rootsift[<span class="st" style="color: #20794D;
background-color: null;
font-style: inherit;">'average'</span>], <span class="st" style="color: #20794D;
background-color: null;
font-style: inherit;">'-x'</span>)</span>
<span id="cb3-4">plt.ylim([<span class="dv" style="color: #AD0000;
background-color: null;
font-style: inherit;">0</span>,<span class="fl" style="color: #AD0000;
background-color: null;
font-style: inherit;">1.05</span>])</span>
<span id="cb3-5">plt.xlabel(<span class="st" style="color: #20794D;
background-color: null;
font-style: inherit;">'Thresholds'</span>)</span>
<span id="cb3-6">plt.ylabel(<span class="st" style="color: #20794D;
background-color: null;
font-style: inherit;">'Recall on GT corrs'</span>)</span>
<span id="cb3-7">plt.grid(<span class="va" style="color: #111111;
background-color: null;
font-style: inherit;">True</span>)</span>
<span id="cb3-8">plt.legend([<span class="st" style="color: #20794D;
background-color: null;
font-style: inherit;">'RootSIFT + MAGSAC++'</span>])</span></code></pre></div>
<div class="cell-output cell-output-stdout">
<pre><code>dict_keys(['WGABS/kremlin', 'WGABS/kyiv', 'WGABS/strahov', 'WGABS/vatutin', 'WGALBS/bridge', 'WGALBS/flood', 'WGALBS/kyiv_dolltheater', 'WGALBS/rovenki', 'WGALBS/stadium', 'WGALBS/submarine', 'WGALBS/submarine2', 'WGALBS/tyn', 'WGALBS/zanky', 'WGBS/kn-church', 'WGLBS/alupka', 'WGLBS/berlin', 'WGLBS/charlottenburg', 'WGLBS/church', 'WGLBS/him', 'WGLBS/maidan', 'WGLBS/ministry', 'WGLBS/silasveta2', 'WGLBS/warsaw', 'WGSBS/kettle', 'WGSBS/kettle2', 'WGSBS/lab', 'WGSBS/lab2', 'WGSBS/window', 'WLABS/dh', 'WLABS/kpi', 'WLABS/kyiv', 'WLABS/ministry', 'average'])</code></pre>
</div>
<div class="cell-output cell-output-stderr">
<pre><code>/opt/homebrew/Caskroom/miniforge/base/envs/python39/lib/python3.9/site-packages/ipykernel/ipkernel.py:283: DeprecationWarning: `should_run_async` will not call `transform_cell` automatically in the future. Please pass the result to `transformed_cell` argument and any exception that happen during thetransform in `preprocessing_exc_tuple` in IPython 7.17 and above.
  and should_run_async(code)</code></pre>
</div>
<div class="cell-output cell-output-display">
<div>
<figure class="figure">
<p><img src="https://ducha-aiki.github.io/wide-baseline-stereo-blog/posts/2021-07-30-Reviving-WxBS-benchmark_files/figure-html/cell-4-output-3.png" class="img-fluid figure-img"></p>
</figure>
</div>
</div>
</div>
<p>I have run several popular recent methods in this <a href="https://colab.research.google.com/drive/1yrCFyEoAc0HyqYCRVvzJDh5kQT2Dc3XA?usp=sharing">Colab</a>. The code is very dirty, that is why I don’t put it here, and just present results.</p>
</section>
<section id="methods-tested" class="level2">
<h2 class="anchored" data-anchor-id="methods-tested">Methods tested</h2>
<section id="local-features-without-learned-matching." class="level3">
<h3 class="anchored" data-anchor-id="local-features-without-learned-matching.">Local features without learned matching.</h3>
<p><a href="https://www.robots.ox.ac.uk/~vgg/publications/2012/Arandjelovic12/arandjelovic12.pdf">RootSIFT</a> (ICCV1999 - CVPR-2012), <a href="https://www.cs.ubc.ca/~lowe/papers/ijcv04.pdf">DoG</a>+<a href="https://arxiv.org/pdf/1705.10872.pdf">HardNet</a> (NeurIPS 2017), <a href="https://arxiv.org/abs/1906.06195">R2D2</a> (NeurIPS 2019), <a href="https://arxiv.org/abs/2006.13566">DISK</a> (NeurIPS 2020).</p>
</section>
<section id="local-features-with-learned-matching." class="level3">
<h3 class="anchored" data-anchor-id="local-features-with-learned-matching.">Local features with learned matching.</h3>
<p><a href="https://arxiv.org/abs/1911.11763">SuperGlue</a> (CVPR2020) uses attention-based graph neural network to match SuperPoint local features.</p>
</section>
<section id="dense-matching" class="level3">
<h3 class="anchored" data-anchor-id="dense-matching">Dense matching</h3>
<p><a href="https://prunetruong.com/research/pdcnet">PDCNet</a> (CVPR 2021). On top of pretrained ImageNet coarse-to-fine feature correlation, PDCNet predicts a dense flow and a confidence for each pixel. Trained on huge dataset with synthetic warps.</p>
<p><a href="https://arxiv.org/abs/2012.01909">Patch2pix</a> (CVPR 2021). Similarly to PDCNet, patch2pix is uses pretrained ImageNet features to get the initial corresponcences. Then the matches are progressively refined or rejected with a regressor network.</p>
<p><a href="https://arxiv.org/abs/2106.07791">DFM</a> (CVPRW 2021). Like PDCNet and Patch2pix, DFM uses ImageNet-pretrained and frozen VGGNet backbone as dense feature extractor. Unlike them, DFM does not requite any training and relies on simple handcrafted coarse-to-fine matching from conv5 to conv1.</p>
</section>
<section id="dense-matching-with-transformers" class="level3">
<h3 class="anchored" data-anchor-id="dense-matching-with-transformers">Dense matching with transformers</h3>
<p><a href="https://arxiv.org/abs/2103.14167">COTR</a>, (ICCV2021). Given the pixel location in a query image, COTR finds a corresponing location in the reference image. It operated on the 16x16 feature map produced by pretrained ImageNet model, which then are fed together with a positional embeddings to a transformer-based regressor. It is also the slowest method among evaluated - because you have to run it <em>per each pixel separately</em>, if want dense correspondences.</p>
<p><a href="https://zju3dv.github.io/loftr/">LoFTR</a> (CVPR2021). LoFTR is very similar to Patch2pix, with the difference of using transformer architecture to for both coarse initial matching and consequent refinement. The second difference, is that, unlike all other methods, which are using ImageNet-pretrained models as a feature backbone, LoFTR trains ResNet18 <em>from scratch</em> .</p>
</section>
</section>
<section id="results-discussion" class="level2">
<h2 class="anchored" data-anchor-id="results-discussion">Results &amp; Discussion</h2>
<div class="quarto-figure quarto-figure-center">
<figure class="figure">
<p><img src="https://ducha-aiki.github.io/wide-baseline-stereo-blog/posts/2021-07-30-Reviving-WxBS-benchmark_files/att_00000.png" class="img-fluid figure-img"></p>
<figcaption>image.png</figcaption>
</figure>
</div>
<p>The clear leader is LoFTR, then we have (patch2pix, COTR and SuperGlue) group, followed by (PDCNet, DISK, R2D2). Then DoG-HardNet and RootSIFT, where latter has barely matched couple of image pairs correctly.</p>
<p>I have no idea, why LoFTR is so good, my guesses would be richer training data and training the backbone from scratch, instead of relying on ImageNet feaures. I am also quite surprised by the great performance of the patch2pix, which probably can be improved even more, if the backbone would be at least fine-tuned.</p>
<p>Another surprise is a great performance of the DFM matcher, which is slightly better than PDCNet and worse (but comparable) than COTR. My hypothesis is that all this methods rely on the same ImageNet backbone: if they fail, all the methods fail.</p>
<p>Regarding SuperGlue - one of my guesses is that original SuperPoint features were not trained for different illuminations. Moreover, in some images, there just not enough matching corners - when I was annotating the correspondences, I have to use also “center of the blobs”.</p>
<p>Finally, I would like to that than even excellent LofTR result is far from the ideal – there are stil tens of pairs, where it fails.</p>
<p>P.S. I would also like to thank Prune Truong and Ufuk Efe who helped me to debug my originally wrong run of the PDCNet and DFM respectively.</p>
<section id="qualiative-examples" class="level3">
<h3 class="anchored" data-anchor-id="qualiative-examples">Qualiative examples</h3>
<p>Here are some examples of how different methods work on the same image pairs.</p>
</section>
<section id="season-viewpoint" class="level3">
<h3 class="anchored" data-anchor-id="season-viewpoint">Season + viewpoint</h3>
<p>LoFTR</p>
<div class="quarto-figure quarto-figure-center">
<figure class="figure">
<p><img src="https://ducha-aiki.github.io/wide-baseline-stereo-blog/posts/2021-07-28-Reviving-WxBS-benchmark_files/att_00006.png" class="img-fluid figure-img"></p>
<figcaption>image.png</figcaption>
</figure>
</div>
<p>SuperGlue</p>
<div class="quarto-figure quarto-figure-center">
<figure class="figure">
<p><img src="https://ducha-aiki.github.io/wide-baseline-stereo-blog/posts/2021-07-28-Reviving-WxBS-benchmark_files/att_00007.png" class="img-fluid figure-img"></p>
<figcaption>image.png</figcaption>
</figure>
</div>
<p>patch2pix</p>
<div class="quarto-figure quarto-figure-center">
<figure class="figure">
<p><img src="https://ducha-aiki.github.io/wide-baseline-stereo-blog/posts/2021-07-28-Reviving-WxBS-benchmark_files/att_00008.png" class="img-fluid figure-img"></p>
<figcaption>image.png</figcaption>
</figure>
</div>
</section>
<section id="thermal-vs-visible-viewpoint" class="level3">
<h3 class="anchored" data-anchor-id="thermal-vs-visible-viewpoint">Thermal vs visible + viewpoint</h3>
<p>LoFTR</p>
<div class="quarto-figure quarto-figure-center">
<figure class="figure">
<p><img src="https://ducha-aiki.github.io/wide-baseline-stereo-blog/posts/2021-07-28-Reviving-WxBS-benchmark_files/att_00011.png" class="img-fluid figure-img"></p>
<figcaption>image.png</figcaption>
</figure>
</div>
<p>SuperGlue</p>
<div class="quarto-figure quarto-figure-center">
<figure class="figure">
<p><img src="https://ducha-aiki.github.io/wide-baseline-stereo-blog/posts/2021-07-28-Reviving-WxBS-benchmark_files/att_00010.png" class="img-fluid figure-img"></p>
<figcaption>image.png</figcaption>
</figure>
</div>
<p>patch2pix</p>
<div class="quarto-figure quarto-figure-center">
<figure class="figure">
<p><img src="https://ducha-aiki.github.io/wide-baseline-stereo-blog/posts/2021-07-28-Reviving-WxBS-benchmark_files/att_00009.png" class="img-fluid figure-img"></p>
<figcaption>image.png</figcaption>
</figure>
</div>
</section>
<section id="season-illumination-viewpoint" class="level3">
<h3 class="anchored" data-anchor-id="season-illumination-viewpoint">Season + illumination + viewpoint</h3>
<p>LoFTR</p>
<div class="quarto-figure quarto-figure-center">
<figure class="figure">
<p><img src="https://ducha-aiki.github.io/wide-baseline-stereo-blog/posts/2021-07-28-Reviving-WxBS-benchmark_files/att_00014.png" class="img-fluid figure-img"></p>
<figcaption>image.png</figcaption>
</figure>
</div>
<p>SuperGlue</p>
<div class="quarto-figure quarto-figure-center">
<figure class="figure">
<p><img src="https://ducha-aiki.github.io/wide-baseline-stereo-blog/posts/2021-07-28-Reviving-WxBS-benchmark_files/att_00016.png" class="img-fluid figure-img"></p>
<figcaption>image.png</figcaption>
</figure>
</div>
<p>patch2pix</p>
<div class="quarto-figure quarto-figure-center">
<figure class="figure">
<p><img src="https://ducha-aiki.github.io/wide-baseline-stereo-blog/posts/2021-07-28-Reviving-WxBS-benchmark_files/att_00017.png" class="img-fluid figure-img"></p>
<figcaption>image.png</figcaption>
</figure>
</div>
</section>
<section id="illumination-viewpoint-change-zoom" class="level3">
<h3 class="anchored" data-anchor-id="illumination-viewpoint-change-zoom">Illumination + viewpoint change (zoom)</h3>
<p>LoFTR</p>
<div class="quarto-figure quarto-figure-center">
<figure class="figure">
<p><img src="https://ducha-aiki.github.io/wide-baseline-stereo-blog/posts/2021-07-28-Reviving-WxBS-benchmark_files/att_00018.png" class="img-fluid figure-img"></p>
<figcaption>image.png</figcaption>
</figure>
</div>
<p>SuperGlue <img src="https://ducha-aiki.github.io/wide-baseline-stereo-blog/posts/2021-07-28-Reviving-WxBS-benchmark_files/att_00019.png" class="img-fluid" alt="image.png"></p>
<p>patch2pix</p>
<div class="quarto-figure quarto-figure-center">
<figure class="figure">
<p><img src="https://ducha-aiki.github.io/wide-baseline-stereo-blog/posts/2021-07-28-Reviving-WxBS-benchmark_files/att_00020.png" class="img-fluid figure-img"></p>
<figcaption>image.png</figcaption>
</figure>
</div>
<p>As you can, see, there are many incorrect correspondences even in those images, where methods are performing well.</p>
<p>Here I will conclude and go to vacation.</p>
<p>In the 2nd part of the post, I would like to evaluate COTR, PDCNet and different dense descriptors directly on the GT correspondences, by giving a point in image A and asking for point in image B.</p>


</section>
</section>

<p>Everything you (didn’t) want to know about image matching</p> ]]></description>
  <guid>https://ducha-aiki.github.io/wide-baseline-stereo-blog/posts/2021-07-30-reviving-wxbs-benchmark.html</guid>
  <pubDate>Thu, 29 Jul 2021 22:00:00 GMT</pubDate>
  <media:content url="https://ducha-aiki.github.io/wide-baseline-stereo-blog/posts/images/loftr_match.png" medium="image" type="image/png" height="214" width="144"/>
</item>
<item>
  <title>Submitting to IMC 2021 with custom matcher</title>
  <dc:creator>Dmytro Mishkin</dc:creator>
  <link>https://ducha-aiki.github.io/wide-baseline-stereo-blog/posts/2021-05-27-submitting-to-imc2021-with-custom-matcher.html</link>
  <description><![CDATA[ 





<section id="intro" class="level2">
<h2 class="anchored" data-anchor-id="intro">Intro</h2>
<p>In the <a href="https://ducha-aiki.github.io/wide-baseline-stereo-blog/posts/2021-05-12-submitting-to-imc2021-step-by-step.html">previous tutorial</a> we have created a pipeline for submitting to the IMC2021. However, we have not covered all the possibilities. First, we haven’t submissted custom matches, instead we just run a standard Lowe SNN test with cross-check. Second, we haven’t evaluate multiview (COLMAP) part. In this tutorial we will do both. I am assuming that you have completed the <a href="https://ducha-aiki.github.io/wide-baseline-stereo-blog/posts/2021-05-12-submitting-to-imc2021-step-by-step.html">previous part</a>. If not, please do, because we will be relying on the already extracted features. Let’s check if the feature are there</p>
<div class="sourceCode" id="cb1" style="background: #f1f3f5;"><pre class="sourceCode bash code-with-copy"><code class="sourceCode bash"><span id="cb1-1"><span class="bu" style="color: null;
background-color: null;
font-style: inherit;">cd</span> imc2021-sample-kornia-submission</span>
<span id="cb1-2"><span class="fu" style="color: #4758AB;
background-color: null;
font-style: inherit;">ls</span> extracted/cv2-dog-affnet-hardnet8/googleurban/edinburgh/</span></code></pre></div>
<pre><code>angles.h5  descriptors.h5  keypoints.h5  scales.h5  scores.h5</code></pre>
<div class="callout callout-style-default callout-warning callout-titled">
<div class="callout-header d-flex align-content-center">
<div class="callout-icon-container">
<i class="callout-icon"></i>
</div>
<div class="callout-title-container flex-fill">
Warning
</div>
</div>
<div class="callout-body-container callout-body">
<p>There is a subtle problem with previous tutorial, so we cannot use pre-extracted features.</p>
</div>
</div>
<p>Specifically, because OpenCV SIFT does not exactly respect max_features parameter and can sometimes output 8002 features, instead of 8000. When we were importing the features alone, benchmark <code>import_features.py</code> script automatically re-sorted features based on the score and clip the extra 1-3 features. However, this functionality is not available for the importing custom matches. I have already corrected <a href="https://ducha-aiki.github.io/wide-baseline-stereo-blog/posts/2021-05-12-submitting-to-imc2021-step-by-step.html">previous post</a>, so you can use it for the re-extration. Or, if you are reading this whole tutorial after May 24, just ignore this.</p>
<p>Now we will install <a href="https://arxiv.org/abs/2006.04250">AdaLAM</a> - <a href="https://ducha-aiki.github.io/wide-baseline-stereo-blog/posts/2021-05-14-imc2020-competition-recap.html">one of the winners of IMC2020 Challenge</a>. It uses keypoint geometry to filter out unreliable matches.</p>
<div class="sourceCode" id="cb3" style="background: #f1f3f5;"><pre class="sourceCode bash code-with-copy"><code class="sourceCode bash"><span id="cb3-1"><span class="ex" style="color: null;
background-color: null;
font-style: inherit;">pip</span> install git+https://github.com/cavalli1234/AdaLAM.git</span></code></pre></div>
<p>Let’s check if it works on the sample image pair. We will read the pre-extracted features for it.</p>
<div id="cell-3" class="cell">
<div class="sourceCode cell-code" id="cb4" style="background: #f1f3f5;"><pre class="sourceCode python code-with-copy"><code class="sourceCode python"><span id="cb4-1"><span class="im" style="color: #00769E;
background-color: null;
font-style: inherit;">import</span> matplotlib.pyplot <span class="im" style="color: #00769E;
background-color: null;
font-style: inherit;">as</span> plt</span>
<span id="cb4-2"><span class="im" style="color: #00769E;
background-color: null;
font-style: inherit;">import</span> numpy <span class="im" style="color: #00769E;
background-color: null;
font-style: inherit;">as</span> np</span>
<span id="cb4-3"><span class="im" style="color: #00769E;
background-color: null;
font-style: inherit;">import</span> cv2</span>
<span id="cb4-4"><span class="im" style="color: #00769E;
background-color: null;
font-style: inherit;">import</span> os</span>
<span id="cb4-5"><span class="im" style="color: #00769E;
background-color: null;
font-style: inherit;">import</span> torch</span>
<span id="cb4-6"><span class="im" style="color: #00769E;
background-color: null;
font-style: inherit;">import</span> kornia <span class="im" style="color: #00769E;
background-color: null;
font-style: inherit;">as</span> K</span>
<span id="cb4-7"><span class="im" style="color: #00769E;
background-color: null;
font-style: inherit;">import</span> kornia.feature <span class="im" style="color: #00769E;
background-color: null;
font-style: inherit;">as</span> KF</span>
<span id="cb4-8"><span class="im" style="color: #00769E;
background-color: null;
font-style: inherit;">import</span> h5py</span>
<span id="cb4-9"><span class="im" style="color: #00769E;
background-color: null;
font-style: inherit;">import</span> json</span>
<span id="cb4-10"><span class="im" style="color: #00769E;
background-color: null;
font-style: inherit;">from</span> PIL <span class="im" style="color: #00769E;
background-color: null;
font-style: inherit;">import</span> Image</span>
<span id="cb4-11"><span class="im" style="color: #00769E;
background-color: null;
font-style: inherit;">from</span> adalam <span class="im" style="color: #00769E;
background-color: null;
font-style: inherit;">import</span> AdalamFilter</span>
<span id="cb4-12"><span class="im" style="color: #00769E;
background-color: null;
font-style: inherit;">from</span> kornia_moons.feature <span class="im" style="color: #00769E;
background-color: null;
font-style: inherit;">import</span> <span class="op" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">*</span></span>
<span id="cb4-13"></span>
<span id="cb4-14"><span class="kw" style="color: #003B4F;
background-color: null;
font-weight: bold;
font-style: inherit;">def</span> load_h5(filename):</span>
<span id="cb4-15">    <span class="co" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">'''Loads dictionary from hdf5 file'''</span></span>
<span id="cb4-16"></span>
<span id="cb4-17">    dict_to_load <span class="op" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">=</span> {}</span>
<span id="cb4-18">    <span class="cf" style="color: #003B4F;
background-color: null;
font-weight: bold;
font-style: inherit;">try</span>:</span>
<span id="cb4-19">        <span class="cf" style="color: #003B4F;
background-color: null;
font-weight: bold;
font-style: inherit;">with</span> h5py.File(filename, <span class="st" style="color: #20794D;
background-color: null;
font-style: inherit;">'r'</span>) <span class="im" style="color: #00769E;
background-color: null;
font-style: inherit;">as</span> f:</span>
<span id="cb4-20">            keys <span class="op" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">=</span> [key <span class="cf" style="color: #003B4F;
background-color: null;
font-weight: bold;
font-style: inherit;">for</span> key <span class="kw" style="color: #003B4F;
background-color: null;
font-weight: bold;
font-style: inherit;">in</span> f.keys()]</span>
<span id="cb4-21">            <span class="cf" style="color: #003B4F;
background-color: null;
font-weight: bold;
font-style: inherit;">for</span> key <span class="kw" style="color: #003B4F;
background-color: null;
font-weight: bold;
font-style: inherit;">in</span> keys:</span>
<span id="cb4-22">                dict_to_load[key] <span class="op" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">=</span> f[key][()]</span>
<span id="cb4-23">    <span class="cf" style="color: #003B4F;
background-color: null;
font-weight: bold;
font-style: inherit;">except</span>:</span>
<span id="cb4-24">        <span class="bu" style="color: null;
background-color: null;
font-style: inherit;">print</span>(<span class="st" style="color: #20794D;
background-color: null;
font-style: inherit;">'Cannot find file </span><span class="sc" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">{}</span><span class="st" style="color: #20794D;
background-color: null;
font-style: inherit;">'</span>.<span class="bu" style="color: null;
background-color: null;
font-style: inherit;">format</span>(filename))</span>
<span id="cb4-25">    <span class="cf" style="color: #003B4F;
background-color: null;
font-weight: bold;
font-style: inherit;">return</span> dict_to_load</span>
<span id="cb4-26"></span>
<span id="cb4-27">PATH_TO_FEATS <span class="op" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">=</span> <span class="st" style="color: #20794D;
background-color: null;
font-style: inherit;">'extracted/cv2-dog-affnet-hardnet8/googleurban/edinburgh/'</span></span>
<span id="cb4-28">kps <span class="op" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">=</span> load_h5(os.path.join(PATH_TO_FEATS, <span class="st" style="color: #20794D;
background-color: null;
font-style: inherit;">'keypoints.h5'</span>))</span>
<span id="cb4-29">angles <span class="op" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">=</span> load_h5(os.path.join(PATH_TO_FEATS, <span class="st" style="color: #20794D;
background-color: null;
font-style: inherit;">'angles.h5'</span>))</span>
<span id="cb4-30">scales <span class="op" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">=</span> load_h5(os.path.join(PATH_TO_FEATS, <span class="st" style="color: #20794D;
background-color: null;
font-style: inherit;">'scales.h5'</span>))</span>
<span id="cb4-31">descs <span class="op" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">=</span> load_h5(os.path.join(PATH_TO_FEATS, <span class="st" style="color: #20794D;
background-color: null;
font-style: inherit;">'descriptors.h5'</span>))</span></code></pre></div>
</div>
<p>I have selected two images, which are matching:</p>
<div id="cell-5" class="cell">
<div class="sourceCode cell-code" id="cb5" style="background: #f1f3f5;"><pre class="sourceCode python code-with-copy"><code class="sourceCode python"><span id="cb5-1">IMG_DIR <span class="op" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">=</span> <span class="st" style="color: #20794D;
background-color: null;
font-style: inherit;">'../imc-2021-data/googleurban/edinburgh/set_100/images/'</span></span>
<span id="cb5-2">img1_key <span class="op" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">=</span> <span class="st" style="color: #20794D;
background-color: null;
font-style: inherit;">'2b5315968bc5468c995b978620879439'</span></span>
<span id="cb5-3">img2_key <span class="op" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">=</span> <span class="st" style="color: #20794D;
background-color: null;
font-style: inherit;">'6264aee21d1b48b7985901c4bedfdbd4'</span></span>
<span id="cb5-4"></span>
<span id="cb5-5">img1 <span class="op" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">=</span> cv2.cvtColor(cv2.imread(os.path.join(IMG_DIR, <span class="ss" style="color: #20794D;
background-color: null;
font-style: inherit;">f'</span><span class="sc" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">{</span>img1_key<span class="sc" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">}</span><span class="ss" style="color: #20794D;
background-color: null;
font-style: inherit;">.png'</span>)), cv2.COLOR_BGR2RGB)</span>
<span id="cb5-6">img2 <span class="op" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">=</span> cv2.cvtColor(cv2.imread(os.path.join(IMG_DIR, <span class="ss" style="color: #20794D;
background-color: null;
font-style: inherit;">f'</span><span class="sc" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">{</span>img2_key<span class="sc" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">}</span><span class="ss" style="color: #20794D;
background-color: null;
font-style: inherit;">.png'</span>)), cv2.COLOR_BGR2RGB)</span>
<span id="cb5-7">plt.imshow(np.concatenate([img1, img2], axis<span class="op" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">=</span><span class="dv" style="color: #AD0000;
background-color: null;
font-style: inherit;">1</span>))</span></code></pre></div>
<div class="cell-output cell-output-display">
<div>
<figure class="figure">
<p><img src="https://ducha-aiki.github.io/wide-baseline-stereo-blog/posts/2021-05-27-submitting-to-IMC2021-with-custom-matcher_files/figure-html/cell-3-output-1.png" class="img-fluid figure-img"></p>
</figure>
</div>
</div>
</div>
<p>We will start with matching and drawing the matches with OpenCV for sanity check.</p>
<div id="cell-7" class="cell">
<div class="sourceCode cell-code" id="cb6" style="background: #f1f3f5;"><pre class="sourceCode python code-with-copy"><code class="sourceCode python"><span id="cb6-1"><span class="kw" style="color: #003B4F;
background-color: null;
font-weight: bold;
font-style: inherit;">def</span> opencv_from_imc(kps, sizes, angles):</span>
<span id="cb6-2">    <span class="cf" style="color: #003B4F;
background-color: null;
font-weight: bold;
font-style: inherit;">return</span> [cv2.KeyPoint(kp[<span class="dv" style="color: #AD0000;
background-color: null;
font-style: inherit;">0</span>], kp[<span class="dv" style="color: #AD0000;
background-color: null;
font-style: inherit;">1</span>], <span class="bu" style="color: null;
background-color: null;
font-style: inherit;">float</span>(s), <span class="bu" style="color: null;
background-color: null;
font-style: inherit;">float</span>(a)) <span class="cf" style="color: #003B4F;
background-color: null;
font-weight: bold;
font-style: inherit;">for</span>  kp, s, a <span class="kw" style="color: #003B4F;
background-color: null;
font-weight: bold;
font-style: inherit;">in</span> <span class="bu" style="color: null;
background-color: null;
font-style: inherit;">zip</span>(kps, sizes, angles)]</span>
<span id="cb6-3"></span>
<span id="cb6-4"><span class="kw" style="color: #003B4F;
background-color: null;
font-weight: bold;
font-style: inherit;">def</span> get_data(kps, angles, scales, descs, img_key):</span>
<span id="cb6-5">    kp1 <span class="op" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">=</span> kps[img_key]</span>
<span id="cb6-6">    s1 <span class="op" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">=</span> scales[img_key]</span>
<span id="cb6-7">    a1 <span class="op" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">=</span> angles[img_key]</span>
<span id="cb6-8">    descs1 <span class="op" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">=</span> descs[img_key]</span>
<span id="cb6-9">    <span class="cf" style="color: #003B4F;
background-color: null;
font-weight: bold;
font-style: inherit;">return</span> kp1, s1, a1, descs1</span>
<span id="cb6-10"> </span>
<span id="cb6-11"><span class="kw" style="color: #003B4F;
background-color: null;
font-weight: bold;
font-style: inherit;">def</span> match(img1_key, img2_key, kps, angles, scales, descs):</span>
<span id="cb6-12">    kp1, s1, a1, descs1 <span class="op" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">=</span> get_data(kps, angles, scales, descs, img1_key)</span>
<span id="cb6-13">    kp2, s2, a2, descs2 <span class="op" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">=</span> get_data(kps, angles, scales, descs, img2_key)</span>
<span id="cb6-14">    dists, idxs <span class="op" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">=</span> KF.match_smnn(torch.from_numpy(descs1), torch.from_numpy(descs2), <span class="fl" style="color: #AD0000;
background-color: null;
font-style: inherit;">0.9</span>)</span>
<span id="cb6-15">    <span class="cf" style="color: #003B4F;
background-color: null;
font-weight: bold;
font-style: inherit;">return</span> dists, idxs</span>
<span id="cb6-16"></span>
<span id="cb6-17"><span class="kw" style="color: #003B4F;
background-color: null;
font-weight: bold;
font-style: inherit;">def</span> draw_matches(img1_key, img2_key, dists, idxs, kps, angles, scales, descs):</span>
<span id="cb6-18">    tentatives <span class="op" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">=</span> cv2_matches_from_kornia(dists, idxs)</span>
<span id="cb6-19">    draw_params <span class="op" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">=</span> <span class="bu" style="color: null;
background-color: null;
font-style: inherit;">dict</span>(matchColor <span class="op" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">=</span> (<span class="dv" style="color: #AD0000;
background-color: null;
font-style: inherit;">255</span>,<span class="dv" style="color: #AD0000;
background-color: null;
font-style: inherit;">255</span>,<span class="dv" style="color: #AD0000;
background-color: null;
font-style: inherit;">0</span>), <span class="co" style="color: #5E5E5E;
background-color: null;
font-style: inherit;"># draw matches in yellow color</span></span>
<span id="cb6-20">                   singlePointColor <span class="op" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">=</span> <span class="va" style="color: #111111;
background-color: null;
font-style: inherit;">None</span>,</span>
<span id="cb6-21">                   matchesMask <span class="op" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">=</span> [<span class="va" style="color: #111111;
background-color: null;
font-style: inherit;">True</span> <span class="cf" style="color: #003B4F;
background-color: null;
font-weight: bold;
font-style: inherit;">for</span> x <span class="kw" style="color: #003B4F;
background-color: null;
font-weight: bold;
font-style: inherit;">in</span> idxs], <span class="co" style="color: #5E5E5E;
background-color: null;
font-style: inherit;"># draw only inliers</span></span>
<span id="cb6-22">                   flags <span class="op" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">=</span> <span class="dv" style="color: #AD0000;
background-color: null;
font-style: inherit;">2</span>)</span>
<span id="cb6-23">    img1 <span class="op" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">=</span> cv2.cvtColor(cv2.imread(os.path.join(IMG_DIR, <span class="ss" style="color: #20794D;
background-color: null;
font-style: inherit;">f'</span><span class="sc" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">{</span>img1_key<span class="sc" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">}</span><span class="ss" style="color: #20794D;
background-color: null;
font-style: inherit;">.png'</span>)), cv2.COLOR_BGR2RGB)</span>
<span id="cb6-24">    img2 <span class="op" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">=</span> cv2.cvtColor(cv2.imread(os.path.join(IMG_DIR, <span class="ss" style="color: #20794D;
background-color: null;
font-style: inherit;">f'</span><span class="sc" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">{</span>img2_key<span class="sc" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">}</span><span class="ss" style="color: #20794D;
background-color: null;
font-style: inherit;">.png'</span>)), cv2.COLOR_BGR2RGB)</span>
<span id="cb6-25">    kp1, s1, a1, _ <span class="op" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">=</span> get_data(kps, angles, scales, descs, img1_key)</span>
<span id="cb6-26">    kp2, s2, a2, descs2 <span class="op" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">=</span> get_data(kps, angles, scales, descs, img2_key)</span>
<span id="cb6-27">    </span>
<span id="cb6-28">    img_out <span class="op" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">=</span> cv2.drawMatches(img1,opencv_from_imc(kp1, s1, a1),</span>
<span id="cb6-29">                              img2,opencv_from_imc(kp2, s2, a2),</span>
<span id="cb6-30">                              tentatives,<span class="va" style="color: #111111;
background-color: null;
font-style: inherit;">None</span>,<span class="op" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">**</span>draw_params)</span>
<span id="cb6-31">    plt.figure()</span>
<span id="cb6-32">    fig, ax <span class="op" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">=</span> plt.subplots(figsize<span class="op" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">=</span>(<span class="dv" style="color: #AD0000;
background-color: null;
font-style: inherit;">15</span>, <span class="dv" style="color: #AD0000;
background-color: null;
font-style: inherit;">15</span>))</span>
<span id="cb6-33">    ax.imshow(img_out, interpolation<span class="op" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">=</span><span class="st" style="color: #20794D;
background-color: null;
font-style: inherit;">'nearest'</span>)</span>
<span id="cb6-34">    <span class="cf" style="color: #003B4F;
background-color: null;
font-weight: bold;
font-style: inherit;">return</span></span>
<span id="cb6-35"></span>
<span id="cb6-36">dists, idxs <span class="op" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">=</span> match(img1_key, img2_key, kps, angles, scales, descs)</span>
<span id="cb6-37">draw_matches(img1_key, img2_key, dists, idxs,  kps, angles, scales, descs)</span></code></pre></div>
<div class="cell-output cell-output-display">
<pre><code>&lt;Figure size 432x288 with 0 Axes&gt;</code></pre>
</div>
<div class="cell-output cell-output-display">
<div>
<figure class="figure">
<p><img src="https://ducha-aiki.github.io/wide-baseline-stereo-blog/posts/2021-05-27-submitting-to-IMC2021-with-custom-matcher_files/figure-html/cell-4-output-2.png" class="img-fluid figure-img"></p>
</figure>
</div>
</div>
</div>
<p>Everything seems to be working quite well. But, as you can see, we have some wrong matches among tentatives. Let’s try AdaLAM</p>
<div id="cell-9" class="cell">
<div class="sourceCode cell-code" id="cb8" style="background: #f1f3f5;"><pre class="sourceCode python code-with-copy"><code class="sourceCode python"><span id="cb8-1"><span class="kw" style="color: #003B4F;
background-color: null;
font-weight: bold;
font-style: inherit;">def</span> match_adalam(img1_key, img2_key, kps, angles, scales, descs):</span>
<span id="cb8-2">    kp1, s1, a1, descs1 <span class="op" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">=</span> get_data(kps, angles, scales, descs, img1_key)</span>
<span id="cb8-3">    kp2, s2, a2, descs2 <span class="op" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">=</span> get_data(kps, angles, scales, descs, img2_key)</span>
<span id="cb8-4">    matcher <span class="op" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">=</span> AdalamFilter()</span>
<span id="cb8-5">    </span>
<span id="cb8-6">    <span class="co" style="color: #5E5E5E;
background-color: null;
font-style: inherit;"># AdaLAM wants image sizes, so we have to read them.</span></span>
<span id="cb8-7">    img1 <span class="op" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">=</span> cv2.cvtColor(cv2.imread(os.path.join(IMG_DIR, <span class="ss" style="color: #20794D;
background-color: null;
font-style: inherit;">f'</span><span class="sc" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">{</span>img1_key<span class="sc" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">}</span><span class="ss" style="color: #20794D;
background-color: null;
font-style: inherit;">.png'</span>)), cv2.COLOR_BGR2RGB)</span>
<span id="cb8-8">    img2 <span class="op" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">=</span> cv2.cvtColor(cv2.imread(os.path.join(IMG_DIR, <span class="ss" style="color: #20794D;
background-color: null;
font-style: inherit;">f'</span><span class="sc" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">{</span>img2_key<span class="sc" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">}</span><span class="ss" style="color: #20794D;
background-color: null;
font-style: inherit;">.png'</span>)), cv2.COLOR_BGR2RGB)</span>
<span id="cb8-9"></span>
<span id="cb8-10">    idxs <span class="op" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">=</span> matcher.match_and_filter(kp1, kp2,</span>
<span id="cb8-11">                            descs1, descs2,</span>
<span id="cb8-12">                            im1shape<span class="op" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">=</span>img1.shape[:<span class="dv" style="color: #AD0000;
background-color: null;
font-style: inherit;">2</span>], </span>
<span id="cb8-13">                            im2shape<span class="op" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">=</span>img2.shape[:<span class="dv" style="color: #AD0000;
background-color: null;
font-style: inherit;">2</span>],</span>
<span id="cb8-14">                            o1<span class="op" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">=</span>a1.reshape(<span class="op" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">-</span><span class="dv" style="color: #AD0000;
background-color: null;
font-style: inherit;">1</span>),</span>
<span id="cb8-15">                            o2<span class="op" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">=</span>a2.reshape(<span class="op" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">-</span><span class="dv" style="color: #AD0000;
background-color: null;
font-style: inherit;">1</span>),</span>
<span id="cb8-16">                            s1<span class="op" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">=</span>s1.reshape(<span class="op" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">-</span><span class="dv" style="color: #AD0000;
background-color: null;
font-style: inherit;">1</span>),</span>
<span id="cb8-17">                            s2<span class="op" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">=</span>s2.reshape(<span class="op" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">-</span><span class="dv" style="color: #AD0000;
background-color: null;
font-style: inherit;">1</span>))</span>
<span id="cb8-18">    <span class="co" style="color: #5E5E5E;
background-color: null;
font-style: inherit;"># AdaLAM does not provide confidence score, so we will create dummy one</span></span>
<span id="cb8-19">    dists <span class="op" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">=</span> torch.ones_like(idxs)[:,<span class="dv" style="color: #AD0000;
background-color: null;
font-style: inherit;">0</span>] </span>
<span id="cb8-20">    <span class="cf" style="color: #003B4F;
background-color: null;
font-weight: bold;
font-style: inherit;">return</span> dists, idxs</span>
<span id="cb8-21">dists, idxs <span class="op" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">=</span> match_adalam(img1_key, img2_key, kps, angles, scales, descs)</span>
<span id="cb8-22">draw_matches(img1_key, img2_key, dists, idxs,  kps, angles, scales, descs)</span></code></pre></div>
<div class="cell-output cell-output-display">
<pre><code>&lt;Figure size 432x288 with 0 Axes&gt;</code></pre>
</div>
<div class="cell-output cell-output-display">
<div>
<figure class="figure">
<p><img src="https://ducha-aiki.github.io/wide-baseline-stereo-blog/posts/2021-05-27-submitting-to-IMC2021-with-custom-matcher_files/figure-html/cell-5-output-2.png" class="img-fluid figure-img"></p>
</figure>
</div>
</div>
</div>
<p>We have significantly more matches, which are also cleaner, although not ideal. There is one more thing: if we are going to submit custom matches, we have to run RANSAC ourself (if needed), here is a quote from the <a href="https://www.cs.ubc.ca/research/image-matching-challenge/2021/submit/">competition rules</a>:</p>
<pre><code>The "geom"/"method" field must be set to "cv2-8pt" if custom matches are enabled, as we assume you tune and run your favourite RANSAC algorithm, if applicable.</code></pre>
<p>Thus, we need to merge RANSAC (we pick inlier threshold from the previous tutorial) and matching:</p>
<div id="cell-11" class="cell">
<div class="sourceCode cell-code" id="cb11" style="background: #f1f3f5;"><pre class="sourceCode python code-with-copy"><code class="sourceCode python"><span id="cb11-1"><span class="im" style="color: #00769E;
background-color: null;
font-style: inherit;">import</span> pydegensac</span>
<span id="cb11-2"></span>
<span id="cb11-3"><span class="kw" style="color: #003B4F;
background-color: null;
font-weight: bold;
font-style: inherit;">def</span> match_adalam_with_degensac(img1_key, img2_key, kps, angles, scales, descs, ds_name<span class="op" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">=</span><span class="st" style="color: #20794D;
background-color: null;
font-style: inherit;">'phototourism'</span>):</span>
<span id="cb11-4">    kp1, s1, a1, descs1 <span class="op" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">=</span> get_data(kps, angles, scales, descs, img1_key)</span>
<span id="cb11-5">    kp2, s2, a2, descs2 <span class="op" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">=</span> get_data(kps, angles, scales, descs, img2_key)</span>
<span id="cb11-6">    matcher <span class="op" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">=</span> AdalamFilter()</span>
<span id="cb11-7">    </span>
<span id="cb11-8">    <span class="co" style="color: #5E5E5E;
background-color: null;
font-style: inherit;"># AdaLAM wants image sizes, so we have to read them.</span></span>
<span id="cb11-9">    img1 <span class="op" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">=</span> cv2.cvtColor(cv2.imread(os.path.join(IMG_DIR, <span class="ss" style="color: #20794D;
background-color: null;
font-style: inherit;">f'</span><span class="sc" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">{</span>img1_key<span class="sc" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">}</span><span class="ss" style="color: #20794D;
background-color: null;
font-style: inherit;">.png'</span>)), cv2.COLOR_BGR2RGB)</span>
<span id="cb11-10">    img2 <span class="op" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">=</span> cv2.cvtColor(cv2.imread(os.path.join(IMG_DIR, <span class="ss" style="color: #20794D;
background-color: null;
font-style: inherit;">f'</span><span class="sc" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">{</span>img2_key<span class="sc" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">}</span><span class="ss" style="color: #20794D;
background-color: null;
font-style: inherit;">.png'</span>)), cv2.COLOR_BGR2RGB)</span>
<span id="cb11-11"></span>
<span id="cb11-12">    idxs <span class="op" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">=</span> matcher.match_and_filter(kp1, kp2,</span>
<span id="cb11-13">                            descs1, descs2,</span>
<span id="cb11-14">                            im1shape<span class="op" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">=</span>img1.shape[:<span class="dv" style="color: #AD0000;
background-color: null;
font-style: inherit;">2</span>], </span>
<span id="cb11-15">                            im2shape<span class="op" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">=</span>img2.shape[:<span class="dv" style="color: #AD0000;
background-color: null;
font-style: inherit;">2</span>],</span>
<span id="cb11-16">                            o1<span class="op" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">=</span>a1.reshape(<span class="op" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">-</span><span class="dv" style="color: #AD0000;
background-color: null;
font-style: inherit;">1</span>),</span>
<span id="cb11-17">                            o2<span class="op" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">=</span>a2.reshape(<span class="op" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">-</span><span class="dv" style="color: #AD0000;
background-color: null;
font-style: inherit;">1</span>),</span>
<span id="cb11-18">                            s1<span class="op" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">=</span>s1.reshape(<span class="op" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">-</span><span class="dv" style="color: #AD0000;
background-color: null;
font-style: inherit;">1</span>),</span>
<span id="cb11-19">                            s2<span class="op" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">=</span>s2.reshape(<span class="op" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">-</span><span class="dv" style="color: #AD0000;
background-color: null;
font-style: inherit;">1</span>)).detach().cpu().numpy()</span>
<span id="cb11-20">    src_pts <span class="op" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">=</span> kp1[idxs[:,<span class="dv" style="color: #AD0000;
background-color: null;
font-style: inherit;">0</span>]]</span>
<span id="cb11-21">    dst_pts <span class="op" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">=</span> kp2[idxs[:,<span class="dv" style="color: #AD0000;
background-color: null;
font-style: inherit;">1</span>]]</span>
<span id="cb11-22">    </span>
<span id="cb11-23">    max_iters <span class="op" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">=</span> <span class="dv" style="color: #AD0000;
background-color: null;
font-style: inherit;">100000</span></span>
<span id="cb11-24">    <span class="cf" style="color: #003B4F;
background-color: null;
font-weight: bold;
font-style: inherit;">if</span> ds_name.lower() <span class="op" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">==</span> <span class="st" style="color: #20794D;
background-color: null;
font-style: inherit;">'phototourism'</span>:</span>
<span id="cb11-25">        inl_th <span class="op" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">=</span> <span class="fl" style="color: #AD0000;
background-color: null;
font-style: inherit;">0.5</span></span>
<span id="cb11-26">    <span class="cf" style="color: #003B4F;
background-color: null;
font-weight: bold;
font-style: inherit;">elif</span> ds_name.lower() <span class="op" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">==</span> <span class="st" style="color: #20794D;
background-color: null;
font-style: inherit;">'pragueparks'</span>:</span>
<span id="cb11-27">        inl_th <span class="op" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">=</span> <span class="fl" style="color: #AD0000;
background-color: null;
font-style: inherit;">1.5</span> </span>
<span id="cb11-28">    <span class="cf" style="color: #003B4F;
background-color: null;
font-weight: bold;
font-style: inherit;">elif</span> ds_name.lower() <span class="op" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">==</span> <span class="st" style="color: #20794D;
background-color: null;
font-style: inherit;">'googleurban'</span>:</span>
<span id="cb11-29">        inl_th <span class="op" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">=</span> <span class="fl" style="color: #AD0000;
background-color: null;
font-style: inherit;">0.75</span></span>
<span id="cb11-30">    <span class="cf" style="color: #003B4F;
background-color: null;
font-weight: bold;
font-style: inherit;">else</span>:</span>
<span id="cb11-31">        <span class="cf" style="color: #003B4F;
background-color: null;
font-weight: bold;
font-style: inherit;">raise</span> <span class="pp" style="color: #AD0000;
background-color: null;
font-style: inherit;">ValueError</span>(<span class="st" style="color: #20794D;
background-color: null;
font-style: inherit;">'Unknown dataset'</span>)</span>
<span id="cb11-32">    F, inliers_mask <span class="op" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">=</span> pydegensac.findFundamentalMatrix(src_pts, dst_pts, inl_th, <span class="fl" style="color: #AD0000;
background-color: null;
font-style: inherit;">0.999999</span>, max_iters)</span>
<span id="cb11-33">    out_idxs <span class="op" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">=</span> idxs[inliers_mask]</span>
<span id="cb11-34">    <span class="co" style="color: #5E5E5E;
background-color: null;
font-style: inherit;"># AdaLAM does not provide confidence score, so we will create dummy one</span></span>
<span id="cb11-35">    dists <span class="op" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">=</span> np.ones_like(out_idxs)[:,<span class="dv" style="color: #AD0000;
background-color: null;
font-style: inherit;">0</span>] </span>
<span id="cb11-36">    <span class="cf" style="color: #003B4F;
background-color: null;
font-weight: bold;
font-style: inherit;">return</span> dists, out_idxs</span>
<span id="cb11-37"></span>
<span id="cb11-38">dists, idxs <span class="op" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">=</span> match_adalam_with_degensac(img1_key, img2_key, kps, angles, scales, descs)</span>
<span id="cb11-39">draw_matches(img1_key, img2_key, dists, idxs,  kps, angles, scales, descs)</span></code></pre></div>
<div class="cell-output cell-output-display">
<pre><code>&lt;Figure size 432x288 with 0 Axes&gt;</code></pre>
</div>
<div class="cell-output cell-output-display">
<div>
<figure class="figure">
<p><img src="https://ducha-aiki.github.io/wide-baseline-stereo-blog/posts/2021-05-27-submitting-to-IMC2021-with-custom-matcher_files/figure-html/cell-6-output-2.png" class="img-fluid figure-img"></p>
</figure>
</div>
</div>
</div>
<p>The last, but not least: we have to transpose output indexes, as benchmark expects them in the shape <code>[2 x B]</code>. Do not ask why :)</p>
<p>So, we are ready to process and save our matches. I have a bit optimized the loading process, but not much. Do not expect blazing speed :) The full script is accesible <a href="https://github.com/ducha-aiki/imc2021-sample-kornia-submission/blob/master/extract_adalam_degensac_matches.py">here</a>.</p>
<div id="cell-14" class="cell">
<div class="sourceCode cell-code" id="cb13" style="background: #f1f3f5;"><pre class="sourceCode python code-with-copy"><code class="sourceCode python"><span id="cb13-1"><span class="im" style="color: #00769E;
background-color: null;
font-style: inherit;">import</span> os</span>
<span id="cb13-2"><span class="im" style="color: #00769E;
background-color: null;
font-style: inherit;">import</span> h5py</span>
<span id="cb13-3"><span class="im" style="color: #00769E;
background-color: null;
font-style: inherit;">from</span> tqdm <span class="im" style="color: #00769E;
background-color: null;
font-style: inherit;">import</span> tqdm</span>
<span id="cb13-4"><span class="im" style="color: #00769E;
background-color: null;
font-style: inherit;">from</span> PIL <span class="im" style="color: #00769E;
background-color: null;
font-style: inherit;">import</span> Image</span>
<span id="cb13-5"></span>
<span id="cb13-6"><span class="kw" style="color: #003B4F;
background-color: null;
font-weight: bold;
font-style: inherit;">def</span> match_adalam_with_degensac(kp1, kp2, s1, s2, a1, a2, descs1, descs2,</span>
<span id="cb13-7">                               h1, w1, h2, w2, ds_name<span class="op" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">=</span><span class="st" style="color: #20794D;
background-color: null;
font-style: inherit;">'phototourism'</span>):</span>
<span id="cb13-8">    matcher <span class="op" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">=</span> AdalamFilter()</span>
<span id="cb13-9">    </span>
<span id="cb13-10">    idxs <span class="op" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">=</span> matcher.match_and_filter(kp1, kp2,</span>
<span id="cb13-11">                            descs1, descs2,</span>
<span id="cb13-12">                            im1shape<span class="op" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">=</span>(h1,w1), </span>
<span id="cb13-13">                            im2shape<span class="op" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">=</span>(h2,w2),</span>
<span id="cb13-14">                            o1<span class="op" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">=</span>a1.reshape(<span class="op" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">-</span><span class="dv" style="color: #AD0000;
background-color: null;
font-style: inherit;">1</span>),</span>
<span id="cb13-15">                            o2<span class="op" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">=</span>a2.reshape(<span class="op" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">-</span><span class="dv" style="color: #AD0000;
background-color: null;
font-style: inherit;">1</span>),</span>
<span id="cb13-16">                            s1<span class="op" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">=</span>s1.reshape(<span class="op" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">-</span><span class="dv" style="color: #AD0000;
background-color: null;
font-style: inherit;">1</span>),</span>
<span id="cb13-17">                            s2<span class="op" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">=</span>s2.reshape(<span class="op" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">-</span><span class="dv" style="color: #AD0000;
background-color: null;
font-style: inherit;">1</span>)).detach().cpu().numpy()</span>
<span id="cb13-18">    <span class="cf" style="color: #003B4F;
background-color: null;
font-weight: bold;
font-style: inherit;">if</span> <span class="bu" style="color: null;
background-color: null;
font-style: inherit;">len</span>(idxs) <span class="op" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">&lt;</span> <span class="dv" style="color: #AD0000;
background-color: null;
font-style: inherit;">7</span>:</span>
<span id="cb13-19">        <span class="cf" style="color: #003B4F;
background-color: null;
font-weight: bold;
font-style: inherit;">return</span> np.empty((<span class="dv" style="color: #AD0000;
background-color: null;
font-style: inherit;">0</span>,<span class="dv" style="color: #AD0000;
background-color: null;
font-style: inherit;">1</span>), dtype<span class="op" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">=</span>np.float32), np.empty((<span class="dv" style="color: #AD0000;
background-color: null;
font-style: inherit;">0</span>,<span class="dv" style="color: #AD0000;
background-color: null;
font-style: inherit;">2</span>), dtype<span class="op" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">=</span>np.int32)</span>
<span id="cb13-20">    src_pts <span class="op" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">=</span> kp1[idxs[:,<span class="dv" style="color: #AD0000;
background-color: null;
font-style: inherit;">0</span>]]</span>
<span id="cb13-21">    dst_pts <span class="op" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">=</span> kp2[idxs[:,<span class="dv" style="color: #AD0000;
background-color: null;
font-style: inherit;">1</span>]]</span>
<span id="cb13-22">    </span>
<span id="cb13-23">    max_iters <span class="op" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">=</span> <span class="dv" style="color: #AD0000;
background-color: null;
font-style: inherit;">100000</span></span>
<span id="cb13-24">    <span class="cf" style="color: #003B4F;
background-color: null;
font-weight: bold;
font-style: inherit;">if</span> ds_name.lower() <span class="op" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">==</span> <span class="st" style="color: #20794D;
background-color: null;
font-style: inherit;">'phototourism'</span>:</span>
<span id="cb13-25">        inl_th <span class="op" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">=</span> <span class="fl" style="color: #AD0000;
background-color: null;
font-style: inherit;">0.5</span></span>
<span id="cb13-26">    <span class="cf" style="color: #003B4F;
background-color: null;
font-weight: bold;
font-style: inherit;">elif</span> ds_name.lower() <span class="op" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">==</span> <span class="st" style="color: #20794D;
background-color: null;
font-style: inherit;">'pragueparks'</span>:</span>
<span id="cb13-27">        inl_th <span class="op" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">=</span> <span class="fl" style="color: #AD0000;
background-color: null;
font-style: inherit;">1.5</span> </span>
<span id="cb13-28">    <span class="cf" style="color: #003B4F;
background-color: null;
font-weight: bold;
font-style: inherit;">elif</span> ds_name.lower() <span class="op" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">==</span> <span class="st" style="color: #20794D;
background-color: null;
font-style: inherit;">'googleurban'</span>:</span>
<span id="cb13-29">        inl_th <span class="op" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">=</span> <span class="fl" style="color: #AD0000;
background-color: null;
font-style: inherit;">0.75</span></span>
<span id="cb13-30">    <span class="cf" style="color: #003B4F;
background-color: null;
font-weight: bold;
font-style: inherit;">else</span>:</span>
<span id="cb13-31">        <span class="cf" style="color: #003B4F;
background-color: null;
font-weight: bold;
font-style: inherit;">raise</span> <span class="pp" style="color: #AD0000;
background-color: null;
font-style: inherit;">ValueError</span>(<span class="st" style="color: #20794D;
background-color: null;
font-style: inherit;">'Unknown dataset'</span>)</span>
<span id="cb13-32">    F, inliers_mask <span class="op" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">=</span> pydegensac.findFundamentalMatrix(src_pts, dst_pts, inl_th, <span class="fl" style="color: #AD0000;
background-color: null;
font-style: inherit;">0.999999</span>, max_iters)</span>
<span id="cb13-33">    out_idxs <span class="op" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">=</span> idxs[inliers_mask]</span>
<span id="cb13-34">    <span class="co" style="color: #5E5E5E;
background-color: null;
font-style: inherit;"># AdaLAM does not provide confidence score, so we will create dummy one</span></span>
<span id="cb13-35">    dists <span class="op" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">=</span> np.ones_like(out_idxs)[:,<span class="dv" style="color: #AD0000;
background-color: null;
font-style: inherit;">0</span>] </span>
<span id="cb13-36">    <span class="cf" style="color: #003B4F;
background-color: null;
font-weight: bold;
font-style: inherit;">return</span> dists, out_idxs</span>
<span id="cb13-37"></span>
<span id="cb13-38"></span>
<span id="cb13-39">INPUT_DIR <span class="op" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">=</span> <span class="st" style="color: #20794D;
background-color: null;
font-style: inherit;">'../imc-2021-data'</span></span>
<span id="cb13-40">OUT_DIR <span class="op" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">=</span> <span class="st" style="color: #20794D;
background-color: null;
font-style: inherit;">'extracted/cv2-dog-affnet-hardnet8'</span></span>
<span id="cb13-41"></span>
<span id="cb13-42">os.makedirs(OUT_DIR, exist_ok<span class="op" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">=</span><span class="va" style="color: #111111;
background-color: null;
font-style: inherit;">True</span>)</span>
<span id="cb13-43">datasets <span class="op" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">=</span> os.listdir(INPUT_DIR)</span>
<span id="cb13-44">datasets <span class="op" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">=</span> [<span class="st" style="color: #20794D;
background-color: null;
font-style: inherit;">'googleurban'</span>]</span>
<span id="cb13-45"><span class="cf" style="color: #003B4F;
background-color: null;
font-weight: bold;
font-style: inherit;">for</span> ds <span class="kw" style="color: #003B4F;
background-color: null;
font-weight: bold;
font-style: inherit;">in</span> datasets:</span>
<span id="cb13-46">    ds_in_path <span class="op" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">=</span> os.path.join(INPUT_DIR, ds)</span>
<span id="cb13-47">    ds_out_path <span class="op" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">=</span> os.path.join(OUT_DIR, ds)</span>
<span id="cb13-48">    os.makedirs(ds_out_path, exist_ok<span class="op" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">=</span><span class="va" style="color: #111111;
background-color: null;
font-style: inherit;">True</span>)</span>
<span id="cb13-49">    seqs <span class="op" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">=</span> os.listdir(ds_in_path)</span>
<span id="cb13-50">    <span class="cf" style="color: #003B4F;
background-color: null;
font-weight: bold;
font-style: inherit;">for</span> seq <span class="kw" style="color: #003B4F;
background-color: null;
font-weight: bold;
font-style: inherit;">in</span> seqs:</span>
<span id="cb13-51">        <span class="cf" style="color: #003B4F;
background-color: null;
font-weight: bold;
font-style: inherit;">if</span> os.path.isdir(os.path.join(ds_in_path, seq, <span class="st" style="color: #20794D;
background-color: null;
font-style: inherit;">'set_100'</span>)):</span>
<span id="cb13-52">            seq_in_path <span class="op" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">=</span> os.path.join(ds_in_path, seq, <span class="st" style="color: #20794D;
background-color: null;
font-style: inherit;">'set_100'</span>, <span class="st" style="color: #20794D;
background-color: null;
font-style: inherit;">'images'</span>)</span>
<span id="cb13-53">        <span class="cf" style="color: #003B4F;
background-color: null;
font-weight: bold;
font-style: inherit;">else</span>:</span>
<span id="cb13-54">            seq_in_path <span class="op" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">=</span> os.path.join(ds_in_path, seq)</span>
<span id="cb13-55">        seq_out_path <span class="op" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">=</span> os.path.join(ds_out_path, seq)</span>
<span id="cb13-56">        kps <span class="op" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">=</span> load_h5(os.path.join(seq_out_path, <span class="st" style="color: #20794D;
background-color: null;
font-style: inherit;">'keypoints.h5'</span>))</span>
<span id="cb13-57">        angles <span class="op" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">=</span> load_h5(os.path.join(seq_out_path, <span class="st" style="color: #20794D;
background-color: null;
font-style: inherit;">'angles.h5'</span>))</span>
<span id="cb13-58">        scales <span class="op" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">=</span> load_h5(os.path.join(seq_out_path, <span class="st" style="color: #20794D;
background-color: null;
font-style: inherit;">'scales.h5'</span>))</span>
<span id="cb13-59">        descs <span class="op" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">=</span> load_h5(os.path.join(seq_out_path, <span class="st" style="color: #20794D;
background-color: null;
font-style: inherit;">'descriptors.h5'</span>))</span>
<span id="cb13-60">        img_fnames <span class="op" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">=</span> <span class="bu" style="color: null;
background-color: null;
font-style: inherit;">sorted</span>(os.listdir(seq_in_path))[::<span class="op" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">-</span><span class="dv" style="color: #AD0000;
background-color: null;
font-style: inherit;">1</span>]</span>
<span id="cb13-61">        num_matches <span class="op" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">=</span> []</span>
<span id="cb13-62">        <span class="cf" style="color: #003B4F;
background-color: null;
font-weight: bold;
font-style: inherit;">with</span> h5py.File(<span class="ss" style="color: #20794D;
background-color: null;
font-style: inherit;">f'</span><span class="sc" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">{</span>seq_out_path<span class="sc" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">}</span><span class="ss" style="color: #20794D;
background-color: null;
font-style: inherit;">/matches_stereo_0.h5'</span>, <span class="st" style="color: #20794D;
background-color: null;
font-style: inherit;">'w'</span>) <span class="im" style="color: #00769E;
background-color: null;
font-style: inherit;">as</span> f_m:</span>
<span id="cb13-63">            <span class="cf" style="color: #003B4F;
background-color: null;
font-weight: bold;
font-style: inherit;">for</span> i1, img1_fname <span class="kw" style="color: #003B4F;
background-color: null;
font-weight: bold;
font-style: inherit;">in</span> tqdm(<span class="bu" style="color: null;
background-color: null;
font-style: inherit;">enumerate</span>(img_fnames)):</span>
<span id="cb13-64">                img1_key <span class="op" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">=</span> os.path.splitext(os.path.basename(img1_fname))[<span class="dv" style="color: #AD0000;
background-color: null;
font-style: inherit;">0</span>]</span>
<span id="cb13-65">                img1_fname_full <span class="op" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">=</span> os.path.join(seq_in_path, img1_fname)</span>
<span id="cb13-66">                img1 <span class="op" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">=</span> Image.<span class="bu" style="color: null;
background-color: null;
font-style: inherit;">open</span>(img1_fname_full)</span>
<span id="cb13-67">                w1, h1 <span class="op" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">=</span> img1.size</span>
<span id="cb13-68">                kp1, s1, a1, descs1 <span class="op" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">=</span> get_data(kps, angles, scales, descs, img1_key)</span>
<span id="cb13-69">                <span class="cf" style="color: #003B4F;
background-color: null;
font-weight: bold;
font-style: inherit;">for</span> img2_fname <span class="kw" style="color: #003B4F;
background-color: null;
font-weight: bold;
font-style: inherit;">in</span> img_fnames[i1<span class="op" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">+</span><span class="dv" style="color: #AD0000;
background-color: null;
font-style: inherit;">1</span>:]:</span>
<span id="cb13-70">                    img2_key <span class="op" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">=</span> os.path.splitext(os.path.basename(img2_fname))[<span class="dv" style="color: #AD0000;
background-color: null;
font-style: inherit;">0</span>]</span>
<span id="cb13-71">                    img2_fname_full <span class="op" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">=</span> os.path.join(seq_in_path, img2_fname)</span>
<span id="cb13-72">                    img2 <span class="op" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">=</span> Image.<span class="bu" style="color: null;
background-color: null;
font-style: inherit;">open</span>(img2_fname_full)</span>
<span id="cb13-73">                    w2, h2 <span class="op" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">=</span> img2.size</span>
<span id="cb13-74">                    match_key <span class="op" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">=</span> <span class="ss" style="color: #20794D;
background-color: null;
font-style: inherit;">f'</span><span class="sc" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">{</span>img1_key<span class="sc" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">}</span><span class="ss" style="color: #20794D;
background-color: null;
font-style: inherit;">-</span><span class="sc" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">{</span>img2_key<span class="sc" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">}</span><span class="ss" style="color: #20794D;
background-color: null;
font-style: inherit;">'</span></span>
<span id="cb13-75">                    kp2, s2, a2, descs2 <span class="op" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">=</span> get_data(kps, angles, scales, descs, img2_key)</span>
<span id="cb13-76">                    _, idxs <span class="op" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">=</span> match_adalam_with_degensac(kp1, kp2, s1, s2, a1, a2, descs1, descs2,</span>
<span id="cb13-77">                                   h1, w1, h2, w2, ds_name<span class="op" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">=</span>ds)</span>
<span id="cb13-78">                    num_matches.append(<span class="bu" style="color: null;
background-color: null;
font-style: inherit;">len</span>(idxs))</span>
<span id="cb13-79">                    <span class="cf" style="color: #003B4F;
background-color: null;
font-weight: bold;
font-style: inherit;">if</span> <span class="bu" style="color: null;
background-color: null;
font-style: inherit;">len</span>(idxs) <span class="op" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">==</span> <span class="dv" style="color: #AD0000;
background-color: null;
font-style: inherit;">0</span>:</span>
<span id="cb13-80">                        idxs <span class="op" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">=</span> np.empty([<span class="dv" style="color: #AD0000;
background-color: null;
font-style: inherit;">0</span>, <span class="dv" style="color: #AD0000;
background-color: null;
font-style: inherit;">2</span>], dtype<span class="op" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">=</span>np.int32)</span>
<span id="cb13-81">                    idxs <span class="op" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">=</span> idxs.T</span>
<span id="cb13-82">                    <span class="cf" style="color: #003B4F;
background-color: null;
font-weight: bold;
font-style: inherit;">assert</span> idxs.shape[<span class="dv" style="color: #AD0000;
background-color: null;
font-style: inherit;">0</span>] <span class="op" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">==</span> <span class="dv" style="color: #AD0000;
background-color: null;
font-style: inherit;">2</span></span>
<span id="cb13-83">                    f_m[match_key] <span class="op" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">=</span> idxs</span>
<span id="cb13-84">            <span class="bu" style="color: null;
background-color: null;
font-style: inherit;">print</span>(<span class="ss" style="color: #20794D;
background-color: null;
font-style: inherit;">f'Finished processing "</span><span class="sc" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">{</span>ds<span class="sc" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">}</span><span class="ss" style="color: #20794D;
background-color: null;
font-style: inherit;">/</span><span class="sc" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">{</span>seq<span class="sc" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">}</span><span class="ss" style="color: #20794D;
background-color: null;
font-style: inherit;">" -&gt; </span><span class="sc" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">{</span>np<span class="sc" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">.</span>array(num_matches)<span class="sc" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">.</span>mean()<span class="sc" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">}</span><span class="ss" style="color: #20794D;
background-color: null;
font-style: inherit;"> matches/image'</span>)</span>
<span id="cb13-85">        <span class="co" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">#We can use a single match file for multiview and stereo, but let's pretend that we have different ones</span></span>
<span id="cb13-86">        copyfile(<span class="ss" style="color: #20794D;
background-color: null;
font-style: inherit;">f'</span><span class="sc" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">{</span>seq_out_path<span class="sc" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">}</span><span class="ss" style="color: #20794D;
background-color: null;
font-style: inherit;">/matches_stereo_0.h5'</span>, <span class="ss" style="color: #20794D;
background-color: null;
font-style: inherit;">f'</span><span class="sc" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">{</span>seq_out_path<span class="sc" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">}</span><span class="ss" style="color: #20794D;
background-color: null;
font-style: inherit;">/matches_multiview.h5'</span>)</span></code></pre></div>
</div>
<p>Matching will take us couple of hours on GPU.</p>
</section>
<section id="creating-config-json-file" class="level2">
<h2 class="anchored" data-anchor-id="creating-config-json-file">Creating config json file</h2>
<p>In addition to features and matches, we should submit a config file, which tells the benchmark, how the features should be matched and which RANSAC (no options, only <code>cv2-8pt</code> for using custom matches) we prefer. So, we will create two config files - one, which standard matcher and second one, which uses AdaLAM.</p>
<p>First part of the config is <code>metadata</code> – information about the method and authors. If your method is under review, you may want to set flag <code>publish_anonymously</code> to <code>True</code>.</p>
<div id="cell-18" class="cell">
<div class="sourceCode cell-code" id="cb14" style="background: #f1f3f5;"><pre class="sourceCode python code-with-copy"><code class="sourceCode python"><span id="cb14-1">metadata_dict <span class="op" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">=</span>  {</span>
<span id="cb14-2">    <span class="st" style="color: #20794D;
background-color: null;
font-style: inherit;">"publish_anonymously"</span>: <span class="va" style="color: #111111;
background-color: null;
font-style: inherit;">False</span>,</span>
<span id="cb14-3">    <span class="st" style="color: #20794D;
background-color: null;
font-style: inherit;">"authors"</span>: <span class="st" style="color: #20794D;
background-color: null;
font-style: inherit;">"Dmytro Mishkin, Milan Pultar and kornia team"</span>,</span>
<span id="cb14-4">    <span class="st" style="color: #20794D;
background-color: null;
font-style: inherit;">"contact_email"</span>: <span class="st" style="color: #20794D;
background-color: null;
font-style: inherit;">"ducha.aiki@gmail.com"</span>,</span>
<span id="cb14-5">    <span class="st" style="color: #20794D;
background-color: null;
font-style: inherit;">"method_name"</span>: <span class="st" style="color: #20794D;
background-color: null;
font-style: inherit;">"CV-DoG-AffNet-HardNet8-AdaLAM-DEGENSAC"</span>,</span>
<span id="cb14-6">    <span class="st" style="color: #20794D;
background-color: null;
font-style: inherit;">"method_description"</span>: </span>
<span id="cb14-7">    <span class="vs" style="color: #20794D;
background-color: null;
font-style: inherit;">r"""OpeCV SIFT keypoints 8000 features, followed by the AffNet normalization </span></span>
<span id="cb14-8"><span class="vs" style="color: #20794D;
background-color: null;
font-style: inherit;">    and HardNet8 descriptor as implemented in kornia.</span></span>
<span id="cb14-9"><span class="vs" style="color: #20794D;
background-color: null;
font-style: inherit;">    Matched using AdaLAM with DEGENSAC"""</span>,</span>
<span id="cb14-10">    <span class="st" style="color: #20794D;
background-color: null;
font-style: inherit;">"link_to_website"</span>: <span class="st" style="color: #20794D;
background-color: null;
font-style: inherit;">"https://github.com/kornia/kornia"</span>,</span>
<span id="cb14-11">    <span class="st" style="color: #20794D;
background-color: null;
font-style: inherit;">"link_to_pdf"</span>: <span class="st" style="color: #20794D;
background-color: null;
font-style: inherit;">"https://arxiv.org/abs/2007.09699"</span></span>
<span id="cb14-12">}</span></code></pre></div>
</div>
<p>Second part is <code>config_common</code>: it tells the benchmark, which keypoints and descriptors you use. We will also need this names when importing our features during tuning on the validation set.</p>
<div id="cell-20" class="cell">
<div class="sourceCode cell-code" id="cb15" style="background: #f1f3f5;"><pre class="sourceCode python code-with-copy"><code class="sourceCode python"><span id="cb15-1">config_common_dict <span class="op" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">=</span>  {<span class="st" style="color: #20794D;
background-color: null;
font-style: inherit;">"json_label"</span>: <span class="st" style="color: #20794D;
background-color: null;
font-style: inherit;">"dog-affnet-hardnet8-degensac-adalam"</span>,</span>
<span id="cb15-2">    <span class="st" style="color: #20794D;
background-color: null;
font-style: inherit;">"keypoint"</span>: <span class="st" style="color: #20794D;
background-color: null;
font-style: inherit;">"cv2dog"</span>,</span>
<span id="cb15-3">    <span class="st" style="color: #20794D;
background-color: null;
font-style: inherit;">"descriptor"</span>: <span class="st" style="color: #20794D;
background-color: null;
font-style: inherit;">"affnethardnet8"</span>,</span>
<span id="cb15-4">    <span class="st" style="color: #20794D;
background-color: null;
font-style: inherit;">"num_keypoints"</span>: <span class="dv" style="color: #AD0000;
background-color: null;
font-style: inherit;">8000</span>}</span></code></pre></div>
</div>
<p>Finally, we have to specify robust geometry estimation method. We have no other choice than <code>cv2-8pt</code>, as pre-filtering with DEGENSAC is already performed.</p>
<div id="cell-24" class="cell">
<div class="sourceCode cell-code" id="cb16" style="background: #f1f3f5;"><pre class="sourceCode python code-with-copy"><code class="sourceCode python"><span id="cb16-1"><span class="im" style="color: #00769E;
background-color: null;
font-style: inherit;">from</span> copy <span class="im" style="color: #00769E;
background-color: null;
font-style: inherit;">import</span> deepcopy</span>
<span id="cb16-2"></span>
<span id="cb16-3"></span>
<span id="cb16-4">geom_template_dict <span class="op" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">=</span>  {<span class="st" style="color: #20794D;
background-color: null;
font-style: inherit;">"method"</span>: <span class="st" style="color: #20794D;
background-color: null;
font-style: inherit;">"cv2-8pt"</span></span>
<span id="cb16-5">            }</span></code></pre></div>
</div>
<p>Let’s assemble and save our base config.</p>
<div id="cell-26" class="cell">
<div class="sourceCode cell-code" id="cb17" style="background: #f1f3f5;"><pre class="sourceCode python code-with-copy"><code class="sourceCode python"><span id="cb17-1"><span class="im" style="color: #00769E;
background-color: null;
font-style: inherit;">import</span> json</span>
<span id="cb17-2"></span>
<span id="cb17-3">base_config <span class="op" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">=</span>  {</span>
<span id="cb17-4">    <span class="st" style="color: #20794D;
background-color: null;
font-style: inherit;">"metadata"</span>: metadata_dict,</span>
<span id="cb17-5">    <span class="st" style="color: #20794D;
background-color: null;
font-style: inherit;">"config_common"</span>: config_common_dict,</span>
<span id="cb17-6">    <span class="st" style="color: #20794D;
background-color: null;
font-style: inherit;">"config_phototourism_stereo"</span>: {</span>
<span id="cb17-7">        <span class="st" style="color: #20794D;
background-color: null;
font-style: inherit;">"use_custom_matches"</span>: <span class="va" style="color: #111111;
background-color: null;
font-style: inherit;">True</span>,</span>
<span id="cb17-8">        <span class="st" style="color: #20794D;
background-color: null;
font-style: inherit;">"custom_matches_name"</span>: <span class="st" style="color: #20794D;
background-color: null;
font-style: inherit;">"adalam"</span>,</span>
<span id="cb17-9">        <span class="st" style="color: #20794D;
background-color: null;
font-style: inherit;">"geom"</span>: deepcopy(geom_template_dict)</span>
<span id="cb17-10">        },</span>
<span id="cb17-11">    <span class="st" style="color: #20794D;
background-color: null;
font-style: inherit;">"config_phototourism_multiview"</span>: {</span>
<span id="cb17-12">        <span class="st" style="color: #20794D;
background-color: null;
font-style: inherit;">"use_custom_matches"</span>: <span class="va" style="color: #111111;
background-color: null;
font-style: inherit;">True</span>,</span>
<span id="cb17-13">        <span class="st" style="color: #20794D;
background-color: null;
font-style: inherit;">"custom_matches_name"</span>: <span class="st" style="color: #20794D;
background-color: null;
font-style: inherit;">"adalam"</span>,</span>
<span id="cb17-14">        <span class="st" style="color: #20794D;
background-color: null;
font-style: inherit;">"colmap"</span>: {}},</span>
<span id="cb17-15">    </span>
<span id="cb17-16">    <span class="st" style="color: #20794D;
background-color: null;
font-style: inherit;">"config_pragueparks_stereo"</span>: {</span>
<span id="cb17-17">        <span class="st" style="color: #20794D;
background-color: null;
font-style: inherit;">"use_custom_matches"</span>: <span class="va" style="color: #111111;
background-color: null;
font-style: inherit;">True</span>,</span>
<span id="cb17-18">        <span class="st" style="color: #20794D;
background-color: null;
font-style: inherit;">"custom_matches_name"</span>: <span class="st" style="color: #20794D;
background-color: null;
font-style: inherit;">"adalam"</span>,</span>
<span id="cb17-19">        <span class="st" style="color: #20794D;
background-color: null;
font-style: inherit;">"geom"</span>: deepcopy(geom_template_dict)</span>
<span id="cb17-20">        },</span>
<span id="cb17-21">    <span class="st" style="color: #20794D;
background-color: null;
font-style: inherit;">"config_pragueparks_multiview"</span>: {</span>
<span id="cb17-22">        <span class="st" style="color: #20794D;
background-color: null;
font-style: inherit;">"use_custom_matches"</span>: <span class="va" style="color: #111111;
background-color: null;
font-style: inherit;">True</span>,</span>
<span id="cb17-23">        <span class="st" style="color: #20794D;
background-color: null;
font-style: inherit;">"custom_matches_name"</span>: <span class="st" style="color: #20794D;
background-color: null;
font-style: inherit;">"adalam"</span>,</span>
<span id="cb17-24">        <span class="st" style="color: #20794D;
background-color: null;
font-style: inherit;">"colmap"</span>: {}},</span>
<span id="cb17-25">    <span class="st" style="color: #20794D;
background-color: null;
font-style: inherit;">"config_googleurban_stereo"</span>: {</span>
<span id="cb17-26">        <span class="st" style="color: #20794D;
background-color: null;
font-style: inherit;">"use_custom_matches"</span>: <span class="va" style="color: #111111;
background-color: null;
font-style: inherit;">True</span>,</span>
<span id="cb17-27">        <span class="st" style="color: #20794D;
background-color: null;
font-style: inherit;">"custom_matches_name"</span>: <span class="st" style="color: #20794D;
background-color: null;
font-style: inherit;">"adalam"</span>,</span>
<span id="cb17-28">        <span class="st" style="color: #20794D;
background-color: null;
font-style: inherit;">"geom"</span>: deepcopy(geom_template_dict)</span>
<span id="cb17-29">        },</span>
<span id="cb17-30">    <span class="st" style="color: #20794D;
background-color: null;
font-style: inherit;">"config_googleurban_multiview"</span>: {</span>
<span id="cb17-31">        <span class="st" style="color: #20794D;
background-color: null;
font-style: inherit;">"use_custom_matches"</span>: <span class="va" style="color: #111111;
background-color: null;
font-style: inherit;">True</span>,</span>
<span id="cb17-32">        <span class="st" style="color: #20794D;
background-color: null;
font-style: inherit;">"custom_matches_name"</span>: <span class="st" style="color: #20794D;
background-color: null;
font-style: inherit;">"adalam"</span>,</span>
<span id="cb17-33">        <span class="st" style="color: #20794D;
background-color: null;
font-style: inherit;">"colmap"</span>: {}}</span>
<span id="cb17-34">}</span></code></pre></div>
</div>
<p>Finally, benchmark expects multiple configs, so we have to create a list, and then we can save our config</p>
<div id="cell-28" class="cell">
<div class="sourceCode cell-code" id="cb18" style="background: #f1f3f5;"><pre class="sourceCode python code-with-copy"><code class="sourceCode python"><span id="cb18-1"><span class="im" style="color: #00769E;
background-color: null;
font-style: inherit;">import</span> json </span>
<span id="cb18-2"></span>
<span id="cb18-3"><span class="cf" style="color: #003B4F;
background-color: null;
font-weight: bold;
font-style: inherit;">with</span> <span class="bu" style="color: null;
background-color: null;
font-style: inherit;">open</span>(<span class="st" style="color: #20794D;
background-color: null;
font-style: inherit;">'base_config_adalam.json'</span>, <span class="st" style="color: #20794D;
background-color: null;
font-style: inherit;">'w'</span>) <span class="im" style="color: #00769E;
background-color: null;
font-style: inherit;">as</span> f:</span>
<span id="cb18-4">    json.dump([base_config], f, indent<span class="op" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">=</span><span class="dv" style="color: #AD0000;
background-color: null;
font-style: inherit;">2</span>)</span></code></pre></div>
</div>
</section>
<section id="preliminary-evaluation" class="level1">
<h1>Preliminary evaluation</h1>
<p>Now let’s check how our features perform on validation set. We have to import our feature to the benchmark and run the benchmark.</p>
<p>I will cheat a little bit here and skip the multiview evaluation. The reason is that it requires <a href="https://colmap.github.io">colmap</a>, which might be not easy to install.</p>
<section id="importing-features" class="level3">
<h3 class="anchored" data-anchor-id="importing-features">Importing features</h3>
<p>Here we have to provide the same <code>keypoint</code> and <code>descriptor</code> names, as we wrote in json config. The rest of arguments are straightforward: path to features, json, etc.</p>
<div class="sourceCode" id="cb19" style="background: #f1f3f5;"><pre class="sourceCode bash code-with-copy"><code class="sourceCode bash"><span id="cb19-1"><span class="bu" style="color: null;
background-color: null;
font-style: inherit;">cd</span> ../image-matching-benchmark/</span>
<span id="cb19-2"><span class="ex" style="color: null;
background-color: null;
font-style: inherit;">python</span> <span class="at" style="color: #657422;
background-color: null;
font-style: inherit;">-utt</span> import_features.py <span class="at" style="color: #657422;
background-color: null;
font-style: inherit;">--kp_name</span> cv2dog <span class="at" style="color: #657422;
background-color: null;
font-style: inherit;">--desc_name</span> affnethardnet8 <span class="at" style="color: #657422;
background-color: null;
font-style: inherit;">--num_keypoints</span> 8000 <span class="at" style="color: #657422;
background-color: null;
font-style: inherit;">--path_features</span> ../imc2021-sample-kornia-submission/extracted/cv2-dog-affnet-hardnet8  <span class="at" style="color: #657422;
background-color: null;
font-style: inherit;">--path_results</span> ../benchmark-results <span class="at" style="color: #657422;
background-color: null;
font-style: inherit;">--subset</span> both  <span class="at" style="color: #657422;
background-color: null;
font-style: inherit;">--is_challenge</span> false <span class="at" style="color: #657422;
background-color: null;
font-style: inherit;">--path_json</span> ../imc2021-sample-kornia-submission/base_config_adalam.json  <span class="at" style="color: #657422;
background-color: null;
font-style: inherit;">--datasets</span> phototourism googleurban  <span class="at" style="color: #657422;
background-color: null;
font-style: inherit;">--match_name</span> adalam</span></code></pre></div>
</section>
<section id="running-the-evaluation" class="level3">
<h3 class="anchored" data-anchor-id="running-the-evaluation">Running the evaluation</h3>
<p>Now we are ready to run the evaluation. Note, that now we are running multiview evaluation as well. So, let us check if colmap is available.</p>
<div class="sourceCode" id="cb20" style="background: #f1f3f5;"><pre class="sourceCode bash code-with-copy"><code class="sourceCode bash"><span id="cb20-1"><span class="ex" style="color: null;
background-color: null;
font-style: inherit;">colmap</span> <span class="at" style="color: #657422;
background-color: null;
font-style: inherit;">--help</span></span></code></pre></div>
<p>The output should be something like that:</p>
<pre><code>COLMAP 3.6 -- Structure-from-Motion and Multi-View Stereo
          (Commit Unknown on Unknown with CUDA)
          </code></pre>
<p>OK, now time to run the evaluation</p>
<div class="sourceCode" id="cb22" style="background: #f1f3f5;"><pre class="sourceCode bash code-with-copy"><code class="sourceCode bash"><span id="cb22-1"><span class="ex" style="color: null;
background-color: null;
font-style: inherit;">python</span> <span class="at" style="color: #657422;
background-color: null;
font-style: inherit;">-utt</span> run.py <span class="at" style="color: #657422;
background-color: null;
font-style: inherit;">--run_mode</span><span class="op" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">=</span>interactive <span class="at" style="color: #657422;
background-color: null;
font-style: inherit;">--json_method</span><span class="op" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">=</span>../imc2021-sample-kornia-submission/base_config_adalam.json <span class="at" style="color: #657422;
background-color: null;
font-style: inherit;">--subset</span><span class="op" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">=</span>val  <span class="at" style="color: #657422;
background-color: null;
font-style: inherit;">--path_data</span> ../imc-2021-data/ <span class="at" style="color: #657422;
background-color: null;
font-style: inherit;">--path_results</span> ../benchmark-results <span class="at" style="color: #657422;
background-color: null;
font-style: inherit;">--is_challenge</span> false</span></code></pre></div>
<p>After a while (an <strong>6 hours</strong> for 32 cores machine), the process will finish and you will see the following log message:</p>
<pre><code> -- Saving to: "packed-val/dog-affnet-hardnet8-degensac-adalam.json"</code></pre>
<p>To compare with results without AdaLAM, let’s also run multiview evaluation for the previous setup. I did not do that, because of the time it takes.</p>
<div class="sourceCode" id="cb24" style="background: #f1f3f5;"><pre class="sourceCode bash code-with-copy"><code class="sourceCode bash"><span id="cb24-1"></span>
<span id="cb24-2"><span class="kw" style="color: #003B4F;
background-color: null;
font-weight: bold;
font-style: inherit;">```</span><span class="fu" style="color: #4758AB;
background-color: null;
font-style: inherit;">bash</span></span>
<span id="cb24-3"><span class="ex" style="color: null;
background-color: null;
font-style: inherit;">python</span> <span class="at" style="color: #657422;
background-color: null;
font-style: inherit;">-utt</span> run.py <span class="at" style="color: #657422;
background-color: null;
font-style: inherit;">--run_mode</span><span class="op" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">=</span>interactive <span class="at" style="color: #657422;
background-color: null;
font-style: inherit;">--json_method</span><span class="op" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">=</span>../imc2021-sample-kornia-submission/base_config.json <span class="at" style="color: #657422;
background-color: null;
font-style: inherit;">--subset</span><span class="op" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">=</span>val  <span class="at" style="color: #657422;
background-color: null;
font-style: inherit;">--path_data</span> ../imc-2021-data/ <span class="at" style="color: #657422;
background-color: null;
font-style: inherit;">--path_results</span> ../benchmark-results <span class="at" style="color: #657422;
background-color: null;
font-style: inherit;">--is_challenge</span> false</span></code></pre></div>
</section>
<section id="reading-results" class="level3">
<h3 class="anchored" data-anchor-id="reading-results">Reading results</h3>
<p>Json file with evaluation results are saved to <code>image-matching-benchmark/packed-val/dog-affnet-hardnet8-degensac.json</code> and <code>image-matching-benchmark/packed-val/dog-affnet-hardnet8-degensac-adalam.json</code> , and some visualizations – to <code>../benchmark-visualization/png</code>.</p>
<p>First, we come back to our <code>imc2021-sample-kornia-submission</code> directory:</p>
<div class="sourceCode" id="cb25" style="background: #f1f3f5;"><pre class="sourceCode bash code-with-copy"><code class="sourceCode bash"><span id="cb25-1"><span class="bu" style="color: null;
background-color: null;
font-style: inherit;">cd</span> ../imc2021-sample-kornia-submission</span></code></pre></div>
<p>Metric, which are used for the competition is mean average accuracy (mAA) at visibility threshold 0.1</p>
<div id="cell-34" class="cell">
<div class="sourceCode cell-code" id="cb26" style="background: #f1f3f5;"><pre class="sourceCode python code-with-copy"><code class="sourceCode python"></code></pre></div>
<div class="cell-output cell-output-display">
<pre><code>{'mean': 0.0, 'std_scenes': 0.0, 'std_runs': 0.0}</code></pre>
</div>
</div>
<div id="cell-35" class="cell">
<div class="sourceCode cell-code" id="cb28" style="background: #f1f3f5;"><pre class="sourceCode python code-with-copy"><code class="sourceCode python"><span id="cb28-1"><span class="im" style="color: #00769E;
background-color: null;
font-style: inherit;">import</span> os</span>
<span id="cb28-2"></span>
<span id="cb28-3">hashname<span class="op" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">=</span><span class="st" style="color: #20794D;
background-color: null;
font-style: inherit;">'dog-affnet-hardnet8-degensac'</span></span>
<span id="cb28-4">res_fname <span class="op" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">=</span> os.path.join(<span class="st" style="color: #20794D;
background-color: null;
font-style: inherit;">'../image-matching-benchmark/packed-val'</span>, <span class="ss" style="color: #20794D;
background-color: null;
font-style: inherit;">f'</span><span class="sc" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">{</span>hashname<span class="sc" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">}</span><span class="ss" style="color: #20794D;
background-color: null;
font-style: inherit;">.json'</span>)</span>
<span id="cb28-5"></span>
<span id="cb28-6"><span class="cf" style="color: #003B4F;
background-color: null;
font-weight: bold;
font-style: inherit;">with</span> <span class="bu" style="color: null;
background-color: null;
font-style: inherit;">open</span>(res_fname, <span class="st" style="color: #20794D;
background-color: null;
font-style: inherit;">'r'</span>) <span class="im" style="color: #00769E;
background-color: null;
font-style: inherit;">as</span> f:</span>
<span id="cb28-7">    results <span class="op" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">=</span> json.load(f)</span>
<span id="cb28-8"></span>
<span id="cb28-9">submission_name <span class="op" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">=</span> results[<span class="st" style="color: #20794D;
background-color: null;
font-style: inherit;">'config'</span>][<span class="st" style="color: #20794D;
background-color: null;
font-style: inherit;">'metadata'</span>][<span class="st" style="color: #20794D;
background-color: null;
font-style: inherit;">'method_name'</span>]</span>
<span id="cb28-10">datasets <span class="op" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">=</span> [<span class="st" style="color: #20794D;
background-color: null;
font-style: inherit;">'phototourism'</span>, <span class="st" style="color: #20794D;
background-color: null;
font-style: inherit;">'pragueparks'</span>, <span class="st" style="color: #20794D;
background-color: null;
font-style: inherit;">'googleurban'</span>]</span>
<span id="cb28-11">tasks <span class="op" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">=</span> [<span class="st" style="color: #20794D;
background-color: null;
font-style: inherit;">'stereo'</span>, <span class="st" style="color: #20794D;
background-color: null;
font-style: inherit;">'multiview'</span>] </span>
<span id="cb28-12">metric <span class="op" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">=</span> <span class="st" style="color: #20794D;
background-color: null;
font-style: inherit;">'qt_auc_10_th_0.1'</span></span>
<span id="cb28-13"></span>
<span id="cb28-14"><span class="cf" style="color: #003B4F;
background-color: null;
font-weight: bold;
font-style: inherit;">for</span> dset <span class="kw" style="color: #003B4F;
background-color: null;
font-weight: bold;
font-style: inherit;">in</span> datasets:</span>
<span id="cb28-15">    mAA_stereo <span class="op" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">=</span> results[dset][<span class="st" style="color: #20794D;
background-color: null;
font-style: inherit;">'results'</span>][<span class="st" style="color: #20794D;
background-color: null;
font-style: inherit;">'allseq'</span>][<span class="st" style="color: #20794D;
background-color: null;
font-style: inherit;">'stereo'</span>][<span class="st" style="color: #20794D;
background-color: null;
font-style: inherit;">'run_avg'</span>][metric][<span class="st" style="color: #20794D;
background-color: null;
font-style: inherit;">'mean'</span>]</span>
<span id="cb28-16">    mAA_colmap <span class="op" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">=</span> results[dset][<span class="st" style="color: #20794D;
background-color: null;
font-style: inherit;">'results'</span>][<span class="st" style="color: #20794D;
background-color: null;
font-style: inherit;">'allseq'</span>][<span class="st" style="color: #20794D;
background-color: null;
font-style: inherit;">'multiview'</span>][<span class="st" style="color: #20794D;
background-color: null;
font-style: inherit;">'run_avg'</span>][<span class="st" style="color: #20794D;
background-color: null;
font-style: inherit;">'bag_avg'</span>][<span class="st" style="color: #20794D;
background-color: null;
font-style: inherit;">'qt_auc_colmap_10'</span>][<span class="st" style="color: #20794D;
background-color: null;
font-style: inherit;">'mean'</span>]</span>
<span id="cb28-17">    <span class="bu" style="color: null;
background-color: null;
font-style: inherit;">print</span> (<span class="ss" style="color: #20794D;
background-color: null;
font-style: inherit;">f'</span><span class="sc" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">{</span>submission_name<span class="sc" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">}</span><span class="ss" style="color: #20794D;
background-color: null;
font-style: inherit;"> stereo mAA for </span><span class="sc" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">{</span>dset<span class="sc" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">}</span><span class="ss" style="color: #20794D;
background-color: null;
font-style: inherit;"> is </span><span class="sc" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">{</span>mAA_stereo<span class="sc" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">:.4f}</span><span class="ss" style="color: #20794D;
background-color: null;
font-style: inherit;">'</span>)</span>
<span id="cb28-18">    <span class="bu" style="color: null;
background-color: null;
font-style: inherit;">print</span> (<span class="ss" style="color: #20794D;
background-color: null;
font-style: inherit;">f'</span><span class="sc" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">{</span>submission_name<span class="sc" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">}</span><span class="ss" style="color: #20794D;
background-color: null;
font-style: inherit;"> multiview mAA for </span><span class="sc" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">{</span>dset<span class="sc" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">}</span><span class="ss" style="color: #20794D;
background-color: null;
font-style: inherit;"> is </span><span class="sc" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">{</span>mAA_colmap<span class="sc" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">:.4f}</span><span class="ss" style="color: #20794D;
background-color: null;
font-style: inherit;">'</span>)</span>
<span id="cb28-19">    <span class="co" style="color: #5E5E5E;
background-color: null;
font-style: inherit;"># Remember, I did not run multiview evaluation for the original submission.</span></span></code></pre></div>
<div class="cell-output cell-output-stdout">
<pre><code>CV-DoG-AffNet-HardNet8 (kornia) stereo mAA for phototourism is 0.7108
CV-DoG-AffNet-HardNet8 (kornia) multiview mAA for phototourism is 0.0000
CV-DoG-AffNet-HardNet8 (kornia) stereo mAA for pragueparks is 0.5850
CV-DoG-AffNet-HardNet8 (kornia) multiview mAA for pragueparks is 0.0000
CV-DoG-AffNet-HardNet8 (kornia) stereo mAA for googleurban is 0.3099
CV-DoG-AffNet-HardNet8 (kornia) multiview mAA for googleurban is 0.0000</code></pre>
</div>
</div>
<div id="cell-36" class="cell">
<div class="sourceCode cell-code" id="cb30" style="background: #f1f3f5;"><pre class="sourceCode python code-with-copy"><code class="sourceCode python"><span id="cb30-1"><span class="im" style="color: #00769E;
background-color: null;
font-style: inherit;">import</span> os</span>
<span id="cb30-2"></span>
<span id="cb30-3">hashname<span class="op" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">=</span><span class="st" style="color: #20794D;
background-color: null;
font-style: inherit;">'dog-affnet-hardnet8-degensac-adalam'</span></span>
<span id="cb30-4">res_fname <span class="op" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">=</span> os.path.join(<span class="st" style="color: #20794D;
background-color: null;
font-style: inherit;">'../image-matching-benchmark/packed-val'</span>, <span class="ss" style="color: #20794D;
background-color: null;
font-style: inherit;">f'</span><span class="sc" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">{</span>hashname<span class="sc" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">}</span><span class="ss" style="color: #20794D;
background-color: null;
font-style: inherit;">.json'</span>)</span>
<span id="cb30-5"></span>
<span id="cb30-6"><span class="cf" style="color: #003B4F;
background-color: null;
font-weight: bold;
font-style: inherit;">with</span> <span class="bu" style="color: null;
background-color: null;
font-style: inherit;">open</span>(res_fname, <span class="st" style="color: #20794D;
background-color: null;
font-style: inherit;">'r'</span>) <span class="im" style="color: #00769E;
background-color: null;
font-style: inherit;">as</span> f:</span>
<span id="cb30-7">    results_adalam <span class="op" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">=</span> json.load(f)</span>
<span id="cb30-8"></span>
<span id="cb30-9">submission_name_adalam <span class="op" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">=</span> results_adalam[<span class="st" style="color: #20794D;
background-color: null;
font-style: inherit;">'config'</span>][<span class="st" style="color: #20794D;
background-color: null;
font-style: inherit;">'metadata'</span>][<span class="st" style="color: #20794D;
background-color: null;
font-style: inherit;">'method_name'</span>]</span>
<span id="cb30-10">datasets <span class="op" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">=</span> [<span class="st" style="color: #20794D;
background-color: null;
font-style: inherit;">'phototourism'</span>, <span class="st" style="color: #20794D;
background-color: null;
font-style: inherit;">'pragueparks'</span>, <span class="st" style="color: #20794D;
background-color: null;
font-style: inherit;">'googleurban'</span>]</span>
<span id="cb30-11">tasks <span class="op" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">=</span> [<span class="st" style="color: #20794D;
background-color: null;
font-style: inherit;">'stereo'</span>] <span class="co" style="color: #5E5E5E;
background-color: null;
font-style: inherit;"># ['stereo', 'multiview']  #Remember, that we skip colmap evaluations</span></span>
<span id="cb30-12">metric <span class="op" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">=</span> <span class="st" style="color: #20794D;
background-color: null;
font-style: inherit;">'qt_auc_10_th_0.1'</span></span>
<span id="cb30-13"><span class="cf" style="color: #003B4F;
background-color: null;
font-weight: bold;
font-style: inherit;">for</span> dset <span class="kw" style="color: #003B4F;
background-color: null;
font-weight: bold;
font-style: inherit;">in</span> datasets:</span>
<span id="cb30-14">    mAA_stereo <span class="op" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">=</span> results_adalam[dset][<span class="st" style="color: #20794D;
background-color: null;
font-style: inherit;">'results'</span>][<span class="st" style="color: #20794D;
background-color: null;
font-style: inherit;">'allseq'</span>][<span class="st" style="color: #20794D;
background-color: null;
font-style: inherit;">'stereo'</span>][<span class="st" style="color: #20794D;
background-color: null;
font-style: inherit;">'run_avg'</span>][metric][<span class="st" style="color: #20794D;
background-color: null;
font-style: inherit;">'mean'</span>]</span>
<span id="cb30-15">    mAA_colmap <span class="op" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">=</span> results_adalam[dset][<span class="st" style="color: #20794D;
background-color: null;
font-style: inherit;">'results'</span>][<span class="st" style="color: #20794D;
background-color: null;
font-style: inherit;">'allseq'</span>][<span class="st" style="color: #20794D;
background-color: null;
font-style: inherit;">'multiview'</span>][<span class="st" style="color: #20794D;
background-color: null;
font-style: inherit;">'run_avg'</span>][<span class="st" style="color: #20794D;
background-color: null;
font-style: inherit;">'bag_avg'</span>][<span class="st" style="color: #20794D;
background-color: null;
font-style: inherit;">'qt_auc_colmap_10'</span>][<span class="st" style="color: #20794D;
background-color: null;
font-style: inherit;">'mean'</span>]</span>
<span id="cb30-16">    <span class="bu" style="color: null;
background-color: null;
font-style: inherit;">print</span> (<span class="ss" style="color: #20794D;
background-color: null;
font-style: inherit;">f'</span><span class="sc" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">{</span>submission_name_adalam<span class="sc" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">}</span><span class="ss" style="color: #20794D;
background-color: null;
font-style: inherit;"> stereo mAA for </span><span class="sc" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">{</span>dset<span class="sc" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">}</span><span class="ss" style="color: #20794D;
background-color: null;
font-style: inherit;"> is </span><span class="sc" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">{</span>mAA_stereo<span class="sc" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">:.4f}</span><span class="ss" style="color: #20794D;
background-color: null;
font-style: inherit;">'</span>)</span>
<span id="cb30-17">    <span class="bu" style="color: null;
background-color: null;
font-style: inherit;">print</span> (<span class="ss" style="color: #20794D;
background-color: null;
font-style: inherit;">f'</span><span class="sc" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">{</span>submission_name_adalam<span class="sc" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">}</span><span class="ss" style="color: #20794D;
background-color: null;
font-style: inherit;"> multiview mAA for </span><span class="sc" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">{</span>dset<span class="sc" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">}</span><span class="ss" style="color: #20794D;
background-color: null;
font-style: inherit;"> is </span><span class="sc" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">{</span>mAA_colmap<span class="sc" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">:.4f}</span><span class="ss" style="color: #20794D;
background-color: null;
font-style: inherit;">'</span>)</span></code></pre></div>
<div class="cell-output cell-output-stdout">
<pre><code>CV-DoG-AffNet-HardNet8-AdaLAM-DEGENSAC stereo mAA for phototourism is 0.7326
CV-DoG-AffNet-HardNet8-AdaLAM-DEGENSAC multiview mAA for phototourism is 0.8566
CV-DoG-AffNet-HardNet8-AdaLAM-DEGENSAC stereo mAA for pragueparks is 0.6624
CV-DoG-AffNet-HardNet8-AdaLAM-DEGENSAC multiview mAA for pragueparks is 0.5584
CV-DoG-AffNet-HardNet8-AdaLAM-DEGENSAC stereo mAA for googleurban is 0.3310
CV-DoG-AffNet-HardNet8-AdaLAM-DEGENSAC multiview mAA for googleurban is 0.1597</code></pre>
</div>
</div>
<p>We can also see results sequence-by-sequence</p>
<div id="cell-38" class="cell">
<div class="sourceCode cell-code" id="cb32" style="background: #f1f3f5;"><pre class="sourceCode python code-with-copy"><code class="sourceCode python"><span id="cb32-1"><span class="im" style="color: #00769E;
background-color: null;
font-style: inherit;">import</span> seaborn <span class="im" style="color: #00769E;
background-color: null;
font-style: inherit;">as</span> sns</span>
<span id="cb32-2"><span class="im" style="color: #00769E;
background-color: null;
font-style: inherit;">import</span> matplotlib.pyplot <span class="im" style="color: #00769E;
background-color: null;
font-style: inherit;">as</span> plt</span>
<span id="cb32-3"><span class="im" style="color: #00769E;
background-color: null;
font-style: inherit;">import</span> numpy <span class="im" style="color: #00769E;
background-color: null;
font-style: inherit;">as</span> np</span>
<span id="cb32-4">sns.set_context(<span class="st" style="color: #20794D;
background-color: null;
font-style: inherit;">'paper'</span>, font_scale<span class="op" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">=</span><span class="fl" style="color: #AD0000;
background-color: null;
font-style: inherit;">1.7</span>)</span>
<span id="cb32-5">seqs <span class="op" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">=</span> []</span>
<span id="cb32-6">mAAs_adalam <span class="op" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">=</span> []</span>
<span id="cb32-7">mAAs <span class="op" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">=</span> []</span>
<span id="cb32-8"></span>
<span id="cb32-9"><span class="cf" style="color: #003B4F;
background-color: null;
font-weight: bold;
font-style: inherit;">for</span> dset <span class="kw" style="color: #003B4F;
background-color: null;
font-weight: bold;
font-style: inherit;">in</span> datasets:</span>
<span id="cb32-10">    <span class="cf" style="color: #003B4F;
background-color: null;
font-weight: bold;
font-style: inherit;">for</span> task <span class="kw" style="color: #003B4F;
background-color: null;
font-weight: bold;
font-style: inherit;">in</span> tasks:</span>
<span id="cb32-11">        <span class="cf" style="color: #003B4F;
background-color: null;
font-weight: bold;
font-style: inherit;">for</span> seq <span class="kw" style="color: #003B4F;
background-color: null;
font-weight: bold;
font-style: inherit;">in</span> results[dset][<span class="st" style="color: #20794D;
background-color: null;
font-style: inherit;">'results'</span>].keys():</span>
<span id="cb32-12">            <span class="cf" style="color: #003B4F;
background-color: null;
font-weight: bold;
font-style: inherit;">if</span> seq <span class="op" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">==</span> <span class="st" style="color: #20794D;
background-color: null;
font-style: inherit;">'allseq'</span>: </span>
<span id="cb32-13">                <span class="cf" style="color: #003B4F;
background-color: null;
font-weight: bold;
font-style: inherit;">continue</span></span>
<span id="cb32-14">            mAA  <span class="op" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">=</span> results[dset][<span class="st" style="color: #20794D;
background-color: null;
font-style: inherit;">'results'</span>][seq][task][<span class="st" style="color: #20794D;
background-color: null;
font-style: inherit;">'run_avg'</span>][metric][<span class="st" style="color: #20794D;
background-color: null;
font-style: inherit;">'mean'</span>]</span>
<span id="cb32-15">            mAA_adalam  <span class="op" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">=</span> results_adalam[dset][<span class="st" style="color: #20794D;
background-color: null;
font-style: inherit;">'results'</span>][seq][task][<span class="st" style="color: #20794D;
background-color: null;
font-style: inherit;">'run_avg'</span>][metric][<span class="st" style="color: #20794D;
background-color: null;
font-style: inherit;">'mean'</span>]</span>
<span id="cb32-16">            </span>
<span id="cb32-17">            mAAs.append(mAA)</span>
<span id="cb32-18">            mAAs_adalam.append(mAA_adalam)</span>
<span id="cb32-19">            </span>
<span id="cb32-20">            seqs.append(seq)</span>
<span id="cb32-21">fig, ax <span class="op" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">=</span> plt.subplots(figsize<span class="op" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">=</span>(<span class="dv" style="color: #AD0000;
background-color: null;
font-style: inherit;">15</span>,<span class="dv" style="color: #AD0000;
background-color: null;
font-style: inherit;">5</span>))</span>
<span id="cb32-22">xticks <span class="op" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">=</span> <span class="dv" style="color: #AD0000;
background-color: null;
font-style: inherit;">2</span><span class="op" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">*</span>np.arange(<span class="bu" style="color: null;
background-color: null;
font-style: inherit;">len</span>(seqs))</span>
<span id="cb32-23">ax.set_xticks(xticks)</span>
<span id="cb32-24">ax.bar(xticks, mAAs)</span>
<span id="cb32-25">ax.bar(xticks<span class="op" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">+</span><span class="fl" style="color: #AD0000;
background-color: null;
font-style: inherit;">0.5</span>, mAAs_adalam)</span>
<span id="cb32-26">ax.set_xticklabels(seqs)</span>
<span id="cb32-27">ax.legend([<span class="st" style="color: #20794D;
background-color: null;
font-style: inherit;">'SMNN'</span>, <span class="st" style="color: #20794D;
background-color: null;
font-style: inherit;">'AdaLAM'</span>])</span>
<span id="cb32-28">ax.set_ylabel(<span class="st" style="color: #20794D;
background-color: null;
font-style: inherit;">'mAA'</span>)</span>
<span id="cb32-29">ax.set_xlabel(<span class="st" style="color: #20794D;
background-color: null;
font-style: inherit;">'Sequence'</span>)</span></code></pre></div>
<div class="cell-output cell-output-display">
<pre><code>Text(0.5, 0, 'Sequence')</code></pre>
</div>
<div class="cell-output cell-output-display">
<div>
<figure class="figure">
<p><img src="https://ducha-aiki.github.io/wide-baseline-stereo-blog/posts/2021-05-27-submitting-to-IMC2021-with-custom-matcher_files/figure-html/cell-18-output-2.png" class="img-fluid figure-img"></p>
</figure>
</div>
</div>
</div>
<p>How do our feature correspondences look like?</p>
</section>
<section id="creating-final-submission" class="level2">
<h2 class="anchored" data-anchor-id="creating-final-submission">Creating final submission</h2>
<p>Its time to create our final submission!</p>
<div id="cell-43" class="cell">
<div class="sourceCode cell-code" id="cb34" style="background: #f1f3f5;"><pre class="sourceCode python code-with-copy"><code class="sourceCode python"><span id="cb34-1"></span>
<span id="cb34-2">configs <span class="op" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">=</span> []</span>
<span id="cb34-3">current_config <span class="op" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">=</span> deepcopy(base_config)</span>
<span id="cb34-4"></span>
<span id="cb34-5">    </span>
<span id="cb34-6">current_config[<span class="st" style="color: #20794D;
background-color: null;
font-style: inherit;">'metadata'</span>][<span class="st" style="color: #20794D;
background-color: null;
font-style: inherit;">'method_name'</span>] <span class="op" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">=</span> <span class="st" style="color: #20794D;
background-color: null;
font-style: inherit;">'KORNIA TUTORIAL AdaLAM CV-DoG-AffNet-HardNet8'</span></span>
<span id="cb34-7"></span>
<span id="cb34-8">label <span class="op" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">=</span> current_config[<span class="st" style="color: #20794D;
background-color: null;
font-style: inherit;">'config_common'</span>][<span class="st" style="color: #20794D;
background-color: null;
font-style: inherit;">'json_label'</span>] </span>
<span id="cb34-9">current_config[<span class="st" style="color: #20794D;
background-color: null;
font-style: inherit;">'config_common'</span>][<span class="st" style="color: #20794D;
background-color: null;
font-style: inherit;">'json_label'</span>]  <span class="op" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">=</span> <span class="ss" style="color: #20794D;
background-color: null;
font-style: inherit;">f'</span><span class="sc" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">{</span>label<span class="sc" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">}</span><span class="ss" style="color: #20794D;
background-color: null;
font-style: inherit;">'</span></span>
<span id="cb34-10">configs.append(current_config)</span>
<span id="cb34-11"></span>
<span id="cb34-12"><span class="bu" style="color: null;
background-color: null;
font-style: inherit;">print</span> (current_config)</span>
<span id="cb34-13"><span class="cf" style="color: #003B4F;
background-color: null;
font-weight: bold;
font-style: inherit;">with</span> <span class="bu" style="color: null;
background-color: null;
font-style: inherit;">open</span>(<span class="st" style="color: #20794D;
background-color: null;
font-style: inherit;">'final_submission_adalam.json'</span>, <span class="st" style="color: #20794D;
background-color: null;
font-style: inherit;">'w'</span>) <span class="im" style="color: #00769E;
background-color: null;
font-style: inherit;">as</span> f:</span>
<span id="cb34-14">    json.dump(configs, f, indent<span class="op" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">=</span><span class="dv" style="color: #AD0000;
background-color: null;
font-style: inherit;">2</span>)</span></code></pre></div>
<div class="cell-output cell-output-stdout">
<pre><code>{'metadata': {'publish_anonymously': False, 'authors': 'Dmytro Mishkin, Milan Pultar and kornia team', 'contact_email': 'ducha.aiki@gmail.com', 'method_name': 'KORNIA TUTORIAL AdaLAM CV-DoG-AffNet-HardNet8', 'method_description': 'OpeCV SIFT keypoints 8000 features, followed by the AffNet normalization \n    and HardNet8 descriptor as implemented in kornia.\n    Matched using AdaLAM with DEGENSAC', 'link_to_website': 'https://github.com/kornia/kornia', 'link_to_pdf': 'https://arxiv.org/abs/2007.09699'}, 'config_common': {'json_label': 'dog-affnet-hardnet8-degensac-adalam', 'keypoint': 'cv2dog', 'descriptor': 'affnethardnet8', 'num_keypoints': 8000}, 'config_phototourism_stereo': {'use_custom_matches': True, 'custom_matches_name': 'adalam', 'geom': {'method': 'cv2-8pt'}}, 'config_phototourism_multiview': {'use_custom_matches': True, 'custom_matches_name': 'adalam', 'colmap': {}}, 'config_pragueparks_stereo': {'use_custom_matches': True, 'custom_matches_name': 'adalam', 'geom': {'method': 'cv2-8pt'}}, 'config_pragueparks_multiview': {'use_custom_matches': True, 'custom_matches_name': 'adalam', 'colmap': {}}, 'config_googleurban_stereo': {'use_custom_matches': True, 'custom_matches_name': 'adalam', 'geom': {'method': 'cv2-8pt'}}, 'config_googleurban_multiview': {'use_custom_matches': True, 'custom_matches_name': 'adalam', 'colmap': {}}}</code></pre>
</div>
</div>
<pre><code>Submission Zip file should have folder structure as follow:
 ├── config.json
 ├── [Dataset 1]
 │   ├── [Sequence 1]
 │   │   ├── keypoints.h5
 │   │   ├── descriptors.h5
 │   │   ├── matches.h5
 │   ├── [Sequence 2]
 │   │   ├── ...
 ├── [Dataset 2]
 │   ├── ...</code></pre>
<p>So we have to just copy our features, add config and zip them.</p>
<div class="sourceCode" id="cb37" style="background: #f1f3f5;"><pre class="sourceCode bash code-with-copy"><code class="sourceCode bash"><span id="cb37-1"><span class="fu" style="color: #4758AB;
background-color: null;
font-style: inherit;">cp</span> final_submission_adalam.json extracted/cv2-dog-affnet-hardnet8/config.json</span>
<span id="cb37-2"><span class="bu" style="color: null;
background-color: null;
font-style: inherit;">cd</span> extracted/cv2-dog-affnet-hardnet8</span>
<span id="cb37-3"><span class="fu" style="color: #4758AB;
background-color: null;
font-style: inherit;">zip</span> <span class="at" style="color: #657422;
background-color: null;
font-style: inherit;">-r</span> submission.zip <span class="pp" style="color: #AD0000;
background-color: null;
font-style: inherit;">*</span></span></code></pre></div>
<p>Last step before the submission - check the submission for correctness with provided script</p>
<div class="sourceCode" id="cb38" style="background: #f1f3f5;"><pre class="sourceCode bash code-with-copy"><code class="sourceCode bash"><span id="cb38-1"><span class="bu" style="color: null;
background-color: null;
font-style: inherit;">cd</span> ../../../image-matching-benchmark</span>
<span id="cb38-2"><span class="ex" style="color: null;
background-color: null;
font-style: inherit;">python</span> <span class="at" style="color: #657422;
background-color: null;
font-style: inherit;">-utt</span> submission_validator.py <span class="at" style="color: #657422;
background-color: null;
font-style: inherit;">--submit_file_path</span> ../imc2021-sample-kornia-submission/extracted/cv2-dog-affnet-hardnet8/submission.zip <span class="at" style="color: #657422;
background-color: null;
font-style: inherit;">--benchmark_repo_path</span> . <span class="at" style="color: #657422;
background-color: null;
font-style: inherit;">--raw_data_path</span> ../imc-2021-data/ <span class="at" style="color: #657422;
background-color: null;
font-style: inherit;">--datasets</span> googleurban phototourism pragueparks</span></code></pre></div>
<p>If everything is correct, you will see:</p>
<pre><code>Validating method 1/1: "dog-affnet-hardnet8-degensac"
['googleurban', 'phototourism', 'pragueparks']
Running: googleurban, stereo track
Running: googleurban, multiview track
Running: phototourism, stereo track
Running: phototourism, multiview track
Running: pragueparks, stereo track
Running: pragueparks, multiview track
Validating key "config_googleurban_stereo"
Validating key "config_googleurban_multiview"
Validating key "config_phototourism_stereo"
Validating key "config_phototourism_multiview"
Validating key "config_pragueparks_stereo"
Validating key "config_pragueparks_multiview"</code></pre>
<p>And file <code>submission_log.txt</code> will appear near our .zip file.</p>
<div class="sourceCode" id="cb40" style="background: #f1f3f5;"><pre class="sourceCode bash code-with-copy"><code class="sourceCode bash"><span id="cb40-1"><span class="fu" style="color: #4758AB;
background-color: null;
font-style: inherit;">cat</span> ../imc2021-sample-kornia-submission/extracted/cv2-dog-affnet-hardnet8/submission_log.txt</span></code></pre></div>
<pre><code>Submission is in proper format, please submit to IMW 2021 website.</code></pre>
<p>That’s all, folks! We can submit! But, please, do not just submit this sample submission - make your own :)</p>


</section>
</section>

<p>Everything you (didn’t) want to know about image matching</p> ]]></description>
  <guid>https://ducha-aiki.github.io/wide-baseline-stereo-blog/posts/2021-05-27-submitting-to-imc2021-with-custom-matcher.html</guid>
  <pubDate>Wed, 26 May 2021 22:00:00 GMT</pubDate>
  <media:content url="https://ducha-aiki.github.io/wide-baseline-stereo-blog/posts/images/wooden_lady.png" medium="image" type="image/png" height="114" width="91"/>
</item>
<item>
  <title>Evaluating OpenCV new RANSACs</title>
  <dc:creator>Dmytro Mishkin</dc:creator>
  <link>https://ducha-aiki.github.io/wide-baseline-stereo-blog/posts/2021-05-17-opencv-new-ransacs.html</link>
  <description><![CDATA[ 





<section id="opencv-ransac-is-dead.-long-live-the-opencv-usac" class="level2">
<h2 class="anchored" data-anchor-id="opencv-ransac-is-dead.-long-live-the-opencv-usac">OpenCV RANSAC is dead. Long live the OpenCV USAC!</h2>
<p>Year ago we published a paper “<a href="https://arxiv.org/abs/2003.01587">Image Matching across Wide Baselines: From Paper to Practice</a>”, which, among other messages, has shown that OpenCV RANSAC for fundamental matrix estimation is terrible: it was super inaccurate and slow. Since then my colleague <a href="https://www.linkedin.com/in/maksym-ivashechkin-770186185">Maksym Ivashechkin</a> has spent a summer 2020 improving OpenCV RANSACs. His work was released as a part of <a href="https://docs.opencv.org/4.5.0/d1/df1/md__build_master-contrib_docs-lin64_opencv_doc_tutorials_calib3d_usac.html">OpenCV 4.5.0 release</a>.</p>
<p>Now it is time to benchmark them. Let’s go!</p>
</section>
<section id="evaluation-methodology" class="level2">
<h2 class="anchored" data-anchor-id="evaluation-methodology">Evaluation methodology</h2>
<p>The benchmark is done on the validation subset of the <a href="https://www.cs.ubc.ca/research/image-matching-challenge/2021/data/">Image Matching Challenge 2021</a> datasets. We have detected RootSIFT features, matched them with optimal mutual SNN ratio test and feed into the tested RANSACs. The resulting fundamental matrixes were transformed into relative poses and compared to the ground truth poses. You can check details in the paper “<a href="https://arxiv.org/abs/2003.01587">Image Matching across Wide Baselines: From Paper to Practice</a>”.</p>
<p>For all RANSACs we first determine the optimal inlier threshold by the grid search, whereas number of iterations (<code>max_iter</code>) was set to a reasonable 100k. Then, after fixing this optimal threshold, we vary number of iterations from 10 to 10M. This gives us an accuracy-time curve.</p>
</section>
<section id="methods-evaluated" class="level2">
<h2 class="anchored" data-anchor-id="methods-evaluated">Methods evaluated</h2>
<p>Non-OpenCV methods:</p>
<ul>
<li><a href="https://pypi.org/project/pydegensac/">DEGENSAC</a> - from <code>pydegensac</code> package, based on the original implementation of the method, proposed in CVPR 2005 paper “Two-View Geometry Estimation Unaffected by a Dominant Plane”. It is the default choise for the Image Matching Challenge 2020 and 2021.</li>
<li><a href="https://pypi.org/project/pydegensac/">PyRANSAC</a> - also from <code>pydegensac</code> package, with flag <code>enable_degeneracy_check=False</code>, which is equivalent to a vanilla LO-RANSAC implementation.</li>
</ul>
<p>OpenCV methods, named after the flag, one needs to pass into <code>cv2.findFundamentalMatrix</code> function:</p>
<ul>
<li>USAC_DEFAULT – <a href="https://cmp.felk.cvut.cz/~chum/papers/chum-DAGM03.pdf">LO-RANSAC</a> + degeneracy tests</li>
<li>USAC_FAST – <a href="https://cmp.felk.cvut.cz/~chum/papers/chum-DAGM03.pdf">LO-RANSAC</a> + degeneracy tests. Fewer iterations in local optimization step than USAC_DEFAULT. Uses RANSAC score to maximize number of inliers and terminate earlier.</li>
<li>USAC_ACCURATE. Implements <a href="https://cmp.felk.cvut.cz/~matas/papers/barath-2018-gc_ransac-cvpr.pdf">Graph-Cut RANSAC</a> + degeneracy tests.</li>
<li>USAC_MAGSAC – <a href="https://arxiv.org/abs/1912.05909">MAGSAC++</a> implementation + degeneracy tests.</li>
<li>RANSAC – OpenCV RANSAC implementation from the previous versions of the library, <em>no</em> degeneracy tests</li>
</ul>
<p>All OpenCV USAC methods also use SPRT-test for speeding-up the evaluation.</p>
</section>
<section id="results" class="level2">
<h2 class="anchored" data-anchor-id="results">Results</h2>
<p>Here are results for all 3 datasets. The lefter and upper is curve, the better. Dashed vertical line marks 1/25 sec (“realtime”) and 0.5 sec (challenge limit) time budget. Legend shows the method name and the optimal inlier threshold for the datasets: Phototourism, GoogleUrban and PragueParks respectively.</p>
<div class="quarto-figure quarto-figure-center">
<figure class="figure">
<p><img src="https://ducha-aiki.github.io/wide-baseline-stereo-blog/posts/2021-05-17-OpenCV-New-RANSACs_files/att_00004.png" class="img-fluid figure-img"></p>
<figcaption>Time-accuracy trade-off on all 3 datasets</figcaption>
</figure>
</div>
<ol type="1">
<li><p>The first and main message – <strong>all new flags are much better than the old OpenCV implementation (green curve, worst results), which still a default option</strong>.</p></li>
<li><p>10k iterations and USAC_ACCURATE (red curve) gives you great results within 0.01 sec</p></li>
<li><p>All OpenCV advanced USACs are better than for the small/medium time budget (&lt; 0.1 sec per image) than pydegensac (blue curve).<br>
</p></li>
<li><p>The best methods for the higher budget are OpenCV USAC_MAGSAC and DEGENSAC from the pydegensac package.</p></li>
<li><p>There is no point is using flag “USAC_FAST” it is always better to use USAC_DEFAULT, USAC_ACCURATE or USAC_MAGSAC.</p></li>
<li><p>USAC_MAGSAC is the only method, which optimal threshold is <em>the same across all datasets</em>. This is a valuable property for practice, as it requires the least tuning.</p></li>
</ol>
<p>If you are interesting in results for an individual datasets, here they are.</p>
<section id="phototourism" class="level3">
<h3 class="anchored" data-anchor-id="phototourism">Phototourism</h3>
<div class="quarto-figure quarto-figure-center">
<figure class="figure">
<p><img src="https://ducha-aiki.github.io/wide-baseline-stereo-blog/posts/2021-05-17-OpenCV-New-RANSACs_files/att_00005.png" title="Time-accuracy trade-off on the Phototourism dataset" class="img-fluid figure-img"></p>
<figcaption>image.png</figcaption>
</figure>
</div>
</section>
<section id="googleurban" class="level3">
<h3 class="anchored" data-anchor-id="googleurban">GoogleUrban</h3>
<div class="quarto-figure quarto-figure-center">
<figure class="figure">
<p><img src="https://ducha-aiki.github.io/wide-baseline-stereo-blog/posts/2021-05-17-OpenCV-New-RANSACs_files/att_00007.png" class="img-fluid figure-img"></p>
<figcaption>Time-accuracy trade-off on the GoogleUrban dataset</figcaption>
</figure>
</div>
</section>
<section id="pragueparks" class="level3">
<h3 class="anchored" data-anchor-id="pragueparks">PragueParks</h3>
<div class="quarto-figure quarto-figure-center">
<figure class="figure">
<p><img src="https://ducha-aiki.github.io/wide-baseline-stereo-blog/posts/2021-05-17-OpenCV-New-RANSACs_files/att_00008.png" class="img-fluid figure-img"></p>
<figcaption>Time-accuracy trade-off on the PragueParks dataset</figcaption>
</figure>
</div>
</section>
</section>
<section id="why-do-i-tune-and-evaluate-on-the-same-set" class="level2">
<h2 class="anchored" data-anchor-id="why-do-i-tune-and-evaluate-on-the-same-set">Why do I tune and evaluate on the same set?</h2>
<p>It is true, that tuning and evaluation of the method on the same dataset does not make any sense. However, let me defend my choice. Here are the arguments:</p>
<ol type="1">
<li><p>I do not want to compromise an integrity of the test set, which is the basis of the on-going competition <a href="https://www.cs.ubc.ca/research/image-matching-challenge/current/">Image Matching Challenge 2021</a> with prize money. That is why I do not want to leak information from the abovementioned test set and this is my primarly optimization objective. I also cannot tune the threshold on the “training subset”, as both GoogleUrban and PragueParks do not have such.</p></li>
<li><p>I am interested more in the rough speed-accuracy trade-off than the precise rankings of the methods. It is quite likely, that those methods, which have an small acuracy gap on the validation set, would switch on the test set – as it happened with DEGENSAC and MAGSAC in our original paper. However, it is very unlikely, that method, which performs poorly on the validation set would magically outperform everyone on the test set. Again, see PyRANSAC vs DEGENSAC in the original paper.</p></li>
<li><p>I clearly state this fact as a limitation and do not publish a paper ;)</p></li>
</ol>
</section>
<section id="conclusion" class="level2">
<h2 class="anchored" data-anchor-id="conclusion">Conclusion</h2>
<p>New OpenCV RANSACs are fast and have comparable accuracy, you can safely pick one of them. However, if you are using pydegensac and have &gt; 0.1 sec time budget, there is no need to switch.</p>
<p>Use proper RANSACs and be happy :)</p>


</section>

<p>Everything you (didn’t) want to know about image matching</p> ]]></description>
  <guid>https://ducha-aiki.github.io/wide-baseline-stereo-blog/posts/2021-05-17-opencv-new-ransacs.html</guid>
  <pubDate>Sun, 16 May 2021 22:00:00 GMT</pubDate>
  <media:content url="https://ducha-aiki.github.io/wide-baseline-stereo-blog/posts/images/ransac_small.png" medium="image" type="image/png" height="102" width="144"/>
</item>
<item>
  <title>Image Matching Challenge 2020 Recap</title>
  <dc:creator>Dmytro Mishkin</dc:creator>
  <link>https://ducha-aiki.github.io/wide-baseline-stereo-blog/posts/2021-05-14-imc2020-competition-recap.html</link>
  <description><![CDATA[ 





<section id="what-is-image-matching-challenge" class="level2">
<h2 class="anchored" data-anchor-id="what-is-image-matching-challenge">What is Image Matching Challenge?</h2>
<p>Image Matching Challenge is an on-going benchmark of the local features, matching methods and RANSACs, held since 2019. Its main idea is to measure a <em>downstream metric</em> of the image matching pipeline, such as <em>camera pose accuracy</em> after a <em>careful hyperparameter tuning</em>.</p>
<div class="quarto-figure quarto-figure-center">
<figure class="figure">
<p><img src="https://ducha-aiki.github.io/wide-baseline-stereo-blog/posts/2021-05-14-IMC2020-competition-recap_files/att_00000.png" class="img-fluid figure-img"></p>
<figcaption>image.png</figcaption>
</figure>
</div>
<p>Before the challenge we have benchmarked existing popular and recent methods, such as (Root)SIFT, ORB, HardNet, SuperPoint, R2D2, D2Net, etc, and published a paper called “<a href="https://arxiv.org/pdf/2003.01587">Image matching across wide baselines: From paper to practice</a>”.</p>
<p>The best results were obtained by a combinination of difference-of-Gaussians (DoG) local feature detector, commonly referred as <a href="https://docs.opencv.org/master/da/df5/tutorial_py_sift_intro.html">SIFT detector</a>, with deep learned patch descriptor such as <a href="https://kornia.readthedocs.io/en/latest/feature.html#kornia.feature.HardNet">HardNet</a>, <a href="https://kornia.readthedocs.io/en/latest/feature.html#kornia.feature.SOSNet">SOSNet</a>, or <a href="https://kornia.readthedocs.io/en/latest/feature.html#kornia.feature.LAFAffNetShapeEstimator">AffNet</a>-<a href="https://kornia.readthedocs.io/en/latest/feature.html#kornia.feature.HardNet">HardNet</a>.</p>
<p>Let’s check what was proposed by the challenge participants.</p>
</section>
<section id="top-solutions-2020-and-follow-ups" class="level2">
<h2 class="anchored" data-anchor-id="top-solutions-2020-and-follow-ups">Top solutions-2020 and follow-ups</h2>
<section id="superglue" class="level3">
<h3 class="anchored" data-anchor-id="superglue">SuperGlue</h3>
<div class="quarto-figure quarto-figure-center">
<figure class="figure">
<p><img src="https://ducha-aiki.github.io/wide-baseline-stereo-blog/posts/2021-05-14-IMC2020-competition-recap_files/att_00001.png" class="img-fluid figure-img"></p>
<figcaption>SuperGlue pipeline</figcaption>
</figure>
</div>
<p><a href="https://psarlin.com/superglue/">SuperGlue</a> is an attention-based graph neural network for matching local features, taking into account both geometry (keypoint location) and appearance (descriptor). Unlike previous works, e.g.&nbsp;<a href="https://arxiv.org/abs/1711.05971">CNe</a>, or OANet below, it does not “scoring-and-cleaning” already established tentative correspondences. Instead, it establishes the correpondences given the local features from two images. SuperGlue won IMC-2020, as well as two other competitions at CVPR 2020. Its inference implementation is available <a href="https://github.com/magicleap/SuperGluePretrainedNetwork">here</a>.</p>
</section>
<section id="oanet" class="level3">
<h3 class="anchored" data-anchor-id="oanet">OANet</h3>
<p><img src="https://ducha-aiki.github.io/wide-baseline-stereo-blog/posts/2021-05-14-IMC2020-competition-recap_files/att_00002.png" class="img-fluid" alt="OANet architecture"> <a href="https://arxiv.org/pdf/1908.04964.pdf">OANet</a> is an specialized neural network architecture for “scoring-and-cleaning” already established tentative correspondences. OANet was run on top of DoG-HardNet local features. Pytorch version of OANet is available <a href="https://github.com/zjhthu/OANet">here</a></p>
</section>
<section id="adalam" class="level3">
<h3 class="anchored" data-anchor-id="adalam">AdaLAM</h3>
<div class="quarto-figure quarto-figure-center">
<figure class="figure">
<p><img src="https://ducha-aiki.github.io/wide-baseline-stereo-blog/posts/2021-05-14-IMC2020-competition-recap_files/att_00003.png" class="img-fluid figure-img"></p>
<figcaption>AdaLAM stages, from original paper</figcaption>
</figure>
</div>
<p><a href="https://arxiv.org/abs/2006.04250">AdaLAM</a> is a <em>handcrafted</em> algorithm for tentative correspondence cleaning, which work comparably or even better than a learning-based approaches. It is based on two core assumptions:</p>
<ul>
<li>Keypoints, which are near from each other are probably corresponding to the neaighboring keypoints in the other images</li>
<li>If keypoints are in correspondence, it means that their orientation and scale are also in correspondence. Check my post <a href="https://ducha-aiki.github.io/wide-baseline-stereo-blog/posts/2020-07-17-affine-correspondences.html">Local affine features: useful side product</a> for an explanation</li>
</ul>
<p>The implementation is avilable <a href="https://github.com/cavalli1234/AdaLAM">here</a>. As DoG+HardNet were used as a local features.</p>
</section>
<section id="disk" class="level3">
<h3 class="anchored" data-anchor-id="disk">DISK</h3>
<div class="quarto-figure quarto-figure-center">
<figure class="figure">
<p><img src="https://ducha-aiki.github.io/wide-baseline-stereo-blog/posts/2021-05-14-IMC2020-competition-recap_files/att_00004.png" class="img-fluid figure-img"></p>
<figcaption>DISK training objectives</figcaption>
</figure>
</div>
<p><a href="https://arxiv.org/pdf/2006.13566.pdf">DISK</a> is local feature, which has two main differences from the rest of competitors (SuperPoint, R2D2):</p>
<ol type="1">
<li>DISK is trained with a reinforcement-learning objective</li>
<li>DISK has UNet-like architecture, unlike VGG-style for the rest of the features.</li>
</ol>
<p>Its implementation is available <a href="https://github.com/cvlab-epfl/disk">here</a>.</p>
</section>
<section id="hynet" class="level3">
<h3 class="anchored" data-anchor-id="hynet">HyNet</h3>
<p><img src="https://ducha-aiki.github.io/wide-baseline-stereo-blog/posts/2021-05-14-IMC2020-competition-recap_files/att_00005.png" class="img-fluid" alt="HyNet architecture"> <a href="https://arxiv.org/pdf/2006.10202.pdf">HyNet</a> is an next stage in the L2Net-HardNet-SOSNet series of local patch descriptors. It is different from the previous works in two ways:</p>
<ol type="1">
<li>BatchNorm and ReLU in the HardNet architecture are replaced with FRN and TLU respecively.</li>
<li>During training, distance to the negative (non-matching) <img src="https://latex.codecogs.com/png.latex?d(%5Ctheta)"> and positive (matching) <img src="https://latex.codecogs.com/png.latex?s(%5Ctheta)"> samples are calculated in a different way for better learning, see image below. <em>u</em> and <em>v</em> denote descriptors is</li>
</ol>
<div class="quarto-figure quarto-figure-center">
<figure class="figure">
<p><img src="https://ducha-aiki.github.io/wide-baseline-stereo-blog/posts/2021-05-14-IMC2020-competition-recap_files/att_00006.png" class="img-fluid figure-img"></p>
<figcaption>Positive and negative distances used fro HyNet training</figcaption>
</figure>
</div>
<p>HyNet submission also used semantic segmentation network to remove the keypoints from the non-matchable areas, such as sky and water.</p>
</section>
<section id="hardnet8" class="level3">
<h3 class="anchored" data-anchor-id="hardnet8">HardNet8</h3>
<div class="quarto-figure quarto-figure-center">
<figure class="figure">
<p><img src="https://ducha-aiki.github.io/wide-baseline-stereo-blog/posts/2021-05-14-IMC2020-competition-recap_files/att_00008.png" class="img-fluid figure-img"></p>
<figcaption>HardNet8 architecture</figcaption>
</figure>
</div>
<p><a href="https://arxiv.org/abs/2007.09699">HardNet8</a> is another improvement of the HardNet architecture:</p>
<ol type="1">
<li>Deeper and wider network</li>
<li>The output is compressed with a PCA.</li>
<li>The training set and hyperparameters are carefully selected.</li>
</ol>
<p>It is available in <a href="https://kornia.readthedocs.io/en/latest/feature.html#kornia.feature.HardNet8">kornia</a></p>
</section>
</section>
<section id="challenge" class="level2">
<h2 class="anchored" data-anchor-id="challenge">2021 challenge</h2>
<p>This year challenge brings 2 new datasets: PragueParks and GoogleUrban.</p>
<section id="the-pragueparks-dataset" class="level3">
<h3 class="anchored" data-anchor-id="the-pragueparks-dataset">The PragueParks dataset</h3>
<div class="quarto-figure quarto-figure-center">
<figure class="figure">
<p><img src="https://ducha-aiki.github.io/wide-baseline-stereo-blog/posts/2021-05-14-IMC2020-competition-recap_files/att_00009.png" class="img-fluid figure-img"></p>
<figcaption>PragueParks dataset samples</figcaption>
</figure>
</div>
<p>The PragueParks dataset contains images from video sequences captured by the organizers with an iPhone 11, in 2021. The iPhone 11 has two cameras, with normal and wide lenses, both of which were used. Note that while the video is high quality, some of the frames suffer from motion blur. These videos were then processed by the commercial 3D reconstruction software <a href="https://www.capturingreality.com/">RealityCapture</a>, which is orders of magnitude faster than COLMAP, while delivering a comparable output in terms of accuracy. Similarly to we did for the “PhotoTourism” dataset, this data is then subsampled in order to generate the subsets used for evaluation.</p>
<p>The dataset contains small-scale scenes like tree, pond, wooden and metal sculptures with different level of zoom, lots of vegetation, and no people. The distribution of its camera poses differs from Phototourism.</p>
</section>
<section id="the-googleurban-dataset" class="level3">
<h3 class="anchored" data-anchor-id="the-googleurban-dataset">The GoogleUrban dataset</h3>
<div class="quarto-figure quarto-figure-center">
<figure class="figure">
<p><img src="https://ducha-aiki.github.io/wide-baseline-stereo-blog/posts/2021-05-14-IMC2020-competition-recap_files/att_00010.png" class="img-fluid figure-img"></p>
<figcaption>GoogleUrban dataset samples</figcaption>
</figure>
</div>
<p>The GoogleUrban dataset contains images used by Google to evaluate localization algorithms, such as those in Google’s Visual Positioning System, which powers Live View on millions on mobile devices. They are obtained from videos collected from different cell phones, on many countries all over the world, often years apart. They contain poses, but not depth maps. Please note that due to legal reasons, this data is released with a restricted license, and must be deleted by the end of the challenge.</p>
</section>
</section>
<section id="submit-your-solution" class="level2">
<h2 class="anchored" data-anchor-id="submit-your-solution">Submit your solution!</h2>
<p>You can check out the tutorial on how to submit to IMC 2021 in the post <a href="https://ducha-aiki.github.io/wide-baseline-stereo-blog/2021/05/12/submitting-to-IMC2021-step-by-step.html">Submitting to Image Matching Challenge 2021</a>. Good luck!</p>


</section>

<p>Everything you (didn’t) want to know about image matching</p> ]]></description>
  <guid>https://ducha-aiki.github.io/wide-baseline-stereo-blog/posts/2021-05-14-imc2020-competition-recap.html</guid>
  <pubDate>Thu, 13 May 2021 22:00:00 GMT</pubDate>
  <media:content url="https://ducha-aiki.github.io/wide-baseline-stereo-blog/posts/images/trevi_small.png" medium="image" type="image/png" height="115" width="144"/>
</item>
<item>
  <title>Submitting to Image Matching Challenge 2021</title>
  <dc:creator>Dmytro Mishkin</dc:creator>
  <link>https://ducha-aiki.github.io/wide-baseline-stereo-blog/posts/2021-05-12-submitting-to-imc2021-step-by-step.html</link>
  <description><![CDATA[ 





<section id="what-is-image-matching-challenge" class="level2">
<h2 class="anchored" data-anchor-id="what-is-image-matching-challenge">What is Image Matching Challenge?</h2>
<p>IMC is a benchmark and <a href="https://www.cs.ubc.ca/research/image-matching-challenge/current/">challenge</a> for the local features (such as SIFT, SuperPoint, etc), matching methods (CNe, SuperGlue, etc.) and robust geometry estimators such as RANSAC, at <a href="https://image-matching-workshop.github.io/">CVPR 2021 Workshop on Image Matching</a>.</p>
<p>I will walk you through the submission process, including writing setting-up an environment, writing a processing script and tuning matching and RANSAC for the best performance. As this is a tutorial, not a research paper, we use local feature descriptor available in <a href="https://github.com/kornia/kornia">kornia</a> for this sample submission.</p>
<p>All the codes and scripts in this tutorial are also avilable at https://github.com/ducha-aiki/imc2021-sample-kornia-submission. Let’s go!</p>
</section>
<section id="setting-up-the-environment" class="level2">
<h2 class="anchored" data-anchor-id="setting-up-the-environment">Setting up the environment</h2>
<p>First, let’s clone the benchmark repository.</p>
<div class="sourceCode" id="cb1" style="background: #f1f3f5;"><pre class="sourceCode bash code-with-copy"><code class="sourceCode bash"><span id="cb1-1"><span class="fu" style="color: #4758AB;
background-color: null;
font-style: inherit;">git</span> clone https://github.com/ubc-vision/image-matching-benchmark</span>
<span id="cb1-2"><span class="bu" style="color: null;
background-color: null;
font-style: inherit;">cd</span> image-matching-benchmark</span>
<span id="cb1-3"><span class="fu" style="color: #4758AB;
background-color: null;
font-style: inherit;">git</span> submodule update <span class="at" style="color: #657422;
background-color: null;
font-style: inherit;">--init</span> <span class="at" style="color: #657422;
background-color: null;
font-style: inherit;">--recursive</span></span></code></pre></div>
<p>Now we need to create conda virtual environment. I assume that you have conda installed, and if not - please, follow instructions <a href="https://conda.io/projects/conda/en/latest/user-guide/install/index.html">here</a></p>
<div class="sourceCode" id="cb2" style="background: #f1f3f5;"><pre class="sourceCode bash code-with-copy"><code class="sourceCode bash"><span id="cb2-1"><span class="ex" style="color: null;
background-color: null;
font-style: inherit;">conda</span> env create <span class="at" style="color: #657422;
background-color: null;
font-style: inherit;">-f</span> system/conda_env_dm.yml</span></code></pre></div>
<p>After successfull environment creation, let’s activate it</p>
<div class="sourceCode" id="cb3" style="background: #f1f3f5;"><pre class="sourceCode bash code-with-copy"><code class="sourceCode bash"><span id="cb3-1"><span class="ex" style="color: null;
background-color: null;
font-style: inherit;">conda</span> activate sfm</span></code></pre></div>
</section>
<section id="downloading-the-data" class="level2">
<h2 class="anchored" data-anchor-id="downloading-the-data">Downloading the data</h2>
<p>IMC-2021 data consists of <a href="https://www.cs.ubc.ca/research/image-matching-challenge/2021/data/">3 datasets</a>: Phototourism, PragueParks and GoogleUrban. You can download first two freely, but need to request a credentials and accept license agreement for the GoogleUrban dataset.</p>
<p>Latter is done by writing an email to image-matching@googlegroups.com and asking for the password. All datasets consist of two parts each: test part, which contains only images and validation part, which comes together with ground truth. Validation ground truth can (and should!) be used for hyperparameter tuning for your submission.</p>
<div class="sourceCode" id="cb4" style="background: #f1f3f5;"><pre class="sourceCode bash code-with-copy"><code class="sourceCode bash"><span id="cb4-1"><span class="bu" style="color: null;
background-color: null;
font-style: inherit;">cd</span> ..</span>
<span id="cb4-2"><span class="fu" style="color: #4758AB;
background-color: null;
font-style: inherit;">mkdir</span> imc-2021-data</span>
<span id="cb4-3"><span class="bu" style="color: null;
background-color: null;
font-style: inherit;">cd</span> imc-2021-data</span>
<span id="cb4-4"><span class="fu" style="color: #4758AB;
background-color: null;
font-style: inherit;">wget</span> https://www.cs.ubc.ca/research/kmyi_data/imc2021-public/imc-2021-test-public-pragueparks.tar.gz</span>
<span id="cb4-5"><span class="fu" style="color: #4758AB;
background-color: null;
font-style: inherit;">wget</span> https://www.cs.ubc.ca/research/kmyi_data/imc2021-public/imc-2021-validation-pragueparks.tar.gz</span>
<span id="cb4-6"><span class="fu" style="color: #4758AB;
background-color: null;
font-style: inherit;">wget</span> https://www.cs.ubc.ca/research/kmyi_data/imc2021-public/imc-2021-test-public-phototourism.tar.gz</span>
<span id="cb4-7"><span class="fu" style="color: #4758AB;
background-color: null;
font-style: inherit;">wget</span> https://www.cs.ubc.ca/research/kmyi_data/imc2021-public/imc-2021-validation-phototourism.tar.gz</span></code></pre></div>
<p>I am assuming that you have requsted an access to the GoogleUrban dataset, downloaded it and put in the same directory, as the rest of the data:</p>
<pre><code>(sfm) mishkdmy@n33:~/dev/imc-2021-data$ ls
imc-2021-test-public-googleurban.tar.gz   imc-2021-validation-googleurban.tar.gz
imc-2021-test-public-phototourism.tar.gz  imc-2021-validation-phototourism.tar.gz
imc-2021-test-public-pragueparks.tar.gz   imc-2021-validation-pragueparks.tar.gz</code></pre>
<p>Now let’s unpack it.</p>
<div class="sourceCode" id="cb6" style="background: #f1f3f5;"><pre class="sourceCode bash code-with-copy"><code class="sourceCode bash"><span id="cb6-1"><span class="cf" style="color: #003B4F;
background-color: null;
font-weight: bold;
font-style: inherit;">for</span> f <span class="kw" style="color: #003B4F;
background-color: null;
font-weight: bold;
font-style: inherit;">in</span> <span class="pp" style="color: #AD0000;
background-color: null;
font-style: inherit;">*</span>.tar.gz <span class="kw" style="color: #003B4F;
background-color: null;
font-weight: bold;
font-style: inherit;">;</span> <span class="cf" style="color: #003B4F;
background-color: null;
font-weight: bold;
font-style: inherit;">do</span> <span class="fu" style="color: #4758AB;
background-color: null;
font-style: inherit;">tar</span> <span class="at" style="color: #657422;
background-color: null;
font-style: inherit;">-xzf</span> <span class="va" style="color: #111111;
background-color: null;
font-style: inherit;">$f</span><span class="kw" style="color: #003B4F;
background-color: null;
font-weight: bold;
font-style: inherit;">;</span> <span class="cf" style="color: #003B4F;
background-color: null;
font-weight: bold;
font-style: inherit;">done</span></span></code></pre></div>
<p>Now the directory should look like this:</p>
<pre><code>(sfm) mishkdmy@n33:~/dev/imc-2021-data$ ls
googleurban                               imc-2021-validation-phototourism.tar.gz
imc-2021-test-public-googleurban.tar.gz   imc-2021-validation-pragueparks.tar.gz
imc-2021-test-public-phototourism.tar.gz  phototourism
imc-2021-test-public-pragueparks.tar.gz   pragueparks
imc-2021-validation-googleurban.tar.gz</code></pre>
</section>
<section id="extracting-the-features" class="level2">
<h2 class="anchored" data-anchor-id="extracting-the-features">Extracting the features</h2>
<p>Let’s start with creating the directory for our scripts (or you can clone it from <a href="https://github.com/ducha-aiki/imc2021-sample-kornia-submission">here</a> )</p>
<div class="sourceCode" id="cb8" style="background: #f1f3f5;"><pre class="sourceCode bash code-with-copy"><code class="sourceCode bash"><span id="cb8-1"><span class="bu" style="color: null;
background-color: null;
font-style: inherit;">cd</span> ..</span>
<span id="cb8-2"><span class="fu" style="color: #4758AB;
background-color: null;
font-style: inherit;">mkdir</span> imc2021-sample-kornia-submission</span>
<span id="cb8-3"><span class="bu" style="color: null;
background-color: null;
font-style: inherit;">cd</span> imc2021-sample-kornia-submission</span></code></pre></div>
<p>Now we will create a script, which extracts <a href="https://arxiv.org/abs/1711.06704">AffNet</a>-<a href="https://arxiv.org/abs/2007.09699">HardNet8</a> descriptors on top of OpenCV <a href="https://www.cs.ubc.ca/~lowe/papers/ijcv04.pdf">SIFT</a> keypoints.</p>
<p>We need to install <a href="https://github.com/pytorch/pytorch">pytorch</a> and <a href="https://github.com/kornia/kornia">kornia</a> for this:</p>
<div class="sourceCode" id="cb9" style="background: #f1f3f5;"><pre class="sourceCode bash code-with-copy"><code class="sourceCode bash"><span id="cb9-1"><span class="ex" style="color: null;
background-color: null;
font-style: inherit;">pip</span> install torch torchvision kornia</span>
<span id="cb9-2"><span class="ex" style="color: null;
background-color: null;
font-style: inherit;">pip</span> install kornia_moons <span class="at" style="color: #657422;
background-color: null;
font-style: inherit;">--no-deps</span></span></code></pre></div>
<p>Great! Now we are ready to extract the features from the images. Required imports and initializations:</p>
<div id="cell-6" class="cell" data-execution_count="1">
<div class="sourceCode cell-code" id="cb10" style="background: #f1f3f5;"><pre class="sourceCode python code-with-copy"><code class="sourceCode python"><span id="cb10-1"><span class="im" style="color: #00769E;
background-color: null;
font-style: inherit;">import</span> matplotlib.pyplot <span class="im" style="color: #00769E;
background-color: null;
font-style: inherit;">as</span> plt</span>
<span id="cb10-2"><span class="im" style="color: #00769E;
background-color: null;
font-style: inherit;">import</span> numpy <span class="im" style="color: #00769E;
background-color: null;
font-style: inherit;">as</span> np</span>
<span id="cb10-3"><span class="im" style="color: #00769E;
background-color: null;
font-style: inherit;">import</span> cv2</span>
<span id="cb10-4"><span class="im" style="color: #00769E;
background-color: null;
font-style: inherit;">import</span> torch</span>
<span id="cb10-5"><span class="im" style="color: #00769E;
background-color: null;
font-style: inherit;">import</span> kornia <span class="im" style="color: #00769E;
background-color: null;
font-style: inherit;">as</span> K</span>
<span id="cb10-6"><span class="im" style="color: #00769E;
background-color: null;
font-style: inherit;">import</span> kornia.feature <span class="im" style="color: #00769E;
background-color: null;
font-style: inherit;">as</span> KF</span>
<span id="cb10-7"><span class="im" style="color: #00769E;
background-color: null;
font-style: inherit;">from</span> kornia_moons.feature <span class="im" style="color: #00769E;
background-color: null;
font-style: inherit;">import</span> <span class="op" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">*</span></span>
<span id="cb10-8"></span>
<span id="cb10-9">device <span class="op" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">=</span> torch.device(<span class="st" style="color: #20794D;
background-color: null;
font-style: inherit;">'cpu'</span>)</span>
<span id="cb10-10"><span class="cf" style="color: #003B4F;
background-color: null;
font-weight: bold;
font-style: inherit;">try</span>:</span>
<span id="cb10-11">    <span class="cf" style="color: #003B4F;
background-color: null;
font-weight: bold;
font-style: inherit;">if</span> torch.cuda.is_available():</span>
<span id="cb10-12">        device <span class="op" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">=</span> torch.device(<span class="st" style="color: #20794D;
background-color: null;
font-style: inherit;">'cuda'</span>)</span>
<span id="cb10-13">        <span class="bu" style="color: null;
background-color: null;
font-style: inherit;">print</span> (<span class="st" style="color: #20794D;
background-color: null;
font-style: inherit;">"GPU mode"</span>)</span>
<span id="cb10-14"><span class="cf" style="color: #003B4F;
background-color: null;
font-weight: bold;
font-style: inherit;">except</span>:</span>
<span id="cb10-15">    <span class="bu" style="color: null;
background-color: null;
font-style: inherit;">print</span> (<span class="st" style="color: #20794D;
background-color: null;
font-style: inherit;">'CPU mode'</span>)</span>
<span id="cb10-16"><span class="co" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">#device = torch.device('cpu')</span></span>
<span id="cb10-17"><span class="co" style="color: #5E5E5E;
background-color: null;
font-style: inherit;"># SIFT (DoG) Detector</span></span>
<span id="cb10-18">sift_det <span class="op" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">=</span>  cv2.SIFT_create(<span class="dv" style="color: #AD0000;
background-color: null;
font-style: inherit;">8000</span>, contrastThreshold<span class="op" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">=-</span><span class="dv" style="color: #AD0000;
background-color: null;
font-style: inherit;">10000</span>, edgeThreshold<span class="op" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">=-</span><span class="dv" style="color: #AD0000;
background-color: null;
font-style: inherit;">10000</span>)</span>
<span id="cb10-19"></span>
<span id="cb10-20"><span class="co" style="color: #5E5E5E;
background-color: null;
font-style: inherit;"># HardNet8 descriptor</span></span>
<span id="cb10-21"></span>
<span id="cb10-22">hardnet8 <span class="op" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">=</span> KF.HardNet8(<span class="va" style="color: #111111;
background-color: null;
font-style: inherit;">True</span>).<span class="bu" style="color: null;
background-color: null;
font-style: inherit;">eval</span>().to(device)</span>
<span id="cb10-23"></span>
<span id="cb10-24"><span class="co" style="color: #5E5E5E;
background-color: null;
font-style: inherit;"># Affine shape estimator</span></span>
<span id="cb10-25">affnet <span class="op" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">=</span> KF.LAFAffNetShapeEstimator(<span class="va" style="color: #111111;
background-color: null;
font-style: inherit;">True</span>).<span class="bu" style="color: null;
background-color: null;
font-style: inherit;">eval</span>().to(device)</span></code></pre></div>
</div>
<p>Now we can define extract_features function. Feel free to modify it for your own features.</p>
<div id="cell-9" class="cell" data-execution_count="2">
<div class="sourceCode cell-code" id="cb11" style="background: #f1f3f5;"><pre class="sourceCode python code-with-copy"><code class="sourceCode python"><span id="cb11-1"><span class="kw" style="color: #003B4F;
background-color: null;
font-weight: bold;
font-style: inherit;">def</span> extract_features(img_fname, detector, affine, descriptor, device, visualize<span class="op" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">=</span><span class="va" style="color: #111111;
background-color: null;
font-style: inherit;">False</span>):</span>
<span id="cb11-2">    img <span class="op" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">=</span> cv2.cvtColor(cv2.imread(img_fname), cv2.COLOR_BGR2RGB)</span>
<span id="cb11-3">    <span class="cf" style="color: #003B4F;
background-color: null;
font-weight: bold;
font-style: inherit;">if</span> visualize:</span>
<span id="cb11-4">        plt.imshow(img)</span>
<span id="cb11-5">    kpts <span class="op" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">=</span> detector.detect(img, <span class="va" style="color: #111111;
background-color: null;
font-style: inherit;">None</span>)[:<span class="dv" style="color: #AD0000;
background-color: null;
font-style: inherit;">8000</span>]</span>
<span id="cb11-6">    <span class="co" style="color: #5E5E5E;
background-color: null;
font-style: inherit;"># We will not train anything, so let's save time and memory by no_grad()</span></span>
<span id="cb11-7">    <span class="cf" style="color: #003B4F;
background-color: null;
font-weight: bold;
font-style: inherit;">with</span> torch.no_grad():</span>
<span id="cb11-8">        timg <span class="op" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">=</span> K.image_to_tensor(img, <span class="va" style="color: #111111;
background-color: null;
font-style: inherit;">False</span>).<span class="bu" style="color: null;
background-color: null;
font-style: inherit;">float</span>()<span class="op" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">/</span><span class="fl" style="color: #AD0000;
background-color: null;
font-style: inherit;">255.</span></span>
<span id="cb11-9">        timg <span class="op" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">=</span> timg.to(device)</span>
<span id="cb11-10">        timg_gray <span class="op" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">=</span> K.rgb_to_grayscale(timg)</span>
<span id="cb11-11">        <span class="co" style="color: #5E5E5E;
background-color: null;
font-style: inherit;"># kornia expects keypoints in the local affine frame format. </span></span>
<span id="cb11-12">        <span class="co" style="color: #5E5E5E;
background-color: null;
font-style: inherit;"># Luckily, kornia_moons has a conversion function</span></span>
<span id="cb11-13">        lafs <span class="op" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">=</span> laf_from_opencv_SIFT_kpts(kpts, device<span class="op" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">=</span>device)</span>
<span id="cb11-14">        lafs_new <span class="op" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">=</span> affine(lafs, timg_gray)</span>
<span id="cb11-15">        <span class="cf" style="color: #003B4F;
background-color: null;
font-weight: bold;
font-style: inherit;">if</span> visualize:</span>
<span id="cb11-16">            visualize_LAF(timg, lafs_new, <span class="dv" style="color: #AD0000;
background-color: null;
font-style: inherit;">0</span>)</span>
<span id="cb11-17">        patches <span class="op" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">=</span> KF.extract_patches_from_pyramid(timg_gray, lafs_new, <span class="dv" style="color: #AD0000;
background-color: null;
font-style: inherit;">32</span>)</span>
<span id="cb11-18">        B, N, CH, H, W <span class="op" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">=</span> patches.size()</span>
<span id="cb11-19">        <span class="co" style="color: #5E5E5E;
background-color: null;
font-style: inherit;"># Descriptor accepts standard tensor [B, CH, H, W], while patches are [B, N, CH, H, W] shape</span></span>
<span id="cb11-20">        <span class="co" style="color: #5E5E5E;
background-color: null;
font-style: inherit;"># So we need to reshape a bit :) </span></span>
<span id="cb11-21">        descs <span class="op" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">=</span> descriptor(patches.view(B <span class="op" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">*</span> N, CH, H, W)).view(B <span class="op" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">*</span> N, <span class="op" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">-</span><span class="dv" style="color: #AD0000;
background-color: null;
font-style: inherit;">1</span>).detach().cpu().numpy()    </span>
<span id="cb11-22">    <span class="cf" style="color: #003B4F;
background-color: null;
font-weight: bold;
font-style: inherit;">return</span> kpts, descs</span></code></pre></div>
</div>
<p>Let’s check how it works on a single image.</p>
<div id="cell-11" class="cell" data-execution_count="3">
<div class="sourceCode cell-code" id="cb12" style="background: #f1f3f5;"><pre class="sourceCode python code-with-copy"><code class="sourceCode python"><span id="cb12-1">img_fname <span class="op" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">=</span> <span class="st" style="color: #20794D;
background-color: null;
font-style: inherit;">'../imc-2021-data/pragueparks/wooden_lady/set_100/images/IMG_9603.MOV_frame000001.jpg'</span></span>
<span id="cb12-2">kpts, descs <span class="op" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">=</span> extract_features(img_fname,  sift_det, affnet, hardnet8, device, <span class="va" style="color: #111111;
background-color: null;
font-style: inherit;">True</span>)</span></code></pre></div>
<div class="cell-output cell-output-display">
<div>
<figure class="figure">
<p><img src="https://ducha-aiki.github.io/wide-baseline-stereo-blog/posts/2021-05-12-submitting-to-IMC2021-step-by-step_files/figure-html/cell-5-output-1.png" class="img-fluid figure-img"></p>
</figure>
</div>
</div>
<div class="cell-output cell-output-display">
<div>
<figure class="figure">
<p><img src="https://ducha-aiki.github.io/wide-baseline-stereo-blog/posts/2021-05-12-submitting-to-IMC2021-step-by-step_files/figure-html/cell-5-output-2.png" class="img-fluid figure-img"></p>
</figure>
</div>
</div>
</div>
<p>So far, so good. Now we need to convert our keypoints from OpenCV format to the benchmark format, which is <code>numpy.array</code> <code>[N x dim]</code></p>
<div id="cell-14" class="cell" data-execution_count="4">
<div class="sourceCode cell-code" id="cb13" style="background: #f1f3f5;"><pre class="sourceCode python code-with-copy"><code class="sourceCode python"><span id="cb13-1"><span class="kw" style="color: #003B4F;
background-color: null;
font-weight: bold;
font-style: inherit;">def</span> convert_kpts_to_imc(cv2_kpts):</span>
<span id="cb13-2">    keypoints <span class="op" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">=</span> np.array([(x.pt[<span class="dv" style="color: #AD0000;
background-color: null;
font-style: inherit;">0</span>], x.pt[<span class="dv" style="color: #AD0000;
background-color: null;
font-style: inherit;">1</span>]) <span class="cf" style="color: #003B4F;
background-color: null;
font-weight: bold;
font-style: inherit;">for</span> x <span class="kw" style="color: #003B4F;
background-color: null;
font-weight: bold;
font-style: inherit;">in</span> cv2_kpts ]).reshape(<span class="op" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">-</span><span class="dv" style="color: #AD0000;
background-color: null;
font-style: inherit;">1</span>, <span class="dv" style="color: #AD0000;
background-color: null;
font-style: inherit;">2</span>)</span>
<span id="cb13-3">    scales <span class="op" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">=</span> np.array([<span class="fl" style="color: #AD0000;
background-color: null;
font-style: inherit;">12.0</span><span class="op" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">*</span> x.size <span class="cf" style="color: #003B4F;
background-color: null;
font-weight: bold;
font-style: inherit;">for</span> x <span class="kw" style="color: #003B4F;
background-color: null;
font-weight: bold;
font-style: inherit;">in</span> cv2_kpts ]).reshape(<span class="op" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">-</span><span class="dv" style="color: #AD0000;
background-color: null;
font-style: inherit;">1</span>, <span class="dv" style="color: #AD0000;
background-color: null;
font-style: inherit;">1</span>)</span>
<span id="cb13-4">    angles <span class="op" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">=</span> np.array([x.angle <span class="cf" style="color: #003B4F;
background-color: null;
font-weight: bold;
font-style: inherit;">for</span> x <span class="kw" style="color: #003B4F;
background-color: null;
font-weight: bold;
font-style: inherit;">in</span> cv2_kpts ]).reshape(<span class="op" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">-</span><span class="dv" style="color: #AD0000;
background-color: null;
font-style: inherit;">1</span>, <span class="dv" style="color: #AD0000;
background-color: null;
font-style: inherit;">1</span>)</span>
<span id="cb13-5">    responses <span class="op" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">=</span> np.array([x.response <span class="cf" style="color: #003B4F;
background-color: null;
font-weight: bold;
font-style: inherit;">for</span> x <span class="kw" style="color: #003B4F;
background-color: null;
font-weight: bold;
font-style: inherit;">in</span> cv2_kpts]).reshape(<span class="op" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">-</span><span class="dv" style="color: #AD0000;
background-color: null;
font-style: inherit;">1</span>, <span class="dv" style="color: #AD0000;
background-color: null;
font-style: inherit;">1</span>)</span>
<span id="cb13-6">    <span class="cf" style="color: #003B4F;
background-color: null;
font-weight: bold;
font-style: inherit;">return</span> keypoints, scales, angles, responses</span></code></pre></div>
</div>
<p>Now we are ready to write a script, which extracts local features for all images in the IMC-2021. The full script is accesible <a href="https://github.com/ducha-aiki/imc2021-sample-kornia-submission/blob/master/extract_features_imc.py">here</a></p>
<div id="cell-16" class="cell">
<div class="sourceCode cell-code" id="cb14" style="background: #f1f3f5;"><pre class="sourceCode python code-with-copy"><code class="sourceCode python"><span id="cb14-1"><span class="im" style="color: #00769E;
background-color: null;
font-style: inherit;">import</span> os</span>
<span id="cb14-2"><span class="im" style="color: #00769E;
background-color: null;
font-style: inherit;">import</span> h5py</span>
<span id="cb14-3"><span class="im" style="color: #00769E;
background-color: null;
font-style: inherit;">from</span> tqdm <span class="im" style="color: #00769E;
background-color: null;
font-style: inherit;">import</span> tqdm</span>
<span id="cb14-4"></span>
<span id="cb14-5"></span>
<span id="cb14-6">INPUT_DIR <span class="op" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">=</span> <span class="st" style="color: #20794D;
background-color: null;
font-style: inherit;">'../imc-2021-data'</span></span>
<span id="cb14-7">OUT_DIR <span class="op" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">=</span> <span class="st" style="color: #20794D;
background-color: null;
font-style: inherit;">'extracted/cv2-dog-affnet-hardnet8'</span></span>
<span id="cb14-8"></span>
<span id="cb14-9">os.makedirs(OUT_DIR, exist_ok<span class="op" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">=</span><span class="va" style="color: #111111;
background-color: null;
font-style: inherit;">True</span>)</span>
<span id="cb14-10">datasets <span class="op" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">=</span> os.listdir(INPUT_DIR)</span>
<span id="cb14-11"><span class="cf" style="color: #003B4F;
background-color: null;
font-weight: bold;
font-style: inherit;">for</span> ds <span class="kw" style="color: #003B4F;
background-color: null;
font-weight: bold;
font-style: inherit;">in</span> datasets:</span>
<span id="cb14-12">    ds_in_path <span class="op" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">=</span> os.path.join(INPUT_DIR, ds)</span>
<span id="cb14-13">    ds_out_path <span class="op" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">=</span> os.path.join(OUT_DIR, ds)</span>
<span id="cb14-14">    os.makedirs(ds_out_path, exist_ok<span class="op" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">=</span><span class="va" style="color: #111111;
background-color: null;
font-style: inherit;">True</span>)</span>
<span id="cb14-15">    seqs <span class="op" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">=</span> os.listdir(ds_in_path)</span>
<span id="cb14-16">    <span class="cf" style="color: #003B4F;
background-color: null;
font-weight: bold;
font-style: inherit;">for</span> seq <span class="kw" style="color: #003B4F;
background-color: null;
font-weight: bold;
font-style: inherit;">in</span> seqs:</span>
<span id="cb14-17">        <span class="cf" style="color: #003B4F;
background-color: null;
font-weight: bold;
font-style: inherit;">if</span> os.path.isdir(os.path.join(ds_in_path, seq, <span class="st" style="color: #20794D;
background-color: null;
font-style: inherit;">'set_100'</span>)):</span>
<span id="cb14-18">            seq_in_path <span class="op" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">=</span> os.path.join(ds_in_path, seq, <span class="st" style="color: #20794D;
background-color: null;
font-style: inherit;">'set_100'</span>, <span class="st" style="color: #20794D;
background-color: null;
font-style: inherit;">'images'</span>)</span>
<span id="cb14-19">        <span class="cf" style="color: #003B4F;
background-color: null;
font-weight: bold;
font-style: inherit;">else</span>:</span>
<span id="cb14-20">            seq_in_path <span class="op" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">=</span> os.path.join(ds_in_path, seq)</span>
<span id="cb14-21">        seq_out_path <span class="op" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">=</span> os.path.join(ds_out_path, seq)</span>
<span id="cb14-22">        os.makedirs(seq_out_path, exist_ok<span class="op" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">=</span><span class="va" style="color: #111111;
background-color: null;
font-style: inherit;">True</span>)</span>
<span id="cb14-23">        img_fnames <span class="op" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">=</span> os.listdir(seq_in_path)</span>
<span id="cb14-24">        num_kp <span class="op" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">=</span> []</span>
<span id="cb14-25">        <span class="cf" style="color: #003B4F;
background-color: null;
font-weight: bold;
font-style: inherit;">with</span> h5py.File(<span class="ss" style="color: #20794D;
background-color: null;
font-style: inherit;">f'</span><span class="sc" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">{</span>seq_out_path<span class="sc" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">}</span><span class="ss" style="color: #20794D;
background-color: null;
font-style: inherit;">/keypoints.h5'</span>, <span class="st" style="color: #20794D;
background-color: null;
font-style: inherit;">'w'</span>) <span class="im" style="color: #00769E;
background-color: null;
font-style: inherit;">as</span> f_kp, <span class="op" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">\</span></span>
<span id="cb14-26">             h5py.File(<span class="ss" style="color: #20794D;
background-color: null;
font-style: inherit;">f'</span><span class="sc" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">{</span>seq_out_path<span class="sc" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">}</span><span class="ss" style="color: #20794D;
background-color: null;
font-style: inherit;">/descriptors.h5'</span>, <span class="st" style="color: #20794D;
background-color: null;
font-style: inherit;">'w'</span>) <span class="im" style="color: #00769E;
background-color: null;
font-style: inherit;">as</span> f_desc, <span class="op" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">\</span></span>
<span id="cb14-27">             h5py.File(<span class="ss" style="color: #20794D;
background-color: null;
font-style: inherit;">f'</span><span class="sc" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">{</span>seq_out_path<span class="sc" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">}</span><span class="ss" style="color: #20794D;
background-color: null;
font-style: inherit;">/scores.h5'</span>, <span class="st" style="color: #20794D;
background-color: null;
font-style: inherit;">'w'</span>) <span class="im" style="color: #00769E;
background-color: null;
font-style: inherit;">as</span> f_score, <span class="op" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">\</span></span>
<span id="cb14-28">             h5py.File(<span class="ss" style="color: #20794D;
background-color: null;
font-style: inherit;">f'</span><span class="sc" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">{</span>seq_out_path<span class="sc" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">}</span><span class="ss" style="color: #20794D;
background-color: null;
font-style: inherit;">/angles.h5'</span> <span class="st" style="color: #20794D;
background-color: null;
font-style: inherit;">'w'</span>) <span class="im" style="color: #00769E;
background-color: null;
font-style: inherit;">as</span> f_ang, <span class="op" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">\</span></span>
<span id="cb14-29">             h5py.File(<span class="ss" style="color: #20794D;
background-color: null;
font-style: inherit;">f'</span><span class="sc" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">{</span>seq_out_path<span class="sc" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">}</span><span class="ss" style="color: #20794D;
background-color: null;
font-style: inherit;">/scales.h5'</span>, <span class="st" style="color: #20794D;
background-color: null;
font-style: inherit;">'w'</span>) <span class="im" style="color: #00769E;
background-color: null;
font-style: inherit;">as</span> f_scale:</span>
<span id="cb14-30">            <span class="cf" style="color: #003B4F;
background-color: null;
font-weight: bold;
font-style: inherit;">for</span> img_fname <span class="kw" style="color: #003B4F;
background-color: null;
font-weight: bold;
font-style: inherit;">in</span> tqdm(img_fnames):</span>
<span id="cb14-31">                img_fname_full <span class="op" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">=</span> os.path.join(seq_in_path, img_fname)</span>
<span id="cb14-32">                key <span class="op" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">=</span> os.path.splitext(os.path.basename(img_fname))[<span class="dv" style="color: #AD0000;
background-color: null;
font-style: inherit;">0</span>]</span>
<span id="cb14-33">                kpts, descs <span class="op" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">=</span> extract_features(img_fname_full,  sift_det, affnet, hardnet8, device, <span class="va" style="color: #111111;
background-color: null;
font-style: inherit;">False</span>)</span>
<span id="cb14-34">                keypoints, scales, angles, responses <span class="op" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">=</span> convert_kpts_to_imc(kpts)</span>
<span id="cb14-35">                f_kp[key] <span class="op" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">=</span> keypoints</span>
<span id="cb14-36">                f_desc[key] <span class="op" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">=</span> descs.reshape(<span class="op" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">-</span><span class="dv" style="color: #AD0000;
background-color: null;
font-style: inherit;">1</span>, <span class="dv" style="color: #AD0000;
background-color: null;
font-style: inherit;">128</span>)</span>
<span id="cb14-37">                f_score[key] <span class="op" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">=</span> responses</span>
<span id="cb14-38">                f_ang[key] <span class="op" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">=</span> angles</span>
<span id="cb14-39">                f_scale[key] <span class="op" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">=</span> scales</span>
<span id="cb14-40">                num_kp.append(<span class="bu" style="color: null;
background-color: null;
font-style: inherit;">len</span>(keypoints))</span>
<span id="cb14-41">            <span class="bu" style="color: null;
background-color: null;
font-style: inherit;">print</span>(<span class="ss" style="color: #20794D;
background-color: null;
font-style: inherit;">f'Finished processing "</span><span class="sc" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">{</span>ds<span class="sc" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">}</span><span class="ss" style="color: #20794D;
background-color: null;
font-style: inherit;">/</span><span class="sc" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">{</span>seq<span class="sc" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">}</span><span class="ss" style="color: #20794D;
background-color: null;
font-style: inherit;">" -&gt; </span><span class="sc" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">{</span>np<span class="sc" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">.</span>array(num_kp)<span class="sc" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">.</span>mean()<span class="sc" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">}</span><span class="ss" style="color: #20794D;
background-color: null;
font-style: inherit;"> features/image'</span>)</span></code></pre></div>
</div>
</section>
<section id="creating-config-json-file" class="level2">
<h2 class="anchored" data-anchor-id="creating-config-json-file">Creating config json file</h2>
<p>In addition to features, we should submit a config file, which tells the benchmark, how the features should be matched and which RANSAC we prefer. In priciple, we can just write an arbitrary config file and submit already, but this may lead to the bad results. Let’s instead generate a config file from python, so we can easily re-generate it. Why would we need this? Quite simple - to try different parameters on the validation set and only then create a final config.</p>
<p>First part of the config is <code>metadata</code> – information about the method and authors. If your method is under review, you may want to set flag <code>publish_anonymously</code> to <code>True</code>.</p>
<div id="cell-19" class="cell" data-execution_count="5">
<div class="sourceCode cell-code" id="cb15" style="background: #f1f3f5;"><pre class="sourceCode python code-with-copy"><code class="sourceCode python"><span id="cb15-1">metadata_dict <span class="op" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">=</span>  {</span>
<span id="cb15-2">    <span class="st" style="color: #20794D;
background-color: null;
font-style: inherit;">"publish_anonymously"</span>: <span class="va" style="color: #111111;
background-color: null;
font-style: inherit;">False</span>,</span>
<span id="cb15-3">    <span class="st" style="color: #20794D;
background-color: null;
font-style: inherit;">"authors"</span>: <span class="st" style="color: #20794D;
background-color: null;
font-style: inherit;">"Dmytro Mishkin, Milan Pultar and kornia team"</span>,</span>
<span id="cb15-4">    <span class="st" style="color: #20794D;
background-color: null;
font-style: inherit;">"contact_email"</span>: <span class="st" style="color: #20794D;
background-color: null;
font-style: inherit;">"ducha.aiki@gmail.com"</span>,</span>
<span id="cb15-5">    <span class="st" style="color: #20794D;
background-color: null;
font-style: inherit;">"method_name"</span>: <span class="st" style="color: #20794D;
background-color: null;
font-style: inherit;">"CV-DoG-AffNet-HardNet8 (kornia)"</span>,</span>
<span id="cb15-6">    <span class="st" style="color: #20794D;
background-color: null;
font-style: inherit;">"method_description"</span>: </span>
<span id="cb15-7">    <span class="vs" style="color: #20794D;
background-color: null;
font-style: inherit;">r"""OpeCV SIFT keypoints 8000 features, followed by the AffNet normalization </span></span>
<span id="cb15-8"><span class="vs" style="color: #20794D;
background-color: null;
font-style: inherit;">    and HardNet8 descriptor as implemented in kornia.</span></span>
<span id="cb15-9"><span class="vs" style="color: #20794D;
background-color: null;
font-style: inherit;">    Matched using the built-in matcher (bidirectional filter with the 'both' strategy, </span></span>
<span id="cb15-10"><span class="vs" style="color: #20794D;
background-color: null;
font-style: inherit;">    hopefully optimal inlier and ratio test thresholds) with DEGENSAC"""</span>,</span>
<span id="cb15-11">    <span class="st" style="color: #20794D;
background-color: null;
font-style: inherit;">"link_to_website"</span>: <span class="st" style="color: #20794D;
background-color: null;
font-style: inherit;">"https://github.com/kornia/kornia"</span>,</span>
<span id="cb15-12">    <span class="st" style="color: #20794D;
background-color: null;
font-style: inherit;">"link_to_pdf"</span>: <span class="st" style="color: #20794D;
background-color: null;
font-style: inherit;">"https://arxiv.org/abs/2007.09699"</span></span>
<span id="cb15-13">}</span></code></pre></div>
</div>
<p>Second part is <code>config_common</code>: it tells the benchmark, which keypoints and descriptors you use. We will also need this names when importing our features during tuning on the validation set.</p>
<div id="cell-21" class="cell" data-execution_count="6">
<div class="sourceCode cell-code" id="cb16" style="background: #f1f3f5;"><pre class="sourceCode python code-with-copy"><code class="sourceCode python"><span id="cb16-1">config_common_dict <span class="op" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">=</span>  {<span class="st" style="color: #20794D;
background-color: null;
font-style: inherit;">"json_label"</span>: <span class="st" style="color: #20794D;
background-color: null;
font-style: inherit;">"dog-affnet-hardnet8-degensac"</span>,</span>
<span id="cb16-2">    <span class="st" style="color: #20794D;
background-color: null;
font-style: inherit;">"keypoint"</span>: <span class="st" style="color: #20794D;
background-color: null;
font-style: inherit;">"cv2dog"</span>,</span>
<span id="cb16-3">    <span class="st" style="color: #20794D;
background-color: null;
font-style: inherit;">"descriptor"</span>: <span class="st" style="color: #20794D;
background-color: null;
font-style: inherit;">"affnethardnet8"</span>,</span>
<span id="cb16-4">    <span class="st" style="color: #20794D;
background-color: null;
font-style: inherit;">"num_keypoints"</span>: <span class="dv" style="color: #AD0000;
background-color: null;
font-style: inherit;">8000</span>}</span></code></pre></div>
</div>
<p>Now comes the information how to match your local features. It may vary from dataset to dataset and also be different for the <code>multiview</code> and <code>stereo</code> mode. That is why we will create a template dictionary and change some parameters later.</p>
<p>Specifically, we have to specify, which distance our descriptor prefers: <code>L2</code>, <code>L1</code> and <code>Hamming</code> are supported.</p>
<p>Then comes the tentative matches filtering. One can pass <code>none</code> for no filtering, <code>snn_ratio_pairwise</code> for Lowe’s SNN ratio and <code>fginn_ratio_pairwise</code> for FGINN. If you are not familiar with filtering strategies, checkout this blogpost: “<a href="https://ducha-aiki.medium.com/how-to-match-to-learn-or-not-to-learn-part-2-1ab52ede2022">How to match: to learn or not to learn?</a>”. The threshold is what have to be tunes.</p>
<p>We will use SNN, because of simplicity. Finally, we would like to make sure that tentative matches are cross-consistent, that is why we will enable <code>symmetric</code> matching.</p>
<p><strong>Warning!</strong> We will use <em>FLANN</em> approximate nearest neighbor matching for speed-up tuning procedure, but it is better to turn it off for the final submission.</p>
<div id="cell-23" class="cell" data-execution_count="7">
<div class="sourceCode cell-code" id="cb17" style="background: #f1f3f5;"><pre class="sourceCode python code-with-copy"><code class="sourceCode python"><span id="cb17-1"><span class="im" style="color: #00769E;
background-color: null;
font-style: inherit;">from</span> copy <span class="im" style="color: #00769E;
background-color: null;
font-style: inherit;">import</span> deepcopy</span>
<span id="cb17-2"></span>
<span id="cb17-3">matcher_template_dict <span class="op" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">=</span> {</span>
<span id="cb17-4">     <span class="st" style="color: #20794D;
background-color: null;
font-style: inherit;">"method"</span>: <span class="st" style="color: #20794D;
background-color: null;
font-style: inherit;">"nn"</span>,</span>
<span id="cb17-5">     <span class="st" style="color: #20794D;
background-color: null;
font-style: inherit;">"distance"</span>: <span class="st" style="color: #20794D;
background-color: null;
font-style: inherit;">"L2"</span>,</span>
<span id="cb17-6">     <span class="st" style="color: #20794D;
background-color: null;
font-style: inherit;">"flann"</span>: <span class="va" style="color: #111111;
background-color: null;
font-style: inherit;">True</span>,</span>
<span id="cb17-7">     <span class="st" style="color: #20794D;
background-color: null;
font-style: inherit;">"num_nn"</span>: <span class="dv" style="color: #AD0000;
background-color: null;
font-style: inherit;">1</span>,</span>
<span id="cb17-8">     <span class="st" style="color: #20794D;
background-color: null;
font-style: inherit;">"filtering"</span>: {</span>
<span id="cb17-9">         <span class="st" style="color: #20794D;
background-color: null;
font-style: inherit;">"type"</span>: <span class="st" style="color: #20794D;
background-color: null;
font-style: inherit;">"snn_ratio_pairwise"</span>,</span>
<span id="cb17-10">         <span class="st" style="color: #20794D;
background-color: null;
font-style: inherit;">"threshold"</span>: <span class="fl" style="color: #AD0000;
background-color: null;
font-style: inherit;">0.90</span></span>
<span id="cb17-11">     },</span>
<span id="cb17-12">     <span class="st" style="color: #20794D;
background-color: null;
font-style: inherit;">"symmetric"</span>: {</span>
<span id="cb17-13">         <span class="st" style="color: #20794D;
background-color: null;
font-style: inherit;">"enabled"</span>: <span class="va" style="color: #111111;
background-color: null;
font-style: inherit;">True</span>,</span>
<span id="cb17-14">         <span class="st" style="color: #20794D;
background-color: null;
font-style: inherit;">"reduce"</span>: <span class="st" style="color: #20794D;
background-color: null;
font-style: inherit;">"both"</span>,</span>
<span id="cb17-15">     }</span>
<span id="cb17-16">}</span></code></pre></div>
</div>
<p>Finally, we have to specify robust geometry estimation method. We will pick the default choise from the previous challenge - <a href="https://github.com/ducha-aiki/pydegensac">DEGENSAC</a>. <code>threshold</code> is what have to be tuned, the rest of parameters are already optimal, or fixed by the competition rules – <code>max_iter</code>.</p>
<div id="cell-25" class="cell" data-execution_count="8">
<div class="sourceCode cell-code" id="cb18" style="background: #f1f3f5;"><pre class="sourceCode python code-with-copy"><code class="sourceCode python"><span id="cb18-1">geom_template_dict <span class="op" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">=</span>  {<span class="st" style="color: #20794D;
background-color: null;
font-style: inherit;">"method"</span>: <span class="st" style="color: #20794D;
background-color: null;
font-style: inherit;">"cmp-degensac-f"</span>,</span>
<span id="cb18-2">                <span class="st" style="color: #20794D;
background-color: null;
font-style: inherit;">"threshold"</span>: <span class="fl" style="color: #AD0000;
background-color: null;
font-style: inherit;">0.5</span>,</span>
<span id="cb18-3">                <span class="st" style="color: #20794D;
background-color: null;
font-style: inherit;">"confidence"</span>: <span class="fl" style="color: #AD0000;
background-color: null;
font-style: inherit;">0.999999</span>,</span>
<span id="cb18-4">                <span class="st" style="color: #20794D;
background-color: null;
font-style: inherit;">"max_iter"</span>: <span class="dv" style="color: #AD0000;
background-color: null;
font-style: inherit;">100000</span>,</span>
<span id="cb18-5">                <span class="st" style="color: #20794D;
background-color: null;
font-style: inherit;">"error_type"</span>: <span class="st" style="color: #20794D;
background-color: null;
font-style: inherit;">"sampson"</span>,</span>
<span id="cb18-6">                <span class="st" style="color: #20794D;
background-color: null;
font-style: inherit;">"degeneracy_check"</span>: <span class="va" style="color: #111111;
background-color: null;
font-style: inherit;">True</span>,</span>
<span id="cb18-7">            }</span></code></pre></div>
</div>
<p>Let’s assemble and save our base config.</p>
<div id="cell-27" class="cell" data-execution_count="9">
<div class="sourceCode cell-code" id="cb19" style="background: #f1f3f5;"><pre class="sourceCode python code-with-copy"><code class="sourceCode python"><span id="cb19-1"><span class="im" style="color: #00769E;
background-color: null;
font-style: inherit;">import</span> json</span>
<span id="cb19-2"></span>
<span id="cb19-3">base_config <span class="op" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">=</span>  {</span>
<span id="cb19-4">    <span class="st" style="color: #20794D;
background-color: null;
font-style: inherit;">"metadata"</span>: metadata_dict,</span>
<span id="cb19-5">    <span class="st" style="color: #20794D;
background-color: null;
font-style: inherit;">"config_common"</span>: config_common_dict,</span>
<span id="cb19-6">    <span class="st" style="color: #20794D;
background-color: null;
font-style: inherit;">"config_phototourism_stereo"</span>: {</span>
<span id="cb19-7">        <span class="st" style="color: #20794D;
background-color: null;
font-style: inherit;">"use_custom_matches"</span>: <span class="va" style="color: #111111;
background-color: null;
font-style: inherit;">False</span>,</span>
<span id="cb19-8">        <span class="st" style="color: #20794D;
background-color: null;
font-style: inherit;">"matcher"</span>: deepcopy(matcher_template_dict),</span>
<span id="cb19-9">        <span class="st" style="color: #20794D;
background-color: null;
font-style: inherit;">"outlier_filter"</span>: { <span class="st" style="color: #20794D;
background-color: null;
font-style: inherit;">"method"</span>: <span class="st" style="color: #20794D;
background-color: null;
font-style: inherit;">"none"</span> },</span>
<span id="cb19-10">        <span class="st" style="color: #20794D;
background-color: null;
font-style: inherit;">"geom"</span>: deepcopy(geom_template_dict)</span>
<span id="cb19-11">        },</span>
<span id="cb19-12">    <span class="st" style="color: #20794D;
background-color: null;
font-style: inherit;">"config_phototourism_multiview"</span>: {</span>
<span id="cb19-13">        <span class="st" style="color: #20794D;
background-color: null;
font-style: inherit;">"use_custom_matches"</span>: <span class="va" style="color: #111111;
background-color: null;
font-style: inherit;">False</span>,</span>
<span id="cb19-14">        <span class="st" style="color: #20794D;
background-color: null;
font-style: inherit;">"matcher"</span>: deepcopy(matcher_template_dict),</span>
<span id="cb19-15">        <span class="st" style="color: #20794D;
background-color: null;
font-style: inherit;">"outlier_filter"</span>: { <span class="st" style="color: #20794D;
background-color: null;
font-style: inherit;">"method"</span>: <span class="st" style="color: #20794D;
background-color: null;
font-style: inherit;">"none"</span> },</span>
<span id="cb19-16">        <span class="st" style="color: #20794D;
background-color: null;
font-style: inherit;">"colmap"</span>: {}},</span>
<span id="cb19-17">    </span>
<span id="cb19-18">    <span class="st" style="color: #20794D;
background-color: null;
font-style: inherit;">"config_pragueparks_stereo"</span>: {</span>
<span id="cb19-19">        <span class="st" style="color: #20794D;
background-color: null;
font-style: inherit;">"use_custom_matches"</span>: <span class="va" style="color: #111111;
background-color: null;
font-style: inherit;">False</span>,</span>
<span id="cb19-20">        <span class="st" style="color: #20794D;
background-color: null;
font-style: inherit;">"matcher"</span>: deepcopy(matcher_template_dict),</span>
<span id="cb19-21">        <span class="st" style="color: #20794D;
background-color: null;
font-style: inherit;">"outlier_filter"</span>: { <span class="st" style="color: #20794D;
background-color: null;
font-style: inherit;">"method"</span>: <span class="st" style="color: #20794D;
background-color: null;
font-style: inherit;">"none"</span> },</span>
<span id="cb19-22">        <span class="st" style="color: #20794D;
background-color: null;
font-style: inherit;">"geom"</span>: deepcopy(geom_template_dict)</span>
<span id="cb19-23">        },</span>
<span id="cb19-24">    <span class="st" style="color: #20794D;
background-color: null;
font-style: inherit;">"config_pragueparks_multiview"</span>: {</span>
<span id="cb19-25">        <span class="st" style="color: #20794D;
background-color: null;
font-style: inherit;">"use_custom_matches"</span>: <span class="va" style="color: #111111;
background-color: null;
font-style: inherit;">False</span>,</span>
<span id="cb19-26">        <span class="st" style="color: #20794D;
background-color: null;
font-style: inherit;">"matcher"</span>: deepcopy(matcher_template_dict),</span>
<span id="cb19-27">        <span class="st" style="color: #20794D;
background-color: null;
font-style: inherit;">"outlier_filter"</span>: { <span class="st" style="color: #20794D;
background-color: null;
font-style: inherit;">"method"</span>: <span class="st" style="color: #20794D;
background-color: null;
font-style: inherit;">"none"</span> },</span>
<span id="cb19-28">        <span class="st" style="color: #20794D;
background-color: null;
font-style: inherit;">"colmap"</span>: {}},</span>
<span id="cb19-29">    <span class="st" style="color: #20794D;
background-color: null;
font-style: inherit;">"config_googleurban_stereo"</span>: {</span>
<span id="cb19-30">        <span class="st" style="color: #20794D;
background-color: null;
font-style: inherit;">"use_custom_matches"</span>: <span class="va" style="color: #111111;
background-color: null;
font-style: inherit;">False</span>,</span>
<span id="cb19-31">        <span class="st" style="color: #20794D;
background-color: null;
font-style: inherit;">"matcher"</span>: deepcopy(matcher_template_dict),</span>
<span id="cb19-32">        <span class="st" style="color: #20794D;
background-color: null;
font-style: inherit;">"outlier_filter"</span>: { <span class="st" style="color: #20794D;
background-color: null;
font-style: inherit;">"method"</span>: <span class="st" style="color: #20794D;
background-color: null;
font-style: inherit;">"none"</span> },</span>
<span id="cb19-33">        <span class="st" style="color: #20794D;
background-color: null;
font-style: inherit;">"geom"</span>: deepcopy(geom_template_dict)</span>
<span id="cb19-34">        },</span>
<span id="cb19-35">    <span class="st" style="color: #20794D;
background-color: null;
font-style: inherit;">"config_googleurban_multiview"</span>: {</span>
<span id="cb19-36">        <span class="st" style="color: #20794D;
background-color: null;
font-style: inherit;">"use_custom_matches"</span>: <span class="va" style="color: #111111;
background-color: null;
font-style: inherit;">False</span>,</span>
<span id="cb19-37">        <span class="st" style="color: #20794D;
background-color: null;
font-style: inherit;">"matcher"</span>: deepcopy(matcher_template_dict),</span>
<span id="cb19-38">        <span class="st" style="color: #20794D;
background-color: null;
font-style: inherit;">"outlier_filter"</span>: { <span class="st" style="color: #20794D;
background-color: null;
font-style: inherit;">"method"</span>: <span class="st" style="color: #20794D;
background-color: null;
font-style: inherit;">"none"</span> },</span>
<span id="cb19-39">        <span class="st" style="color: #20794D;
background-color: null;
font-style: inherit;">"colmap"</span>: {}}</span>
<span id="cb19-40">}</span></code></pre></div>
</div>
<p>Finally, benchmark expects multiple configs, so we have to create a list, and then we can save our config</p>
<div id="cell-29" class="cell" data-execution_count="10">
<div class="sourceCode cell-code" id="cb20" style="background: #f1f3f5;"><pre class="sourceCode python code-with-copy"><code class="sourceCode python"><span id="cb20-1"><span class="im" style="color: #00769E;
background-color: null;
font-style: inherit;">import</span> json </span>
<span id="cb20-2"></span>
<span id="cb20-3"><span class="cf" style="color: #003B4F;
background-color: null;
font-weight: bold;
font-style: inherit;">with</span> <span class="bu" style="color: null;
background-color: null;
font-style: inherit;">open</span>(<span class="st" style="color: #20794D;
background-color: null;
font-style: inherit;">'base_config.json'</span>, <span class="st" style="color: #20794D;
background-color: null;
font-style: inherit;">'w'</span>) <span class="im" style="color: #00769E;
background-color: null;
font-style: inherit;">as</span> f:</span>
<span id="cb20-4">    json.dump([base_config], f, indent<span class="op" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">=</span><span class="dv" style="color: #AD0000;
background-color: null;
font-style: inherit;">2</span>)</span></code></pre></div>
</div>
</section>
<section id="preliminary-evaluation" class="level1">
<h1>Preliminary evaluation</h1>
<p>Now let’s check how our features perform on validation set. We have to import our feature to the benchmark and run the benchmark.</p>
<p>I will cheat a little bit here and skip the multiview evaluation. The reason is that it requires <a href="https://colmap.github.io">colmap</a>, which might be not easy to install.</p>
<section id="importing-features" class="level3">
<h3 class="anchored" data-anchor-id="importing-features">Importing features</h3>
<p>Here we have to provide the same <code>keypoint</code> and <code>descriptor</code> names, as we wrote in json config. The rest of arguments are straightforward: path to features, json, etc.</p>
<div class="sourceCode" id="cb21" style="background: #f1f3f5;"><pre class="sourceCode bash code-with-copy"><code class="sourceCode bash"><span id="cb21-1"><span class="bu" style="color: null;
background-color: null;
font-style: inherit;">cd</span> ../image-matching-benchmark/</span>
<span id="cb21-2"><span class="ex" style="color: null;
background-color: null;
font-style: inherit;">python</span> <span class="at" style="color: #657422;
background-color: null;
font-style: inherit;">-utt</span> import_features.py <span class="at" style="color: #657422;
background-color: null;
font-style: inherit;">--kp_name</span> cv2dog <span class="at" style="color: #657422;
background-color: null;
font-style: inherit;">--desc_name</span> affnethardnet8 <span class="at" style="color: #657422;
background-color: null;
font-style: inherit;">--num_keypoints</span> 8000 <span class="at" style="color: #657422;
background-color: null;
font-style: inherit;">--path_features</span> ../imc2021-sample-kornia-submission/extracted/cv2-dog-affnet-hardnet8  <span class="at" style="color: #657422;
background-color: null;
font-style: inherit;">--path_results</span> ../benchmark-results <span class="at" style="color: #657422;
background-color: null;
font-style: inherit;">--subset</span> both  <span class="at" style="color: #657422;
background-color: null;
font-style: inherit;">--is_challenge</span> false <span class="at" style="color: #657422;
background-color: null;
font-style: inherit;">--path_json</span> ../imc2021-sample-kornia-submission/base_config.json  <span class="at" style="color: #657422;
background-color: null;
font-style: inherit;">--datasets</span> phototourism googleurban pragueparks</span></code></pre></div>
</section>
<section id="running-the-evaluation" class="level3">
<h3 class="anchored" data-anchor-id="running-the-evaluation">Running the evaluation</h3>
<p>Now we are ready to run the evaluation</p>
<div class="sourceCode" id="cb22" style="background: #f1f3f5;"><pre class="sourceCode bash code-with-copy"><code class="sourceCode bash"><span id="cb22-1"><span class="ex" style="color: null;
background-color: null;
font-style: inherit;">python</span> <span class="at" style="color: #657422;
background-color: null;
font-style: inherit;">-utt</span> run.py <span class="at" style="color: #657422;
background-color: null;
font-style: inherit;">--run_mode</span><span class="op" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">=</span>interactive <span class="at" style="color: #657422;
background-color: null;
font-style: inherit;">--json_method</span><span class="op" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">=</span>../imc2021-sample-kornia-submission/base_config.json <span class="at" style="color: #657422;
background-color: null;
font-style: inherit;">--subset</span><span class="op" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">=</span>val  <span class="at" style="color: #657422;
background-color: null;
font-style: inherit;">--eval_multiview</span><span class="op" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">=</span>False <span class="at" style="color: #657422;
background-color: null;
font-style: inherit;">--path_data</span> ../imc-2021-data/ <span class="at" style="color: #657422;
background-color: null;
font-style: inherit;">--path_results</span> ../benchmark-results <span class="at" style="color: #657422;
background-color: null;
font-style: inherit;">--is_challenge</span> false</span></code></pre></div>
<p>After a while (an 30 min for 32 cores machine), the process will finish and you will see the following log message:</p>
<pre><code> -- Saving to: "packed-val/dog-affnet-hardnet8-degensac.json"</code></pre>
</section>
<section id="reading-results" class="level3">
<h3 class="anchored" data-anchor-id="reading-results">Reading results</h3>
<p>Json file with evaluation results is saved to <code>image-matching-benchmark/packed-val/dog-affnet-hardnet8-degensac.json</code>, and some visualizations – to <code>../benchmark-visualization/png</code>.</p>
<p>First, we come back to our <code>imc2021-sample-kornia-submission</code> directory:</p>
<div class="sourceCode" id="cb24" style="background: #f1f3f5;"><pre class="sourceCode bash code-with-copy"><code class="sourceCode bash"><span id="cb24-1"><span class="bu" style="color: null;
background-color: null;
font-style: inherit;">cd</span> ../imc2021-sample-kornia-submission</span></code></pre></div>
<p>Metric, which are used for the competition is mean average accuracy (mAA) at visibility threshold 0.1</p>
<div id="cell-34" class="cell" data-execution_count="14">
<div class="sourceCode cell-code" id="cb25" style="background: #f1f3f5;"><pre class="sourceCode python code-with-copy"><code class="sourceCode python"><span id="cb25-1"><span class="im" style="color: #00769E;
background-color: null;
font-style: inherit;">import</span> os</span>
<span id="cb25-2"></span>
<span id="cb25-3">hashname<span class="op" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">=</span><span class="st" style="color: #20794D;
background-color: null;
font-style: inherit;">'dog-affnet-hardnet8-degensac'</span></span>
<span id="cb25-4">res_fname <span class="op" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">=</span> os.path.join(<span class="st" style="color: #20794D;
background-color: null;
font-style: inherit;">'../image-matching-benchmark/packed-val'</span>, <span class="ss" style="color: #20794D;
background-color: null;
font-style: inherit;">f'</span><span class="sc" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">{</span>hashname<span class="sc" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">}</span><span class="ss" style="color: #20794D;
background-color: null;
font-style: inherit;">.json'</span>)</span>
<span id="cb25-5"></span>
<span id="cb25-6"><span class="cf" style="color: #003B4F;
background-color: null;
font-weight: bold;
font-style: inherit;">with</span> <span class="bu" style="color: null;
background-color: null;
font-style: inherit;">open</span>(res_fname, <span class="st" style="color: #20794D;
background-color: null;
font-style: inherit;">'r'</span>) <span class="im" style="color: #00769E;
background-color: null;
font-style: inherit;">as</span> f:</span>
<span id="cb25-7">    results <span class="op" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">=</span> json.load(f)</span>
<span id="cb25-8"></span>
<span id="cb25-9">submission_name <span class="op" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">=</span> results[<span class="st" style="color: #20794D;
background-color: null;
font-style: inherit;">'config'</span>][<span class="st" style="color: #20794D;
background-color: null;
font-style: inherit;">'metadata'</span>][<span class="st" style="color: #20794D;
background-color: null;
font-style: inherit;">'method_name'</span>]</span>
<span id="cb25-10">datasets <span class="op" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">=</span> [<span class="st" style="color: #20794D;
background-color: null;
font-style: inherit;">'phototourism'</span>, <span class="st" style="color: #20794D;
background-color: null;
font-style: inherit;">'pragueparks'</span>, <span class="st" style="color: #20794D;
background-color: null;
font-style: inherit;">'googleurban'</span>]</span>
<span id="cb25-11">tasks <span class="op" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">=</span> [<span class="st" style="color: #20794D;
background-color: null;
font-style: inherit;">'stereo'</span>] <span class="co" style="color: #5E5E5E;
background-color: null;
font-style: inherit;"># ['stereo', 'multiview']  #Remember, that we skip colmap evaluations</span></span>
<span id="cb25-12">metric <span class="op" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">=</span> <span class="st" style="color: #20794D;
background-color: null;
font-style: inherit;">'qt_auc_10_th_0.1'</span></span>
<span id="cb25-13"><span class="cf" style="color: #003B4F;
background-color: null;
font-weight: bold;
font-style: inherit;">for</span> dset <span class="kw" style="color: #003B4F;
background-color: null;
font-weight: bold;
font-style: inherit;">in</span> datasets:</span>
<span id="cb25-14">    <span class="cf" style="color: #003B4F;
background-color: null;
font-weight: bold;
font-style: inherit;">for</span> task <span class="kw" style="color: #003B4F;
background-color: null;
font-weight: bold;
font-style: inherit;">in</span> tasks:</span>
<span id="cb25-15">        mAA <span class="op" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">=</span> results[dset][<span class="st" style="color: #20794D;
background-color: null;
font-style: inherit;">'results'</span>][<span class="st" style="color: #20794D;
background-color: null;
font-style: inherit;">'allseq'</span>][task][<span class="st" style="color: #20794D;
background-color: null;
font-style: inherit;">'run_avg'</span>][metric][<span class="st" style="color: #20794D;
background-color: null;
font-style: inherit;">'mean'</span>]</span>
<span id="cb25-16">        <span class="bu" style="color: null;
background-color: null;
font-style: inherit;">print</span> (<span class="ss" style="color: #20794D;
background-color: null;
font-style: inherit;">f'</span><span class="sc" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">{</span>submission_name<span class="sc" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">}</span><span class="ss" style="color: #20794D;
background-color: null;
font-style: inherit;"> </span><span class="sc" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">{</span>task<span class="sc" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">}</span><span class="ss" style="color: #20794D;
background-color: null;
font-style: inherit;"> mAA for </span><span class="sc" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">{</span>dset<span class="sc" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">}</span><span class="ss" style="color: #20794D;
background-color: null;
font-style: inherit;"> is </span><span class="sc" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">{</span>mAA<span class="sc" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">:.4f}</span><span class="ss" style="color: #20794D;
background-color: null;
font-style: inherit;">'</span>)</span></code></pre></div>
<div class="cell-output cell-output-stdout">
<pre><code>CV-DoG-AffNet-HardNet8 (kornia) stereo mAA for phototourism is 0.7108
CV-DoG-AffNet-HardNet8 (kornia) stereo mAA for pragueparks is 0.5850
CV-DoG-AffNet-HardNet8 (kornia) stereo mAA for googleurban is 0.3099</code></pre>
</div>
</div>
<p>We can also see results sequence-by-sequence</p>
<div id="cell-36" class="cell" data-execution_count="14">
<div class="sourceCode cell-code" id="cb27" style="background: #f1f3f5;"><pre class="sourceCode python code-with-copy"><code class="sourceCode python"><span id="cb27-1"><span class="im" style="color: #00769E;
background-color: null;
font-style: inherit;">import</span> seaborn <span class="im" style="color: #00769E;
background-color: null;
font-style: inherit;">as</span> sns</span>
<span id="cb27-2">sns.set_context(<span class="st" style="color: #20794D;
background-color: null;
font-style: inherit;">'paper'</span>, font_scale<span class="op" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">=</span><span class="fl" style="color: #AD0000;
background-color: null;
font-style: inherit;">1.7</span>)</span>
<span id="cb27-3">seqs <span class="op" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">=</span> []</span>
<span id="cb27-4">mAAs <span class="op" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">=</span> []</span>
<span id="cb27-5"><span class="cf" style="color: #003B4F;
background-color: null;
font-weight: bold;
font-style: inherit;">for</span> dset <span class="kw" style="color: #003B4F;
background-color: null;
font-weight: bold;
font-style: inherit;">in</span> datasets:</span>
<span id="cb27-6">    <span class="cf" style="color: #003B4F;
background-color: null;
font-weight: bold;
font-style: inherit;">for</span> task <span class="kw" style="color: #003B4F;
background-color: null;
font-weight: bold;
font-style: inherit;">in</span> tasks:</span>
<span id="cb27-7">        <span class="cf" style="color: #003B4F;
background-color: null;
font-weight: bold;
font-style: inherit;">for</span> seq <span class="kw" style="color: #003B4F;
background-color: null;
font-weight: bold;
font-style: inherit;">in</span> results[dset][<span class="st" style="color: #20794D;
background-color: null;
font-style: inherit;">'results'</span>].keys():</span>
<span id="cb27-8">            <span class="cf" style="color: #003B4F;
background-color: null;
font-weight: bold;
font-style: inherit;">if</span> seq <span class="op" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">==</span> <span class="st" style="color: #20794D;
background-color: null;
font-style: inherit;">'allseq'</span>: </span>
<span id="cb27-9">                <span class="cf" style="color: #003B4F;
background-color: null;
font-weight: bold;
font-style: inherit;">continue</span></span>
<span id="cb27-10">            mAA  <span class="op" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">=</span> results[dset][<span class="st" style="color: #20794D;
background-color: null;
font-style: inherit;">'results'</span>][seq][task][<span class="st" style="color: #20794D;
background-color: null;
font-style: inherit;">'run_avg'</span>][metric][<span class="st" style="color: #20794D;
background-color: null;
font-style: inherit;">'mean'</span>]</span>
<span id="cb27-11">            mAAs.append(mAA)</span>
<span id="cb27-12">            seqs.append(seq)</span>
<span id="cb27-13">fig, ax <span class="op" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">=</span> plt.subplots(figsize<span class="op" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">=</span>(<span class="dv" style="color: #AD0000;
background-color: null;
font-style: inherit;">15</span>,<span class="dv" style="color: #AD0000;
background-color: null;
font-style: inherit;">5</span>))</span>
<span id="cb27-14">xticks <span class="op" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">=</span> <span class="dv" style="color: #AD0000;
background-color: null;
font-style: inherit;">2</span><span class="op" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">*</span>np.arange(<span class="bu" style="color: null;
background-color: null;
font-style: inherit;">len</span>(seqs))</span>
<span id="cb27-15">ax.set_xticks(xticks)</span>
<span id="cb27-16">ax.bar(xticks, mAAs)</span>
<span id="cb27-17">ax.set_xticklabels(seqs)</span>
<span id="cb27-18">ax.set_ylabel(<span class="st" style="color: #20794D;
background-color: null;
font-style: inherit;">'mAA'</span>)</span>
<span id="cb27-19">ax.set_xlabel(<span class="st" style="color: #20794D;
background-color: null;
font-style: inherit;">'Sequence'</span>)</span></code></pre></div>
<div class="cell-output cell-output-display" data-execution_count="14">
<pre><code>Text(0.5, 0, 'Sequence')</code></pre>
</div>
<div class="cell-output cell-output-display">
<div>
<figure class="figure">
<p><img src="https://ducha-aiki.github.io/wide-baseline-stereo-blog/posts/2021-05-12-submitting-to-IMC2021-step-by-step_files/figure-html/cell-16-output-2.png" class="img-fluid figure-img"></p>
</figure>
</div>
</div>
</div>
<p>How do our feature correspondences look like?</p>
<div id="cell-38" class="cell" data-execution_count="16">
<div class="sourceCode cell-code" id="cb29" style="background: #f1f3f5;"><pre class="sourceCode python code-with-copy"><code class="sourceCode python"><span id="cb29-1"><span class="im" style="color: #00769E;
background-color: null;
font-style: inherit;">import</span> cv2</span>
<span id="cb29-2"><span class="kw" style="color: #003B4F;
background-color: null;
font-weight: bold;
font-style: inherit;">def</span> plot_image_grid(images, ncols<span class="op" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">=</span><span class="va" style="color: #111111;
background-color: null;
font-style: inherit;">None</span>, cmap<span class="op" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">=</span><span class="st" style="color: #20794D;
background-color: null;
font-style: inherit;">'gray'</span>):</span>
<span id="cb29-3">    <span class="co" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">#Taken from https://stackoverflow.com/a/66961099/1983544</span></span>
<span id="cb29-4">    <span class="co" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">'''Plot a grid of images'''</span></span>
<span id="cb29-5">    <span class="cf" style="color: #003B4F;
background-color: null;
font-weight: bold;
font-style: inherit;">if</span> <span class="kw" style="color: #003B4F;
background-color: null;
font-weight: bold;
font-style: inherit;">not</span> ncols:</span>
<span id="cb29-6">        factors <span class="op" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">=</span> [i <span class="cf" style="color: #003B4F;
background-color: null;
font-weight: bold;
font-style: inherit;">for</span> i <span class="kw" style="color: #003B4F;
background-color: null;
font-weight: bold;
font-style: inherit;">in</span> <span class="bu" style="color: null;
background-color: null;
font-style: inherit;">range</span>(<span class="dv" style="color: #AD0000;
background-color: null;
font-style: inherit;">1</span>, <span class="bu" style="color: null;
background-color: null;
font-style: inherit;">len</span>(images)<span class="op" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">+</span><span class="dv" style="color: #AD0000;
background-color: null;
font-style: inherit;">1</span>) <span class="cf" style="color: #003B4F;
background-color: null;
font-weight: bold;
font-style: inherit;">if</span> <span class="bu" style="color: null;
background-color: null;
font-style: inherit;">len</span>(images) <span class="op" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">%</span> i <span class="op" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">==</span> <span class="dv" style="color: #AD0000;
background-color: null;
font-style: inherit;">0</span>]</span>
<span id="cb29-7">        ncols <span class="op" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">=</span> factors[<span class="bu" style="color: null;
background-color: null;
font-style: inherit;">len</span>(factors) <span class="op" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">//</span> <span class="dv" style="color: #AD0000;
background-color: null;
font-style: inherit;">2</span>] <span class="cf" style="color: #003B4F;
background-color: null;
font-weight: bold;
font-style: inherit;">if</span> <span class="bu" style="color: null;
background-color: null;
font-style: inherit;">len</span>(factors) <span class="cf" style="color: #003B4F;
background-color: null;
font-weight: bold;
font-style: inherit;">else</span> <span class="bu" style="color: null;
background-color: null;
font-style: inherit;">len</span>(images) <span class="op" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">//</span> <span class="dv" style="color: #AD0000;
background-color: null;
font-style: inherit;">4</span> <span class="op" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">+</span> <span class="dv" style="color: #AD0000;
background-color: null;
font-style: inherit;">1</span></span>
<span id="cb29-8">    nrows <span class="op" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">=</span> <span class="bu" style="color: null;
background-color: null;
font-style: inherit;">int</span>(<span class="bu" style="color: null;
background-color: null;
font-style: inherit;">len</span>(images) <span class="op" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">/</span> ncols) <span class="op" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">+</span> <span class="bu" style="color: null;
background-color: null;
font-style: inherit;">int</span>(<span class="bu" style="color: null;
background-color: null;
font-style: inherit;">len</span>(images) <span class="op" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">%</span> ncols)</span>
<span id="cb29-9">    imgs <span class="op" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">=</span> [images[i] <span class="cf" style="color: #003B4F;
background-color: null;
font-weight: bold;
font-style: inherit;">if</span> <span class="bu" style="color: null;
background-color: null;
font-style: inherit;">len</span>(images) <span class="op" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">&gt;</span> i <span class="cf" style="color: #003B4F;
background-color: null;
font-weight: bold;
font-style: inherit;">else</span> <span class="va" style="color: #111111;
background-color: null;
font-style: inherit;">None</span> <span class="cf" style="color: #003B4F;
background-color: null;
font-weight: bold;
font-style: inherit;">for</span> i <span class="kw" style="color: #003B4F;
background-color: null;
font-weight: bold;
font-style: inherit;">in</span> <span class="bu" style="color: null;
background-color: null;
font-style: inherit;">range</span>(nrows <span class="op" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">*</span> ncols)]</span>
<span id="cb29-10">    f, axes <span class="op" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">=</span> plt.subplots(nrows, ncols, figsize<span class="op" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">=</span>(<span class="dv" style="color: #AD0000;
background-color: null;
font-style: inherit;">3</span><span class="op" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">*</span>ncols, <span class="dv" style="color: #AD0000;
background-color: null;
font-style: inherit;">2</span><span class="op" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">*</span>nrows))</span>
<span id="cb29-11">    axes <span class="op" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">=</span> axes.flatten()[:<span class="bu" style="color: null;
background-color: null;
font-style: inherit;">len</span>(imgs)]</span>
<span id="cb29-12">    <span class="cf" style="color: #003B4F;
background-color: null;
font-weight: bold;
font-style: inherit;">for</span> img, ax <span class="kw" style="color: #003B4F;
background-color: null;
font-weight: bold;
font-style: inherit;">in</span> <span class="bu" style="color: null;
background-color: null;
font-style: inherit;">zip</span>(imgs, axes.flatten()): </span>
<span id="cb29-13">        <span class="cf" style="color: #003B4F;
background-color: null;
font-weight: bold;
font-style: inherit;">if</span> np.<span class="bu" style="color: null;
background-color: null;
font-style: inherit;">any</span>(img):</span>
<span id="cb29-14">            <span class="cf" style="color: #003B4F;
background-color: null;
font-weight: bold;
font-style: inherit;">if</span> <span class="bu" style="color: null;
background-color: null;
font-style: inherit;">len</span>(img.shape) <span class="op" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">&gt;</span> <span class="dv" style="color: #AD0000;
background-color: null;
font-style: inherit;">2</span> <span class="kw" style="color: #003B4F;
background-color: null;
font-weight: bold;
font-style: inherit;">and</span> img.shape[<span class="dv" style="color: #AD0000;
background-color: null;
font-style: inherit;">2</span>] <span class="op" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">==</span> <span class="dv" style="color: #AD0000;
background-color: null;
font-style: inherit;">1</span>:</span>
<span id="cb29-15">                img <span class="op" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">=</span> img.squeeze()</span>
<span id="cb29-16">            ax.imshow(img, cmap<span class="op" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">=</span>cmap)</span>
<span id="cb29-17">imgs <span class="op" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">=</span> []</span>
<span id="cb29-18">VIS_DIR <span class="op" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">=</span> <span class="ss" style="color: #20794D;
background-color: null;
font-style: inherit;">f'../benchmark-visualization/png/</span><span class="sc" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">{</span>hashname<span class="sc" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">}</span><span class="ss" style="color: #20794D;
background-color: null;
font-style: inherit;">'</span></span>
<span id="cb29-19"><span class="cf" style="color: #003B4F;
background-color: null;
font-weight: bold;
font-style: inherit;">for</span> dset <span class="kw" style="color: #003B4F;
background-color: null;
font-weight: bold;
font-style: inherit;">in</span> os.listdir(VIS_DIR):</span>
<span id="cb29-20">    dset_dir <span class="op" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">=</span> os.path.join(VIS_DIR, dset)</span>
<span id="cb29-21">    <span class="cf" style="color: #003B4F;
background-color: null;
font-weight: bold;
font-style: inherit;">for</span> seq <span class="kw" style="color: #003B4F;
background-color: null;
font-weight: bold;
font-style: inherit;">in</span> os.listdir(dset_dir):</span>
<span id="cb29-22">        seq_dir <span class="op" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">=</span> os.path.join(dset_dir, seq, <span class="st" style="color: #20794D;
background-color: null;
font-style: inherit;">'stereo'</span>)</span>
<span id="cb29-23">        <span class="cf" style="color: #003B4F;
background-color: null;
font-weight: bold;
font-style: inherit;">for</span> img_fname <span class="kw" style="color: #003B4F;
background-color: null;
font-weight: bold;
font-style: inherit;">in</span> os.listdir(seq_dir):</span>
<span id="cb29-24">            full_fname <span class="op" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">=</span> os.path.join(seq_dir, img_fname)</span>
<span id="cb29-25">            img <span class="op" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">=</span> cv2.resize(cv2.cvtColor(cv2.imread(full_fname), cv2.COLOR_BGR2RGB), (<span class="dv" style="color: #AD0000;
background-color: null;
font-style: inherit;">200</span>,<span class="dv" style="color: #AD0000;
background-color: null;
font-style: inherit;">100</span>))</span>
<span id="cb29-26">            imgs.append(img)</span>
<span id="cb29-27">plot_image_grid(imgs)</span></code></pre></div>
<div class="cell-output cell-output-display">
<div>
<figure class="figure">
<p><img src="https://ducha-aiki.github.io/wide-baseline-stereo-blog/posts/2021-05-12-submitting-to-IMC2021-step-by-step_files/figure-html/cell-17-output-1.png" class="img-fluid figure-img"></p>
</figure>
</div>
</div>
</div>
</section>
<section id="hyperparameters-tuning" class="level2">
<h2 class="anchored" data-anchor-id="hyperparameters-tuning">Hyperparameters tuning</h2>
<p>Let’s now tune our hyperparameters, specifically, matching threshold and RANSAC inlier ratio. First we tune RANSAC and then matching ratio. We need to generate jsons for each configuration.</p>
<div id="cell-40" class="cell" data-execution_count="12">
<div class="sourceCode cell-code" id="cb30" style="background: #f1f3f5;"><pre class="sourceCode python code-with-copy"><code class="sourceCode python"><span id="cb30-1">inl_ths <span class="op" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">=</span> [<span class="fl" style="color: #AD0000;
background-color: null;
font-style: inherit;">0.25</span>, <span class="fl" style="color: #AD0000;
background-color: null;
font-style: inherit;">0.5</span>, <span class="fl" style="color: #AD0000;
background-color: null;
font-style: inherit;">0.75</span>, <span class="fl" style="color: #AD0000;
background-color: null;
font-style: inherit;">1.0</span>, <span class="fl" style="color: #AD0000;
background-color: null;
font-style: inherit;">1.5</span>, <span class="fl" style="color: #AD0000;
background-color: null;
font-style: inherit;">2.0</span>, <span class="fl" style="color: #AD0000;
background-color: null;
font-style: inherit;">2.5</span>, <span class="fl" style="color: #AD0000;
background-color: null;
font-style: inherit;">3.0</span>]</span>
<span id="cb30-2">configs <span class="op" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">=</span> []</span>
<span id="cb30-3"></span>
<span id="cb30-4"><span class="cf" style="color: #003B4F;
background-color: null;
font-weight: bold;
font-style: inherit;">for</span> inl_th <span class="kw" style="color: #003B4F;
background-color: null;
font-weight: bold;
font-style: inherit;">in</span> inl_ths:</span>
<span id="cb30-5">    current_config <span class="op" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">=</span> deepcopy(base_config)</span>
<span id="cb30-6">    <span class="cf" style="color: #003B4F;
background-color: null;
font-weight: bold;
font-style: inherit;">for</span> dset <span class="kw" style="color: #003B4F;
background-color: null;
font-weight: bold;
font-style: inherit;">in</span> [<span class="st" style="color: #20794D;
background-color: null;
font-style: inherit;">'phototourism'</span>, <span class="st" style="color: #20794D;
background-color: null;
font-style: inherit;">'pragueparks'</span>, <span class="st" style="color: #20794D;
background-color: null;
font-style: inherit;">'googleurban'</span>]:</span>
<span id="cb30-7">        current_config[<span class="ss" style="color: #20794D;
background-color: null;
font-style: inherit;">f'config_</span><span class="sc" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">{</span>dset<span class="sc" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">}</span><span class="ss" style="color: #20794D;
background-color: null;
font-style: inherit;">_stereo'</span>][<span class="st" style="color: #20794D;
background-color: null;
font-style: inherit;">'geom'</span>][<span class="st" style="color: #20794D;
background-color: null;
font-style: inherit;">'threshold'</span>] <span class="op" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">=</span> inl_th</span>
<span id="cb30-8">    label <span class="op" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">=</span> current_config[<span class="st" style="color: #20794D;
background-color: null;
font-style: inherit;">'config_common'</span>][<span class="st" style="color: #20794D;
background-color: null;
font-style: inherit;">'json_label'</span>] </span>
<span id="cb30-9">    current_config[<span class="st" style="color: #20794D;
background-color: null;
font-style: inherit;">'config_common'</span>][<span class="st" style="color: #20794D;
background-color: null;
font-style: inherit;">'json_label'</span>]  <span class="op" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">=</span> <span class="ss" style="color: #20794D;
background-color: null;
font-style: inherit;">f'</span><span class="sc" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">{</span>label<span class="sc" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">}</span><span class="ss" style="color: #20794D;
background-color: null;
font-style: inherit;">-inlth-</span><span class="sc" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">{</span>inl_th<span class="sc" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">}</span><span class="ss" style="color: #20794D;
background-color: null;
font-style: inherit;">'</span></span>
<span id="cb30-10">    configs.append(current_config)</span>
<span id="cb30-11"></span>
<span id="cb30-12"><span class="cf" style="color: #003B4F;
background-color: null;
font-weight: bold;
font-style: inherit;">with</span> <span class="bu" style="color: null;
background-color: null;
font-style: inherit;">open</span>(<span class="st" style="color: #20794D;
background-color: null;
font-style: inherit;">'ransac_tuning.json'</span>, <span class="st" style="color: #20794D;
background-color: null;
font-style: inherit;">'w'</span>) <span class="im" style="color: #00769E;
background-color: null;
font-style: inherit;">as</span> f:</span>
<span id="cb30-13">    json.dump(configs, f, indent<span class="op" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">=</span><span class="dv" style="color: #AD0000;
background-color: null;
font-style: inherit;">2</span>)</span></code></pre></div>
</div>
<div class="sourceCode" id="cb31" style="background: #f1f3f5;"><pre class="sourceCode bash code-with-copy"><code class="sourceCode bash"><span id="cb31-1"><span class="bu" style="color: null;
background-color: null;
font-style: inherit;">cd</span> ../image-matching-benchmark/</span>
<span id="cb31-2"><span class="ex" style="color: null;
background-color: null;
font-style: inherit;">python</span> <span class="at" style="color: #657422;
background-color: null;
font-style: inherit;">-utt</span> run.py <span class="at" style="color: #657422;
background-color: null;
font-style: inherit;">--run_mode</span><span class="op" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">=</span>interactive <span class="at" style="color: #657422;
background-color: null;
font-style: inherit;">--json_method</span><span class="op" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">=</span>../imc2021-sample-kornia-submission/base_config.json <span class="at" style="color: #657422;
background-color: null;
font-style: inherit;">--subset</span><span class="op" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">=</span>val  <span class="at" style="color: #657422;
background-color: null;
font-style: inherit;">--eval_multiview</span><span class="op" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">=</span>False <span class="at" style="color: #657422;
background-color: null;
font-style: inherit;">--path_data</span> ../imc-2021-data/ <span class="at" style="color: #657422;
background-color: null;
font-style: inherit;">--path_results</span> ../benchmark-results <span class="at" style="color: #657422;
background-color: null;
font-style: inherit;">--is_challenge</span> false</span>
<span id="cb31-3"><span class="bu" style="color: null;
background-color: null;
font-style: inherit;">cd</span> ../imc2021-sample-kornia-submission</span></code></pre></div>
<p>Now we will write a function, which reads results and picks the best threshold per dataset</p>
<div id="cell-42" class="cell" data-execution_count="15">
<div class="sourceCode cell-code" id="cb32" style="background: #f1f3f5;"><pre class="sourceCode python code-with-copy"><code class="sourceCode python"><span id="cb32-1">hashname<span class="op" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">=</span><span class="st" style="color: #20794D;
background-color: null;
font-style: inherit;">'dog-affnet-hardnet8-degensac'</span></span>
<span id="cb32-2">res_dict <span class="op" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">=</span> {}</span>
<span id="cb32-3">datasets <span class="op" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">=</span> [<span class="st" style="color: #20794D;
background-color: null;
font-style: inherit;">'phototourism'</span>, <span class="st" style="color: #20794D;
background-color: null;
font-style: inherit;">'pragueparks'</span>, <span class="st" style="color: #20794D;
background-color: null;
font-style: inherit;">'googleurban'</span>]</span>
<span id="cb32-4"><span class="cf" style="color: #003B4F;
background-color: null;
font-weight: bold;
font-style: inherit;">for</span> dset <span class="kw" style="color: #003B4F;
background-color: null;
font-weight: bold;
font-style: inherit;">in</span> datasets:</span>
<span id="cb32-5">    res_dict[dset] <span class="op" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">=</span> {}</span>
<span id="cb32-6"></span>
<span id="cb32-7">task <span class="op" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">=</span> <span class="st" style="color: #20794D;
background-color: null;
font-style: inherit;">'stereo'</span></span>
<span id="cb32-8">metric <span class="op" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">=</span> <span class="st" style="color: #20794D;
background-color: null;
font-style: inherit;">'qt_auc_10_th_0.1'</span></span>
<span id="cb32-9"><span class="cf" style="color: #003B4F;
background-color: null;
font-weight: bold;
font-style: inherit;">for</span> inl_th <span class="kw" style="color: #003B4F;
background-color: null;
font-weight: bold;
font-style: inherit;">in</span> inl_ths:</span>
<span id="cb32-10">    res_fname <span class="op" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">=</span> os.path.join(<span class="st" style="color: #20794D;
background-color: null;
font-style: inherit;">'../image-matching-benchmark/packed-val'</span>, <span class="ss" style="color: #20794D;
background-color: null;
font-style: inherit;">f'</span><span class="sc" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">{</span>hashname<span class="sc" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">}</span><span class="ss" style="color: #20794D;
background-color: null;
font-style: inherit;">-inlth-</span><span class="sc" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">{</span>inl_th<span class="sc" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">}</span><span class="ss" style="color: #20794D;
background-color: null;
font-style: inherit;">.json'</span>)</span>
<span id="cb32-11">    <span class="cf" style="color: #003B4F;
background-color: null;
font-weight: bold;
font-style: inherit;">try</span>:</span>
<span id="cb32-12">        <span class="cf" style="color: #003B4F;
background-color: null;
font-weight: bold;
font-style: inherit;">with</span> <span class="bu" style="color: null;
background-color: null;
font-style: inherit;">open</span>(res_fname, <span class="st" style="color: #20794D;
background-color: null;
font-style: inherit;">'r'</span>) <span class="im" style="color: #00769E;
background-color: null;
font-style: inherit;">as</span> f:</span>
<span id="cb32-13">            results <span class="op" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">=</span> json.load(f)</span>
<span id="cb32-14">    <span class="cf" style="color: #003B4F;
background-color: null;
font-weight: bold;
font-style: inherit;">except</span>:</span>
<span id="cb32-15">        <span class="cf" style="color: #003B4F;
background-color: null;
font-weight: bold;
font-style: inherit;">continue</span></span>
<span id="cb32-16">    submission_name <span class="op" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">=</span> results[<span class="st" style="color: #20794D;
background-color: null;
font-style: inherit;">'config'</span>][<span class="st" style="color: #20794D;
background-color: null;
font-style: inherit;">'metadata'</span>][<span class="st" style="color: #20794D;
background-color: null;
font-style: inherit;">'method_name'</span>]</span>
<span id="cb32-17">    res_dict[inl_th] <span class="op" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">=</span> {}</span>
<span id="cb32-18">    <span class="cf" style="color: #003B4F;
background-color: null;
font-weight: bold;
font-style: inherit;">for</span> dset <span class="kw" style="color: #003B4F;
background-color: null;
font-weight: bold;
font-style: inherit;">in</span> datasets:</span>
<span id="cb32-19">        mAA <span class="op" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">=</span> results[dset][<span class="st" style="color: #20794D;
background-color: null;
font-style: inherit;">'results'</span>][<span class="st" style="color: #20794D;
background-color: null;
font-style: inherit;">'allseq'</span>][task][<span class="st" style="color: #20794D;
background-color: null;
font-style: inherit;">'run_avg'</span>][metric][<span class="st" style="color: #20794D;
background-color: null;
font-style: inherit;">'mean'</span>]</span>
<span id="cb32-20">        res_dict[dset][inl_th] <span class="op" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">=</span> mAA</span>
<span id="cb32-21">fig, ax <span class="op" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">=</span> plt.subplots(figsize<span class="op" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">=</span>(<span class="dv" style="color: #AD0000;
background-color: null;
font-style: inherit;">5</span>,<span class="dv" style="color: #AD0000;
background-color: null;
font-style: inherit;">5</span>))</span>
<span id="cb32-22">colors <span class="op" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">=</span> [<span class="st" style="color: #20794D;
background-color: null;
font-style: inherit;">'r'</span>,<span class="st" style="color: #20794D;
background-color: null;
font-style: inherit;">'b'</span>,<span class="st" style="color: #20794D;
background-color: null;
font-style: inherit;">'k'</span>]</span>
<span id="cb32-23">final_ths <span class="op" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">=</span> {}</span>
<span id="cb32-24"><span class="cf" style="color: #003B4F;
background-color: null;
font-weight: bold;
font-style: inherit;">for</span> i, dset <span class="kw" style="color: #003B4F;
background-color: null;
font-weight: bold;
font-style: inherit;">in</span> <span class="bu" style="color: null;
background-color: null;
font-style: inherit;">enumerate</span>(datasets):</span>
<span id="cb32-25">    inl_ths <span class="op" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">=</span> []</span>
<span id="cb32-26">    mAAs <span class="op" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">=</span> []</span>
<span id="cb32-27">    <span class="cf" style="color: #003B4F;
background-color: null;
font-weight: bold;
font-style: inherit;">for</span> inl_th, mAA <span class="kw" style="color: #003B4F;
background-color: null;
font-weight: bold;
font-style: inherit;">in</span> res_dict[dset].items():</span>
<span id="cb32-28">        inl_ths.append(inl_th)</span>
<span id="cb32-29">        mAAs.append(mAA)</span>
<span id="cb32-30">    best_th_idx <span class="op" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">=</span> np.argmax(np.array(mAAs))</span>
<span id="cb32-31">    best_th <span class="op" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">=</span> inl_ths[best_th_idx]</span>
<span id="cb32-32">    best_mAA <span class="op" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">=</span> mAAs[best_th_idx]</span>
<span id="cb32-33">    <span class="bu" style="color: null;
background-color: null;
font-style: inherit;">print</span> (<span class="ss" style="color: #20794D;
background-color: null;
font-style: inherit;">f'Best </span><span class="sc" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">{</span>dset<span class="sc" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">}</span><span class="ss" style="color: #20794D;
background-color: null;
font-style: inherit;"> mAA = </span><span class="sc" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">{</span>best_mAA<span class="sc" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">:.4f}</span><span class="ss" style="color: #20794D;
background-color: null;
font-style: inherit;"> with inl_th = </span><span class="sc" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">{</span>best_th<span class="sc" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">}</span><span class="ss" style="color: #20794D;
background-color: null;
font-style: inherit;">'</span>)</span>
<span id="cb32-34">    ax.plot(inl_ths, mAAs, label<span class="op" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">=</span>dset, color<span class="op" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">=</span>colors[i])</span>
<span id="cb32-35">    ax.plot(best_th, best_mAA, label <span class="op" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">=</span> <span class="ss" style="color: #20794D;
background-color: null;
font-style: inherit;">f'</span><span class="sc" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">{</span>dset<span class="sc" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">}</span><span class="ss" style="color: #20794D;
background-color: null;
font-style: inherit;">-best'</span>, marker<span class="op" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">=</span><span class="st" style="color: #20794D;
background-color: null;
font-style: inherit;">'x'</span>, linestyle<span class="op" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">=</span><span class="st" style="color: #20794D;
background-color: null;
font-style: inherit;">''</span>, color<span class="op" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">=</span>colors[i])</span>
<span id="cb32-36">    final_ths[dset] <span class="op" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">=</span> best_th</span>
<span id="cb32-37">    </span>
<span id="cb32-38">ax.legend()</span>
<span id="cb32-39">ax.set_ylabel(<span class="st" style="color: #20794D;
background-color: null;
font-style: inherit;">'mAA'</span>)</span>
<span id="cb32-40">ax.set_xlabel(<span class="st" style="color: #20794D;
background-color: null;
font-style: inherit;">'DEGENSAC inlier threshold'</span>)</span>
<span id="cb32-41">    </span></code></pre></div>
<div class="cell-output cell-output-stdout">
<pre><code>Best phototourism mAA = 0.7108 with inl_th = 0.5
Best pragueparks mAA = 0.6700 with inl_th = 1.5
Best googleurban mAA = 0.3116 with inl_th = 0.75</code></pre>
</div>
<div class="cell-output cell-output-display" data-execution_count="15">
<pre><code>Text(0.5, 0, 'DEGENSAC inlier threshold')</code></pre>
</div>
<div class="cell-output cell-output-display">
<div>
<figure class="figure">
<p><img src="https://ducha-aiki.github.io/wide-baseline-stereo-blog/posts/2021-05-12-submitting-to-IMC2021-step-by-step_files/figure-html/cell-19-output-3.png" class="img-fluid figure-img"></p>
</figure>
</div>
</div>
</div>
</section>
<section id="creating-final-submission" class="level2">
<h2 class="anchored" data-anchor-id="creating-final-submission">Creating final submission</h2>
<p>Its time to create our final submission!</p>
<div id="cell-44" class="cell" data-execution_count="16">
<div class="sourceCode cell-code" id="cb35" style="background: #f1f3f5;"><pre class="sourceCode python code-with-copy"><code class="sourceCode python"><span id="cb35-1"></span>
<span id="cb35-2">configs <span class="op" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">=</span> []</span>
<span id="cb35-3">current_config <span class="op" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">=</span> deepcopy(base_config)</span>
<span id="cb35-4"><span class="cf" style="color: #003B4F;
background-color: null;
font-weight: bold;
font-style: inherit;">for</span> dset <span class="kw" style="color: #003B4F;
background-color: null;
font-weight: bold;
font-style: inherit;">in</span> [<span class="st" style="color: #20794D;
background-color: null;
font-style: inherit;">'phototourism'</span>, <span class="st" style="color: #20794D;
background-color: null;
font-style: inherit;">'pragueparks'</span>, <span class="st" style="color: #20794D;
background-color: null;
font-style: inherit;">'googleurban'</span>]:</span>
<span id="cb35-5">    current_config[<span class="ss" style="color: #20794D;
background-color: null;
font-style: inherit;">f'config_</span><span class="sc" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">{</span>dset<span class="sc" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">}</span><span class="ss" style="color: #20794D;
background-color: null;
font-style: inherit;">_stereo'</span>][<span class="st" style="color: #20794D;
background-color: null;
font-style: inherit;">'geom'</span>][<span class="st" style="color: #20794D;
background-color: null;
font-style: inherit;">'threshold'</span>] <span class="op" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">=</span> final_ths[dset]</span>
<span id="cb35-6">    <span class="co" style="color: #5E5E5E;
background-color: null;
font-style: inherit;"># I did a little bit of tuning offline for multiview, so we will put it here</span></span>
<span id="cb35-7">    current_config[<span class="ss" style="color: #20794D;
background-color: null;
font-style: inherit;">f'config_</span><span class="sc" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">{</span>dset<span class="sc" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">}</span><span class="ss" style="color: #20794D;
background-color: null;
font-style: inherit;">_multiview'</span>][<span class="st" style="color: #20794D;
background-color: null;
font-style: inherit;">'matcher'</span>][<span class="st" style="color: #20794D;
background-color: null;
font-style: inherit;">'filtering'</span>][<span class="st" style="color: #20794D;
background-color: null;
font-style: inherit;">'threshold'</span>] <span class="op" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">=</span> <span class="fl" style="color: #AD0000;
background-color: null;
font-style: inherit;">0.95</span></span>
<span id="cb35-8">    </span>
<span id="cb35-9">    <span class="co" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">#Remember, that we should not forget to turn FLANN ofd</span></span>
<span id="cb35-10">    current_config[<span class="ss" style="color: #20794D;
background-color: null;
font-style: inherit;">f'config_</span><span class="sc" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">{</span>dset<span class="sc" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">}</span><span class="ss" style="color: #20794D;
background-color: null;
font-style: inherit;">_multiview'</span>][<span class="st" style="color: #20794D;
background-color: null;
font-style: inherit;">'matcher'</span>][<span class="st" style="color: #20794D;
background-color: null;
font-style: inherit;">'flann'</span>] <span class="op" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">=</span> <span class="va" style="color: #111111;
background-color: null;
font-style: inherit;">False</span></span>
<span id="cb35-11">    current_config[<span class="ss" style="color: #20794D;
background-color: null;
font-style: inherit;">f'config_</span><span class="sc" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">{</span>dset<span class="sc" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">}</span><span class="ss" style="color: #20794D;
background-color: null;
font-style: inherit;">_stereo'</span>][<span class="st" style="color: #20794D;
background-color: null;
font-style: inherit;">'matcher'</span>][<span class="st" style="color: #20794D;
background-color: null;
font-style: inherit;">'flann'</span>] <span class="op" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">=</span> <span class="va" style="color: #111111;
background-color: null;
font-style: inherit;">False</span></span>
<span id="cb35-12">    </span>
<span id="cb35-13">    </span>
<span id="cb35-14">current_config[<span class="st" style="color: #20794D;
background-color: null;
font-style: inherit;">'metadata'</span>][<span class="st" style="color: #20794D;
background-color: null;
font-style: inherit;">'method_name'</span>] <span class="op" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">=</span> <span class="st" style="color: #20794D;
background-color: null;
font-style: inherit;">'KORNIA TUTORIAL CV-DoG-AffNet-HardNet8'</span></span>
<span id="cb35-15"></span>
<span id="cb35-16">label <span class="op" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">=</span> current_config[<span class="st" style="color: #20794D;
background-color: null;
font-style: inherit;">'config_common'</span>][<span class="st" style="color: #20794D;
background-color: null;
font-style: inherit;">'json_label'</span>] </span>
<span id="cb35-17">current_config[<span class="st" style="color: #20794D;
background-color: null;
font-style: inherit;">'config_common'</span>][<span class="st" style="color: #20794D;
background-color: null;
font-style: inherit;">'json_label'</span>]  <span class="op" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">=</span> <span class="ss" style="color: #20794D;
background-color: null;
font-style: inherit;">f'</span><span class="sc" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">{</span>label<span class="sc" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">}</span><span class="ss" style="color: #20794D;
background-color: null;
font-style: inherit;">'</span></span>
<span id="cb35-18">configs.append(current_config)</span>
<span id="cb35-19"></span>
<span id="cb35-20"><span class="bu" style="color: null;
background-color: null;
font-style: inherit;">print</span> (current_config)</span>
<span id="cb35-21"><span class="cf" style="color: #003B4F;
background-color: null;
font-weight: bold;
font-style: inherit;">with</span> <span class="bu" style="color: null;
background-color: null;
font-style: inherit;">open</span>(<span class="st" style="color: #20794D;
background-color: null;
font-style: inherit;">'final_submission.json'</span>, <span class="st" style="color: #20794D;
background-color: null;
font-style: inherit;">'w'</span>) <span class="im" style="color: #00769E;
background-color: null;
font-style: inherit;">as</span> f:</span>
<span id="cb35-22">    json.dump(configs, f, indent<span class="op" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">=</span><span class="dv" style="color: #AD0000;
background-color: null;
font-style: inherit;">2</span>)</span></code></pre></div>
<div class="cell-output cell-output-stdout">
<pre><code>{'metadata': {'publish_anonymously': False, 'authors': 'Dmytro Mishkin, Milan Pultar and kornia team', 'contact_email': 'ducha.aiki@gmail.com', 'method_name': 'KORNIA TUTORIAL CV-DoG-AffNet-HardNet8', 'method_description': "OpeCV SIFT keypoints 8000 features, followed by the AffNet normalization \n    and HardNet8 descriptor as implemented in kornia.\n    Matched using the built-in matcher (bidirectional filter with the 'both' strategy, \n    hopefully optimal inlier and ratio test thresholds) with DEGENSAC", 'link_to_website': 'https://github.com/kornia/kornia', 'link_to_pdf': 'https://arxiv.org/abs/2007.09699'}, 'config_common': {'json_label': 'dog-affnet-hardnet8-degensac', 'keypoint': 'cv2dog', 'descriptor': 'affnethardnet8', 'num_keypoints': 8000}, 'config_phototourism_stereo': {'use_custom_matches': False, 'matcher': {'method': 'nn', 'distance': 'L2', 'flann': False, 'num_nn': 1, 'filtering': {'type': 'snn_ratio_pairwise', 'threshold': 0.9}, 'symmetric': {'enabled': True, 'reduce': 'both'}}, 'outlier_filter': {'method': 'none'}, 'geom': {'method': 'cmp-degensac-f', 'threshold': 0.5, 'confidence': 0.999999, 'max_iter': 100000, 'error_type': 'sampson', 'degeneracy_check': True}}, 'config_phototourism_multiview': {'use_custom_matches': False, 'matcher': {'method': 'nn', 'distance': 'L2', 'flann': False, 'num_nn': 1, 'filtering': {'type': 'snn_ratio_pairwise', 'threshold': 0.95}, 'symmetric': {'enabled': True, 'reduce': 'both'}}, 'outlier_filter': {'method': 'none'}, 'colmap': {}}, 'config_pragueparks_stereo': {'use_custom_matches': False, 'matcher': {'method': 'nn', 'distance': 'L2', 'flann': False, 'num_nn': 1, 'filtering': {'type': 'snn_ratio_pairwise', 'threshold': 0.9}, 'symmetric': {'enabled': True, 'reduce': 'both'}}, 'outlier_filter': {'method': 'none'}, 'geom': {'method': 'cmp-degensac-f', 'threshold': 1.5, 'confidence': 0.999999, 'max_iter': 100000, 'error_type': 'sampson', 'degeneracy_check': True}}, 'config_pragueparks_multiview': {'use_custom_matches': False, 'matcher': {'method': 'nn', 'distance': 'L2', 'flann': False, 'num_nn': 1, 'filtering': {'type': 'snn_ratio_pairwise', 'threshold': 0.95}, 'symmetric': {'enabled': True, 'reduce': 'both'}}, 'outlier_filter': {'method': 'none'}, 'colmap': {}}, 'config_googleurban_stereo': {'use_custom_matches': False, 'matcher': {'method': 'nn', 'distance': 'L2', 'flann': False, 'num_nn': 1, 'filtering': {'type': 'snn_ratio_pairwise', 'threshold': 0.9}, 'symmetric': {'enabled': True, 'reduce': 'both'}}, 'outlier_filter': {'method': 'none'}, 'geom': {'method': 'cmp-degensac-f', 'threshold': 0.75, 'confidence': 0.999999, 'max_iter': 100000, 'error_type': 'sampson', 'degeneracy_check': True}}, 'config_googleurban_multiview': {'use_custom_matches': False, 'matcher': {'method': 'nn', 'distance': 'L2', 'flann': False, 'num_nn': 1, 'filtering': {'type': 'snn_ratio_pairwise', 'threshold': 0.95}, 'symmetric': {'enabled': True, 'reduce': 'both'}}, 'outlier_filter': {'method': 'none'}, 'colmap': {}}}</code></pre>
</div>
</div>
<pre><code>Submission Zip file should have folder structure as follow:
 ├── config.json
 ├── [Dataset 1]
 │   ├── [Sequence 1]
 │   │   ├── keypoints.h5
 │   │   ├── descriptors.h5
 │   │   ├── matches.h5
 │   ├── [Sequence 2]
 │   │   ├── ...
 ├── [Dataset 2]
 │   ├── ...</code></pre>
<p>So we have to just copy our features, add config and zip them.</p>
<div class="sourceCode" id="cb38" style="background: #f1f3f5;"><pre class="sourceCode bash code-with-copy"><code class="sourceCode bash"><span id="cb38-1"><span class="fu" style="color: #4758AB;
background-color: null;
font-style: inherit;">cp</span> final_submission.json extracted/cv2-dog-affnet-hardnet8/config.json</span>
<span id="cb38-2"><span class="bu" style="color: null;
background-color: null;
font-style: inherit;">cd</span> extracted/cv2-dog-affnet-hardnet8</span>
<span id="cb38-3"><span class="fu" style="color: #4758AB;
background-color: null;
font-style: inherit;">zip</span> <span class="at" style="color: #657422;
background-color: null;
font-style: inherit;">-r</span> submission.zip <span class="pp" style="color: #AD0000;
background-color: null;
font-style: inherit;">*</span></span></code></pre></div>
<p>Last step before the submission - check the submission for correctness with provided script</p>
<div class="sourceCode" id="cb39" style="background: #f1f3f5;"><pre class="sourceCode bash code-with-copy"><code class="sourceCode bash"><span id="cb39-1"><span class="bu" style="color: null;
background-color: null;
font-style: inherit;">cd</span> ../../../image-matching-benchmark</span>
<span id="cb39-2"><span class="ex" style="color: null;
background-color: null;
font-style: inherit;">python</span> <span class="at" style="color: #657422;
background-color: null;
font-style: inherit;">-utt</span> submission_validator.py <span class="at" style="color: #657422;
background-color: null;
font-style: inherit;">--submit_file_path</span> ../imc2021-sample-kornia-submission/extracted/cv2-dog-affnet-hardnet8/submission.zip <span class="at" style="color: #657422;
background-color: null;
font-style: inherit;">--benchmark_repo_path</span> . <span class="at" style="color: #657422;
background-color: null;
font-style: inherit;">--raw_data_path</span> ../imc-2021-data/ <span class="at" style="color: #657422;
background-color: null;
font-style: inherit;">--datasets</span> googleurban phototourism pragueparks</span></code></pre></div>
<p>If everything is correct, you will see:</p>
<pre><code>Validating method 1/1: "dog-affnet-hardnet8-degensac"
['googleurban', 'phototourism', 'pragueparks']
Running: googleurban, stereo track
Running: googleurban, multiview track
Running: phototourism, stereo track
Running: phototourism, multiview track
Running: pragueparks, stereo track
Running: pragueparks, multiview track
Validating key "config_googleurban_stereo"
Validating key "config_googleurban_multiview"
Validating key "config_phototourism_stereo"
Validating key "config_phototourism_multiview"
Validating key "config_pragueparks_stereo"
Validating key "config_pragueparks_multiview"</code></pre>
<p>And file <code>submission_log.txt</code> will appear near our .zip file.</p>
<div class="sourceCode" id="cb41" style="background: #f1f3f5;"><pre class="sourceCode bash code-with-copy"><code class="sourceCode bash"><span id="cb41-1"><span class="fu" style="color: #4758AB;
background-color: null;
font-style: inherit;">cat</span> ../imc2021-sample-kornia-submission/extracted/cv2-dog-affnet-hardnet8/submission_log.txt</span></code></pre></div>
<pre><code>Submission is in proper format, please submit to IMW 2021 website.</code></pre>
<p>That’s all, folks! We can submit! But, please, do not just submit this sample submission - make your own :)</p>


</section>
</section>

<p>Everything you (didn’t) want to know about image matching</p> ]]></description>
  <guid>https://ducha-aiki.github.io/wide-baseline-stereo-blog/posts/2021-05-12-submitting-to-imc2021-step-by-step.html</guid>
  <pubDate>Tue, 11 May 2021 22:00:00 GMT</pubDate>
  <media:content url="https://ducha-aiki.github.io/wide-baseline-stereo-blog/posts/images/wooden_lady.png" medium="image" type="image/png" height="114" width="91"/>
</item>
<item>
  <title>WxBS step by step</title>
  <dc:creator>Dmytro Mishkin</dc:creator>
  <link>https://ducha-aiki.github.io/wide-baseline-stereo-blog/posts/2021-02-11-wxbs-step-by-step.html</link>
  <description><![CDATA[ 





<section id="delving-deeper-into-wxbs-algorithm-steps" class="level2">
<h2 class="anchored" data-anchor-id="delving-deeper-into-wxbs-algorithm-steps">Delving deeper into WxBS algorithm steps</h2>
<div class="quarto-figure quarto-figure-center">
<figure class="figure">
<p><img src="https://ducha-aiki.github.io/wide-baseline-stereo-blog/posts/00_intro_files/matching-filtering.png" class="img-fluid figure-img"></p>
<figcaption>Commonly used wide baseline stereo pipeline</figcaption>
</figure>
</div>
<p>The wide baseline stereo problem is commonly addressed by a family of algorithms, the general structure of which is shown in Figure above. We will be referring to it as the WxBS pipeline or the WxBS algorithm. Let us describe it in more detail and discuss the reasoning behind each block.</p>
<ol type="1">
<li>A set of the local features (Also known as keypoints, local regions, distinguished regions, salient regions, salient points, etc.) is detected in each image independently. In automated systems the local features are usually low level structures like corners, blobs and so on. However, they can also be more high level semantic structures, as we used in the example in <a href="https://ducha-aiki.github.io/wide-baseline-stereo-blog/posts/2021-01-09-wxbs-in-simple-terms.html">intro</a>: “a long staircase on the left side”, “the top of the lampost” and so on. An important detail is that detection is typically done in each image separately. Why is it the case? If the task is to match only a single image pair, that would be an unnecessary restriction. It is even benefitial to process the images jointly, as a human would do, by placing both images side-by-side and looking at them back and forth. However, the wide baseline stereo task rarely arises by itself, more often it is only a part of a bigger system, e.g.&nbsp;visual localization or 3D recontruction from the collection of images. Therefore, one needs to match an image to not the one, but multiple other images. That is why it is benefitial to perform feature extraction only once per image and then load the stored results. Moreover, independent feature extraction is a task which is easy to parallelize and that is typically done in most of libraries and frameworks for the WxBS. One could be wondering if the local feature detection process is necessary at all? Indeed, it is possible to avoid feature detection and consider all the pixels as “detections”. The problem with such approach is the high computational and memory complexity – even a small 800x600 image contains half a million pixels, which need to be matched to half a million pixels in another image.</li>
<li>A region around the local feature to be described is selected. If one considers a keypoint to be literally a point, it is impractical to distinguish between them based only on coordinates and, maybe, the single RGB value of the pixel. On the other extreme, part of the image, which really far from the current keypoint helps little to nothing in terms of finding a correspondence. Thus, a reasonable trade-off needs to be made. Keypoint therefore can be think of as the “reference point of the distinguished region”, e.g.&nbsp;a center of the blob. It worth mention that some detectors return a region by default, so this step is omitted, or, to be precise, included into step 1 “local features detection”. However, it is useful to have it discussed separately .</li>
<li>A patch around each local feature is described with a local feature descriptor, i.e.&nbsp;converted to a compact format. Such procedure also should be robust to changes in acquisition conditions so that descriptors related to the same 3D points are similar and dissimilar otherwise. The local feature descriptors are then used for the efficient generation of tentative correspondences. Could one skip this stage? Yes, but as with the local feature detection, the skipping is not desirable from a computational point of view – the benefits are discussed in the next stage – matching. Local feature detection, measurement region selection and description together convert an image into a sparse representation, which is suitable for the correspondence search. Such representation is more robust to the acquisition conditions and can be further indexed if used for image retrieval.</li>
<li>Tentative correspondences between detected features are established and then filtered. The simplest and common way to generate tentative correspondences is to perform a nearest neighbor search in the descriptor space. The commonly used descriptors are the binary or float point vectors, which allows to employ various algorithms for approximate nearest neighbor search and trade a small amount of accuracy for orders of magnitude speed-up. Such correspondences need to be filtered, that is why the are called “tentative” or “putative” – a significant percantage of them is incorrect. There are many reasons for that – imprerfection of the local feature detector, descriptor, matching process and simply the fact that some parts of the scene are visible only on one image, but not another.</li>
<li>The geometric relationship between the two images is recovered, which is the final goal of the whole process. In addition, the tentative correspondences, which are not consistent with the found geometry, are called outliers and are discarded. The most common way of robust model estimation in the presense of outliers is called RANSAC – random sample consensus. There are other methods as well, e.g.&nbsp;re-weighted least squares, but RANSAC predominantely used in practice.</li>
</ol>


</section>

<p>Everything you (didn’t) want to know about image matching</p> ]]></description>
  <guid>https://ducha-aiki.github.io/wide-baseline-stereo-blog/posts/2021-02-11-wxbs-step-by-step.html</guid>
  <pubDate>Wed, 10 Feb 2021 23:00:00 GMT</pubDate>
  <media:content url="https://ducha-aiki.github.io/wide-baseline-stereo-blog/posts/images/matching-filtering_small.png" medium="image" type="image/png" height="26" width="144"/>
</item>
<item>
  <title>Wide multiple baseline stereo in simple terms</title>
  <dc:creator>Dmytro Mishkin</dc:creator>
  <link>https://ducha-aiki.github.io/wide-baseline-stereo-blog/posts/2021-01-09-wxbs-in-simple-terms.html</link>
  <description><![CDATA[ 





<section id="what-is-wide-multiple-baseline-stereo" class="level2">
<h2 class="anchored" data-anchor-id="what-is-wide-multiple-baseline-stereo">What is wide multiple baseline stereo?</h2>
<div class="quarto-figure quarto-figure-center">
<figure class="figure">
<p><img src="https://ducha-aiki.github.io/wide-baseline-stereo-blog/posts/2020-03-27-intro_files/att_00001.png" class="img-fluid figure-img"></p>
<figcaption>Two images of the same place, taken from different viewpoint and in different times. The building: Kyiv doll theater.</figcaption>
</figure>
</div>
<p>Imagine you have a nice photo you took in autumn and would like to take one in summer, from the same spot. How would you achieve that? You go to the place and you start to compare what you see on the camera screen and on the printed photo. Specifically, you would probably try to locate the same objects, e.g., “that high lamppost” or “this wall clock”. Then one would estimate how differently they are arranged on the old photo and camera screen. For example, by checking whether the lamppost is occluding the clock on the tower or not. That would give an idea of how you should move your camera.</p>
<p>Now, what if you are not allowed to take that photo with you, because it is a museum photo and taking pictures is prohibited there. Instead you can create a description of it. In that case, it is likely that you would try to make a list of features and objects in the photo together with the descriptions, which are sufficient to distinguish the objects. For example, “a long staircase on the left side”, “The nice building with a dark roof and two towers” or “the top of the lampost”. It would be useful to also describe where these objects and features are pictured in the photo, “The lamp posts are on the left, the closest to the viewer is in front of the left tower with a clock. The clock tower is not the part of the building and stands on its own”. Then when arriving, you would try to find those objects, match them to the description you have, and try to estimate where you should go. You repeat the procedure until the camera screen shows a picture which is fitting the description you have and the image you have in your memory.</p>
<p>Congratulations! You just have successfully registered the two images, which have a significant difference in viewpoint, appearance, and illumination. In the process of doing so, you were solving multiple times the wide multiple-baseline stereo problem (WxBS) – estimating the relative camera pose from a pair of images, different in many aspects, yet depicting the same scene.</p>
<p>Let us write down the steps, which we took.</p>
<ol type="1">
<li><p>Identify salient objects and features in the images – “trees”, “statues”, “tip of the tower”, etc in images.</p></li>
<li><p>Describe the objects and features, taking into account their neighborhood: “statue with a blue left ear”.</p></li>
<li><p>Establish potential correspondences between features in the different images, based on their descriptors.</p></li>
<li><p>Estimate, which direction one should move the camera to align the objects and features.</p></li>
</ol>
<p>That is it! detailed explanation of the each of the steps is <a href="https://ducha-aiki.github.io/wide-baseline-stereo-blog/posts/2021-02-11-wxbs-step-by-step.html">in this post</a>. If you are interested in the formal definition, <a href="https://ducha-aiki.github.io/wide-baseline-stereo-blog/posts/2020-07-09-wxbs.html">check here</a>, and the <a href="https://ducha-aiki.github.io/wide-baseline-stereo-blog/posts/2020-03-27-intro.html">history of the WxBS is here</a>.</p>


</section>

<p>Everything you (didn’t) want to know about image matching</p> ]]></description>
  <guid>https://ducha-aiki.github.io/wide-baseline-stereo-blog/posts/2021-01-09-wxbs-in-simple-terms.html</guid>
  <pubDate>Fri, 08 Jan 2021 23:00:00 GMT</pubDate>
  <media:content url="https://ducha-aiki.github.io/wide-baseline-stereo-blog/posts/images/doll-theater.jpg" medium="image" type="image/jpeg"/>
</item>
<item>
  <title>Lessons learned and future directions</title>
  <dc:creator>Dmytro Mishkin</dc:creator>
  <link>https://ducha-aiki.github.io/wide-baseline-stereo-blog/posts/2020-11-26-lessons-and-future-directions.html</link>
  <description><![CDATA[ 





<p>I would like to share some lessons I have learned about wide baseline stereo and propose some research directions, which are worth exploring in the short and longer term.</p>
<section id="lessons-learned" class="level1">
<h1>Lessons learned</h1>
<section id="benchmark-metrics-implementation-dataset" class="level2">
<h2 class="anchored" data-anchor-id="benchmark-metrics-implementation-dataset">1. Benchmark = metrics + implementation + dataset</h2>
<p>In our paper “<a href="https://arxiv.org/abs/2003.01587">Image Matching across Wide Baselines: From Paper to Practice</a>” we focused on the first two parts. Specifically, metrics – if they are not “downstream” metrics, the improvements in the single component might not translate to the overall system improvements. And implementation – implementing the simplest possible setup is, of course, a valuable tool, but one have to also incorporate the best known practices, e.g.&nbsp;matching, RANSAC tuning as so on.</p>
<p>I hope, that we have delivered that message to the community. But the last component – the dataset – we have, perhaps, overlooked a bit ourself. The problem with dataset limitations, e.g.&nbsp;lack of illumination or seasonal changes is not that one does not properly address. Usually benchmark papers are adressing their limitations quite clearly. The problem is that researchers (including myself) have tendency to work on improving results, which are easily measurable, therefore the implicitly designing the methods, which solve only some specific problem, encoded in the form of the dataset.</p>
</section>
<section id="trying-to-work-for-the-most-general-case-might-be-detrimental-for-the-practical-applications." class="level2">
<h2 class="anchored" data-anchor-id="trying-to-work-for-the-most-general-case-might-be-detrimental-for-the-practical-applications.">2. Trying to work for the “most general” case might be detrimental for the practical applications.</h2>
<p>This is kind of opposite side of the lesson 1. For example, the classical SIFT matching is rotation-invariant, because that was assumed to be the requirement to work in the “real world”. However, the practice of image retrieval and then most of learned local features like R2D2, SuperPoint, DELF and so on, showed that “up-is-up” is a reasonable assumption to built on. The rotational invariance in lots of scenarios hurt more than helps.</p>
</section>
<section id="when-borrow-idea-from-classical-paper-adapt-it" class="level2">
<h2 class="anchored" data-anchor-id="when-borrow-idea-from-classical-paper-adapt-it">3. When borrow idea from classical paper, adapt it</h2>
<p><a href="https://proceedings.neurips.cc/paper/2017/hash/831caa1b600f852b7844499430ecac17-Abstract.html">HardNet</a> borrows the idea of the using <a href="https://medium.com/@ducha.aiki/how-to-match-to-learn-or-not-to-learn-part-2-1ab52ede2022">second nearest neighbor</a>(SNN) descriptor from the SIFT. However, using SNN ratio for descriptor learning leads to inferior results. We had to modify it to the triplet margin loss.</p>
<p>## 4. Classic handcrafted algorithms are not dead and can be improved a lot</p>
<p>Modern versions of RANSAC are still necessary for two view matching even when so complex methods as SuperGlue as used. Moreover, they still have quite a lot things for improvement, as proved by my colleague, <a href="https://scholar.google.hu/citations?user=U9-D8DYAAAAJ&amp;hl=en">Daniel Barath</a></p>
<p>They also have a benefit, that once you have an idea, you can make a paper out it faster, as you don’t need gathering data and training the model.</p>
</section>
</section>
<section id="future-research-directions" class="level1">
<h1>Future research directions</h1>
<section id="application-specific-local-features" class="level2">
<h2 class="anchored" data-anchor-id="application-specific-local-features">1. Application-specific local features</h2>
<p>One thing about the current local features is that they are kind of universal. SIFT works reasonably well for lots of applications: 3d reconstruction, SLAM, image retrieval, etc. The learned features, like SuperPoint or R2D2, although are biased towards the data they are trained on, still don’t have anything domain specific.</p>
<p>Let me explain. There are different qualities about the local feature (detectors). It can be more or less robust to nuisance factors like the illumination and camera position. It can be more or less precisely localized. It can be more or less dense and/or evenly distributed over the image.</p>
<p>For example, in image retrieval, one does not really care about precise localization, the robustness is much more important. For the 3d reconstruction one would like to have a lot of 3d points to get the reasonable reconstruction. On the other hand, for the SLAM/relocalization application, sparse features would be more advantageous because of smaller memory footprint and computational cost.</p>
<p>There are, actually, some steps in that direction. Let me name a few.</p>
<ol type="1">
<li><p><a href="https://arxiv.org/pdf/2007.13172.pdf">HOW local features</a> designed for the image retrieval <img src="https://ducha-aiki.github.io/wide-baseline-stereo-blog/posts/2020-11-26-lessons-and-future-directions_files/att_00000.png" class="img-fluid" alt="Robust, but poorly localized HOW local features for imare retrieval, from ‘Learning and aggregating deep local descriptors for instance-level recognition’ paper"></p></li>
<li><p><a href="http://rpg.ifi.uzh.ch/docs/3DV19_Cieslewski.pdf">SIPs: Succinct Interest Points from Unsupervised Inlierness Probability Learning</a> sparse local features for the SLAM. <!--![A small number of keypoints might be enough for estimating the camera pose. Figure from SIPs paper](2020-11-26-lessons-and-future-directions_files/att_00001.png) --></p></li>
<li><p><a href="https://arxiv.org/abs/1912.00623">Reinforced Feature Points: Optimizing Feature Detection and Description for a High-Level Task</a> <img src="https://ducha-aiki.github.io/wide-baseline-stereo-blog/posts/2020-11-26-lessons-and-future-directions_files/att_00002.png" class="img-fluid" alt="Optimizing local features for non-differentiable high-level task, Figure from Reinforced Feature Points"></p></li>
</ol>
<p>I believe, that it is only beginning and we are yet to experience AlphaZero moment for the local features.</p>
</section>
<section id="rethinking-overall-wide-baseline-stereo-pipeline-optimized-for-the-specific-application" class="level2">
<h2 class="anchored" data-anchor-id="rethinking-overall-wide-baseline-stereo-pipeline-optimized-for-the-specific-application">2. Rethinking overall wide baseline stereo pipeline, optimized for the specific application</h2>
<p>It is often perceived, that image matching across the unordered collection of the images is a task with quadratic complexity w.r.t. number of images. Some operations can be done separately, e.g.&nbsp;feature detection, but others, like feature matching and RANSAC cannot. Right?</p>
<p>Not necessarily. It turns out, that one can avoid running feature matching and RANSAC for more than 90% of image pairs with clever preprocessing, ordering and re-using results from the previous matching. Moreover, in order to do the whole task faster (matching image collections), one may need to introduce additional steps, which are not necessary, or slowing things down for the two images case.</p>
<p>That’s what we done for the <a href="https://arxiv.org/abs/2011.11986">intial pose estimation for the global SfM</a> in our paper, which reduced the matching runtime from 200 hours to 29.</p>
<!-- ![Table from the Efficient Initial Pose-graph Generation for Global SfM](2020-11-26-lessons-and-future-directions_files/att_00003.png) -->
<p>Another example would be “<a href="https://arxiv.org/abs/1911.11763">SuperGlue: Learning Feature Matching with Graph Neural Networks</a>”, where authors abandoned traditional descriptor matching and instead leveraged all the information for keypoint and descriptors from both images altogether.</p>
</section>
<section id="rethinking-and-improving-the-process-of-training-data-generation-for-the-wbs" class="level2">
<h2 class="anchored" data-anchor-id="rethinking-and-improving-the-process-of-training-data-generation-for-the-wbs">3. Rethinking and improving the process of training data generation for the WBS</h2>
<p>So far, all the local feature papers I have seen rely on one of the ground truth source.</p>
<ol type="1">
<li>SfM data, obtained with COLMAP with, possibly, a cleaned depth information.</li>
<li>Affine and color augmentation.</li>
<li>Synthetic images (e.g.&nbsp;corners for SuperPoint).</li>
</ol>
<p>There are several problems with them.</p>
<p><strong>SfM data</strong> assumes that the data is matchable by the existing methods, at least for a some extent. That might not always be true for cross-seasonal, medical and other kind of data. It is also not applicable for historical photographies and other types of data. Moreover, SfM data takes quite a time for compute and space to store. I believe that we may do better.</p>
<p><strong>Affine and color augmentation</strong> can take us only this far – we actually want our detectors and descriptors to be robust to the changes, which we don’t know how to simulate/augment.</p>
<p><strong>Synthetic images</strong> as they are in, say, <a href="https://carla.org/">CARLA</a> simulator lack fine details and photorealism. However, I am optimistic about using neural renderers and learned wide baseline stereo is a GAN-like self-improving loop.</p>
</section>
<section id="matching-with-on-demand-view-synthesis-revisited" class="level2">
<h2 class="anchored" data-anchor-id="matching-with-on-demand-view-synthesis-revisited">4. Matching with On-Demand View Synthesis revisited</h2>
<p>I like the “on-demand” principle a lot and I think we can explore it much more that we are now. So far we have either <a href="https://ducha-aiki.github.io/wide-baseline-stereo-blog/posts/2020-08-06-affine-view-synthesis.html">affine view synthesis (ASIFT, MODS)</a>, or <a href="https://people.ee.ethz.ch/~timofter/publications/Anoosheh-ICRA-2019.pdf">GAN-based stylizations</a> for the day-night matching.</p>
<p>That is why I am glad to see papers like <a href="https://arxiv.org/abs/2008.09497">Single-Image Depth Prediction Makes Feature Matching Easier</a>, which generate normalized views based on depth in order to help the matching.</p>
<p>Why not go further? Combine viewpoint, illumination, season, sensor synthesis?</p>
</section>
<section id="moar-inputs" class="level2">
<h2 class="anchored" data-anchor-id="moar-inputs">5. Moar inputs!</h2>
<p>I have mentioned above that monocular depth may help the feature matching or <a href="https://ducha-aiki.github.io/wide-baseline-stereo-blog/2020/07/17/affine-correspondences.html#Back-to-wide-baseline-stereo">camera pose estimation</a>. However, why stop here?</p>
<p>Let’s use other networks as well, especially given that we will need them on robot or vehicle anyway. Semantic segmentation? Yes, please. Surface normals? Why not? Intrinsic images? Йой, най буде! {% fn 1 %}</p>
<p>{{ “Ukrainian, means let is be” | fndetail: 1 }} <img src="https://ducha-aiki.github.io/wide-baseline-stereo-blog/posts/2020-11-26-lessons-and-future-directions_files/att_00004.png" class="img-fluid" alt="image.png"></p>
</section>
<section id="what-do-you-think-would-be-good-idea-for-the-wbs-research" class="level2">
<h2 class="anchored" data-anchor-id="what-do-you-think-would-be-good-idea-for-the-wbs-research">What do you think would be good idea for the WBS research?</h2>
<p>Let me know in comments/twitter. I am also going to update this page from time to time</p>


</section>
</section>

<p>Everything you (didn’t) want to know about image matching</p> ]]></description>
  <guid>https://ducha-aiki.github.io/wide-baseline-stereo-blog/posts/2020-11-26-lessons-and-future-directions.html</guid>
  <pubDate>Wed, 25 Nov 2020 23:00:00 GMT</pubDate>
  <media:content url="https://ducha-aiki.github.io/wide-baseline-stereo-blog/posts/images/yoy_cat.jpg" medium="image" type="image/jpeg"/>
</item>
<item>
  <title>Benchmarking Image Retrieval for Visual Localization</title>
  <dc:creator>Dmytro Mishkin</dc:creator>
  <link>https://ducha-aiki.github.io/wide-baseline-stereo-blog/posts/2020-11-25-review-of-retrieval-for-localization.html</link>
  <description><![CDATA[ 





<p>I would like to share my thoughts on 3DV 2020 paper “<a href="https://arxiv.org/pdf/2011.11946.pdf">Benchmarking Image Retrieval for Visual Localization</a>” by Pion et.al.</p>
<div class="quarto-figure quarto-figure-center">
<figure class="figure">
<p><img src="https://ducha-aiki.github.io/wide-baseline-stereo-blog/posts/2020-11-25-review-of-retrieval-for-localization_files/att_00000.png" class="img-fluid figure-img"></p>
<figcaption>Overview of the benchmark pipeline</figcaption>
</figure>
</div>
<section id="what-is-the-paper-about" class="level1">
<h1>What is the paper about?</h1>
<p>How one would approach visual localization? The most viable way to do it is hierarchical approach, similar to image retrieval with spatial verification.</p>
<p>You get the query image, retrieve the most similar images to it from some database by some efficient method, e.g.&nbsp;global descriptor search. Then given the top-k images you estimate the pose of the query image by doing two-view matching, or some other method.</p>
<p>The question is – how much influence the quality of the image retrieval has? Should you spend more or less time improving it? That is the questions, paper trying to answer.</p>
<p>Authors design 3 re-localization systems.</p>
<ol type="1">
<li><p>“Task 1” system estimates the query image pose as average of the short-list images pose, weighted by the similarity to the query image.</p></li>
<li><p>“Task 2a” system performs pairwise two-view matching between short-list images and query triangulates the query image pose using 3d map built from the successully matches images.</p></li>
<li><p>“Task 2b” pre-builds the 3D map from the database images offline. At the inference time, local feature 2d-3d matching is done on the shortlist images.</p></li>
</ol>
</section>
<section id="what-is-benchmarked" class="level1">
<h1>What is benchmarked?</h1>
<p>The paper compares</p>
<ul>
<li>SIFT-based DenseVLAD</li>
</ul>
<p>and CNN-based</p>
<ul>
<li><a href="https://arxiv.org/abs/1511.07247">NetVLAD</a></li>
<li><a href="https://arxiv.org/pdf/1906.07589.pdf">APGeM</a></li>
<li><a href="https://arxiv.org/abs/2001.05027">DELG</a></li>
</ul>
<p>What is important (and adequately mentioned in the paper, although I would prefer the disclamer in the each figure) is that <strong>all CNN-based methods have different architectures AND training data</strong>. Basically, the paper uses author-released models. Thus one cannot say if APGeM is better or worse than NetVLAD as method, because they were trained on the very different data. However, I also understand that one cannot easily afford to re-implement and re-train everything.</p>
<p>As the sanity check paper provides the results on the <a href="http://cmp.felk.cvut.cz/revisitop/">Revisited Oxford and Paris</a> image retrieval benchmark. <img src="https://ducha-aiki.github.io/wide-baseline-stereo-blog/posts/2020-11-25-review-of-retrieval-for-localization_files/att_00001.png" class="img-fluid" alt="Retrieval metrics on Revised Oxford and Paris"></p>
</section>
<section id="summary-of-results" class="level1">
<h1>Summary of results</h1>
<p>Paper contains a lot of information and I definitely recommend you to read it. Nevertheless, let me try to summarize paper messages and then my take on it.</p>
<ol type="1">
<li><p>For the task1 (similarity-weighted pose) there is no clear winner. (SIFT)-DenseVLAD works the best for the daytime datasets. Probably DenseVLAD is good because it is not invariant and if it can match images, they are really close -&gt; high pose accuracy. For the night both DeLG and AP-GeM are good. As paper guesses, that it because they are only ones, which were trained on night images as well. <img src="https://ducha-aiki.github.io/wide-baseline-stereo-blog/posts/2020-11-25-review-of-retrieval-for-localization_files/att_00002.png" class="img-fluid" alt="image.png"></p></li>
<li><p>There is almost no difference between CNN-based methods for the task2a and task2b (retrieval -&gt; local features matching). This indicates that the limit is the mostly in the number of images and local features.</p></li>
</ol>
<div class="quarto-figure quarto-figure-center">
<figure class="figure">
<p><img src="https://ducha-aiki.github.io/wide-baseline-stereo-blog/posts/2020-11-25-review-of-retrieval-for-localization_files/att_00003.png" class="img-fluid figure-img"></p>
<figcaption>image.png</figcaption>
</figure>
</div>
<div class="quarto-figure quarto-figure-center">
<figure class="figure">
<p><img src="https://ducha-aiki.github.io/wide-baseline-stereo-blog/posts/2020-11-25-review-of-retrieval-for-localization_files/att_00004.png" class="img-fluid figure-img"></p>
<figcaption>image.png</figcaption>
</figure>
</div>
</section>
<section id="my-take-away-messages" class="level1">
<h1>My take-away messages</h1>
<section id="image-relocalization-seems-to-be-is-more-real-world-and-engineering-task-than-image-retrieval." class="level3">
<h3 class="anchored" data-anchor-id="image-relocalization-seems-to-be-is-more-real-world-and-engineering-task-than-image-retrieval.">Image Relocalization seems to be is more real-world and engineering task, than image retrieval.</h3>
<p>And that it why it actually ALREADY WORKS, because if there some weak spot, it is compensated by the system design. Thhe same conclusion from our <a href="https://arxiv.org/abs/2003.01587">IMC paper</a>, experiment with ground truth – if you have 1k images for the 3d model, you can use as bad features, as you want. The COLMAP will recover anyway</p>
<div class="quarto-figure quarto-figure-center">
<figure class="figure">
<p><img src="https://ducha-aiki.github.io/wide-baseline-stereo-blog/posts/2020-11-25-review-of-retrieval-for-localization_files/att_00009.png" class="img-fluid figure-img"></p>
<figcaption>image.png</figcaption>
</figure>
</div>
<p>The retrieval, on the other hand is more interesting to work on, because it is kind of deliberately hard and you can do some fancy stuff, which do not matter in the real world.</p>
</section>
<section id="task1-global-descriptor-only-system-are-quite-useless-now" class="level3">
<h3 class="anchored" data-anchor-id="task1-global-descriptor-only-system-are-quite-useless-now">Task1 (global descriptor-only) system are quite useless now</h3>
<p>Unless we are speaking about the quite dense image representation. I mean, top-accuracy is 35% vs almost 100% for those, which include local features.</p>
<p>Good news: it has a LOT of space for the improvement to work on.</p>
</section>
<section id="for-the-task-2a-and-2b-robust-global-descriptors-are-a-way-to-do-the-retrieval-sorry-vlad." class="level3">
<h3 class="anchored" data-anchor-id="for-the-task-2a-and-2b-robust-global-descriptors-are-a-way-to-do-the-retrieval-sorry-vlad.">For the task 2a and 2b, robust global descriptors are a way to do the retrieval, sorry VLAD.</h3>
<p>The precision will come from the local features. Which I like a lot, because VLAD is more complex to train and initalize, I never liked it (nothing personal).</p>
</section>
<section id="for-the-task2a-and-2b-we-need-new-metrics-e.g.-precisition-x-mb-memory-footprint" class="level3">
<h3 class="anchored" data-anchor-id="for-the-task2a-and-2b-we-need-new-metrics-e.g.-precisition-x-mb-memory-footprint">For the task2a and 2b we need new metrics, e.g.&nbsp;precisition @ X Mb memory footprint</h3>
<p>Because otherwise, the task is easily solved by the brute force – either by photo taking, or, at least with image syntesis, see <a href="https://hal.inria.fr/hal-01616660/document">24/7 place recognition by view synthesis</a>.</p>
<p>Such steps are already taken in the paper <a href="https://arxiv.org/abs/2007.13172">Learning and aggregating deep local descriptors for instance-level recognition</a> – see the table with memory footprint.</p>
<div class="quarto-figure quarto-figure-center">
<figure class="figure">
<p><img src="https://ducha-aiki.github.io/wide-baseline-stereo-blog/posts/2020-11-25-review-of-retrieval-for-localization_files/att_00006.png" class="img-fluid figure-img"></p>
<figcaption>image.png</figcaption>
</figure>
</div>
<p>That is how one could have an interesting research challenge, also having some grounds in the real-world – to work in mobile phones. Otherwise, any method would work, if the database is dense enough.</p>
</section>
<section id="robust-local-features-matter-for-illumination-changes" class="level3">
<h3 class="anchored" data-anchor-id="robust-local-features-matter-for-illumination-changes">Robust local features matter for illumination changes</h3>
<p>It is a bit hidden in the Appendix, so go directly to the Figure 9. It clearly shows that localization performance is bounded by SIFT, if it is used for two view matching, making retrieval improvements irrelevant. When R2D2 or D2Net are used for matching instead, the overall results for night-time are much better.</p>
<div class="quarto-figure quarto-figure-center">
<figure class="figure">
<p><img src="https://ducha-aiki.github.io/wide-baseline-stereo-blog/posts/2020-11-25-review-of-retrieval-for-localization_files/att_00007.png" class="img-fluid figure-img"></p>
<figcaption>image.png</figcaption>
</figure>
</div>
<p>That is in line with my small visual benchmark I did recently.</p>
<p>https://twitter.com/ducha_aiki/status/1330495426865344515</p>
<div class="quarto-figure quarto-figure-center">
<figure class="figure">
<p><img src="https://ducha-aiki.github.io/wide-baseline-stereo-blog/posts/2020-11-25-review-of-retrieval-for-localization_files/att_00008.png" class="img-fluid figure-img"></p>
<figcaption>image.png</figcaption>
</figure>
</div>
<p>That’s all, folks! Now please, check the <a href="https://arxiv.org/abs/2011.11946.pdf">paper</a> and the <a href="https://github.com/naver/kapture-localization">code</a> they provided.</p>


</section>
</section>

<p>Everything you (didn’t) want to know about image matching</p> ]]></description>
  <guid>https://ducha-aiki.github.io/wide-baseline-stereo-blog/posts/2020-11-25-review-of-retrieval-for-localization.html</guid>
  <pubDate>Tue, 24 Nov 2020 23:00:00 GMT</pubDate>
  <media:content url="https://ducha-aiki.github.io/wide-baseline-stereo-blog/posts/images/retrieval-for-loc.png" medium="image" type="image/png" height="64" width="144"/>
</item>
<item>
  <title>Revisiting Brown patch dataset and benchmark</title>
  <dc:creator>Dmytro Mishkin</dc:creator>
  <link>https://ducha-aiki.github.io/wide-baseline-stereo-blog/posts/2020-09-23-local-descriptors-validation.html</link>
  <description><![CDATA[ 





<section id="in-this-post" class="level3">
<h3 class="anchored" data-anchor-id="in-this-post">In this post</h3>
<ol type="1">
<li>Why one needs good development set? What is wrong with existing sets for local patch descriptor learning?</li>
<li>One should validate in the same way, as it is used in production.</li>
<li>Brown patch revisited – implementation details</li>
<li>Local patch descriptors evaluation results.</li>
</ol>
</section>
<section id="really-quick-intro-into-local-patch-descriptors" class="level2">
<h2 class="anchored" data-anchor-id="really-quick-intro-into-local-patch-descriptors">Really quick intro into local patch descriptors</h2>
<p>Local patch descriptor is the thing, which helps you to automatically decide, if two patches in the pair of images correspond to the same point in a real world, or not. It should be robust to illumination, viewpoint and other changes.</p>
<div class="quarto-figure quarto-figure-center">
<figure class="figure">
<p><img src="https://ducha-aiki.github.io/wide-baseline-stereo-blog/posts/2020-09-16-local-descriptors-validation_files/att_00003.png" class="img-fluid figure-img"></p>
<figcaption>The task of local descriptor, neural network here, is to decide if two patches belong to the same point, or nor. Image taken from SoSNet decriptor blogpost by Vassileios Balntas https://medium.com/scape-technologies/mapping-the-world-part-4-sosnet-to-the-rescue-5383671713e7</figcaption>
</figure>
</div>
<p>There are lots of ways how to implement a local patch descriptor: engineered and learned. Local patch descriptor is the crucial component of the <a href="https://ducha-aiki.github.io/wide-baseline-stereo-blog/posts/2020-03-27-intro.html">wide baseline stereo pipeline</a> and a popular computer vision research topic.</p>
</section>
<section id="why-do-you-need-development-set" class="level2">
<h2 class="anchored" data-anchor-id="why-do-you-need-development-set">Why do you need development set?</h2>
<p>Good data is crucial for any machine learning problem – everyone now knows that. One needs high quality training set for training a good model. One also needs good test set, to know, what is <em>real</em> performance. However, there is one more, often forgotten, crucial component – <strong>validation</strong> or <strong>development</strong> set. We use it to decide hyperparameters and validate design choices we make. It should be different from both training and test sets, yet, be good predictor of test set performance. Moreover, it should allow fast iterations, so be not too small.</p>
<p>While such set is commonly called <a href="https://en.wikipedia.org/wiki/Training,_validation,_and_test_sets">validation set</a>, I do like Andrew Ng’s term “<a href="https://cs230.stanford.edu/files/C2M1.pdf">development</a>” set more - because it helps to <em>develop</em> your model.</p>
</section>
<section id="existing-datasets-for-local-patch-descriptors" class="level1">
<h1>Existing datasets for local patch descriptors</h1>
<p>So, what are the development set options for local patch descriptors?</p>
<section id="brown-phototourism." class="level3">
<h3 class="anchored" data-anchor-id="brown-phototourism.">Brown PhotoTourism.</h3>
<div class="quarto-figure quarto-figure-center">
<figure class="figure">
<p><img src="https://ducha-aiki.github.io/wide-baseline-stereo-blog/posts/2020-09-16-local-descriptors-validation_files/brown_patches.png" class="img-fluid figure-img"></p>
<figcaption>Patches from 3 subsets of Brown Phototourism dataset</figcaption>
</figure>
</div>
<p>The most commonly and successfully used dataset for local descriptor learning is PhotoTourism, created in 2008. Here is its <a href="http://matthewalunbrown.com/patchdata/patchdata.html">description by authors</a>:</p>
<blockquote class="blockquote">
<p>The dataset consists of corresponding patches sampled from 3D reconstructions of the Statue of Liberty (New York), Notre Dame (Paris) and Half Dome (Yosemite).</p>
</blockquote>
<p>It also comes with evaluation protocol: patch pairs are labeled as “same” or “different” and the false positive rate at recall of 95% (FPR95) is reported. The variable, used to build ROC curve is descriptor distance between two patches.</p>
<p>Advantages:</p>
<ul>
<li>It contains local patches, extracted for two types of local feature detector – DoG (SIFT) and Harris corners.</li>
<li>It is extracted from images, contraining non-planar structures and the geometrical noise present is caused by the local feature detector, not added artificially.<br>
</li>
<li>Descriptors, trained on the dataset, show very good performance , therefore the data itself is good.</li>
</ul>
<p>Disadvantages:</p>
<ul>
<li>when used as a benchmark, it shows unrealistic results: SIFT is 40x worse than deep learned descriptor. In practice, the difference is much smaller.</li>
</ul>
</section>
<section id="hpatches" class="level3">
<h3 class="anchored" data-anchor-id="hpatches">HPatches</h3>
<p><a href="https://github.com/hpatches/hpatches-dataset">HPatches</a>, where H stands for the “<a href="https://en.wikipedia.org/wiki/Homography">homography</a>” was proposed to overcome the problem of unrealisting metric and, seemingly, too easy data, used in Phototourism dataset.</p>
<p>It was constructed in a different way than a Phototourism. First, local features were detected in the “reference” image and then reprojected to other images in sequences. Reprojection is prossible, because all the images are photographies of the planes – graffity, drawing, print, etc, or are all taken from the same position. After the reprojection, some amount of geometrical noise – rotation, translation, scaling, was added to the local features and the patches were extracted.</p>
<p>This process is illustration on the picture below (both taken from the <a href="https://github.com/hpatches/hpatches-dataset">HPatches website</a>).</p>
<p><img src="https://ducha-aiki.github.io/wide-baseline-stereo-blog/posts/2020-09-16-local-descriptors-validation_files/images_hard.png" class="img-fluid" alt="Visualization of the hard patches locations in the target images."> <img src="https://ducha-aiki.github.io/wide-baseline-stereo-blog/posts/2020-09-16-local-descriptors-validation_files/patches_hard.png" class="img-fluid" alt="Extracted hard patches from the example sequence."></p>
<p>HPAtches also provide 3 testing protocol, evaluating mean average precision (mAP) for 3 different tasks: patch verification – similar to Brown Phototourism, image matching and patch retrieval. The variable, used to build mAP is descriptor distance between two patches.</p>
<p>Advantages:</p>
<ul>
<li>Unlike PhotoTourism patch verification, image matching and patch retrieval tasks are not saturated.</li>
<li>HPatches contains illumination split, allowing the evaluation of descriptor robustness to illumination changes.</li>
</ul>
<p>Disadvantages:</p>
<ul>
<li>patches “misregistration” noise is of artificial nature, although paper claims that it has similar statistics</li>
<li>no non-planar structure</li>
<li>performance in HPaptches does not really correlate with the downstream performance </li>
</ul>
<div class="quarto-figure quarto-figure-center">
<figure class="figure">
<p><img src="https://ducha-aiki.github.io/wide-baseline-stereo-blog/posts/2020-09-16-local-descriptors-validation_files/hpatches_vs_IMC.png" title="Performance in HPaptches does not really correlate with the downstream performance on the Image Matching Benchmark" class="img-fluid figure-img"></p>
<figcaption>image.png</figcaption>
</figure>
</div>
</section>
<section id="amospatches" class="level3">
<h3 class="anchored" data-anchor-id="amospatches">AMOSPatches</h3>
<p><a href="https://github.com/pultarmi/AMOS_patches">AMOS patches</a> is “HPatches illumination on steroids, without geometrical noise”. It has the same advantanges and disadvantages, as HPatches and is mostly focused on illumination and weather changes.</p>
<p><img src="https://ducha-aiki.github.io/wide-baseline-stereo-blog/posts/2020-09-16-local-descriptors-validation_files/att_00002.png" class="img-fluid" alt="Some images, contributing to AMOSPatches"> <img src="https://ducha-aiki.github.io/wide-baseline-stereo-blog/posts/2020-09-16-local-descriptors-validation_files/amos_patches_small.png" class="img-fluid" alt="Some patches from AMOS Patches dataset"></p>
</section>
<section id="photosynth" class="level3">
<h3 class="anchored" data-anchor-id="photosynth">PhotoSynth</h3>
<p><a href="https://github.com/rmitra/PS-Dataset">PhotoSynth</a> can be described and something in the middle between Phototour and HPatches. It contains patches, sampled from planar scenes, as well as from non-planar scenes.</p>
<p>At first glance, it should be great for the test and training purposes. However, there are several issues with it. First, pre-trained HardNetPS descriptor, released together with the dataset, works well on HPatches, but poor in practice.</p>
<p>Second, a couple of colleagues has tried to train the descriptor on top of it, as it was significantly worse than the authors reference model. Moreover, there is no testing/training code protocol available together with dataset.</p>
<p>So, while PhotoSynth might be a good dataset in principle, it definitely needs more love and work.</p>
<div class="quarto-figure quarto-figure-center">
<figure class="figure">
<p><img src="https://ducha-aiki.github.io/wide-baseline-stereo-blog/posts/2020-09-16-local-descriptors-validation_files/att_00000.png" title="Images, contributed to PS dataset" class="img-fluid figure-img"></p>
<figcaption>image.png</figcaption>
</figure>
</div>
<div class="quarto-figure quarto-figure-center">
<figure class="figure">
<p><img src="https://ducha-aiki.github.io/wide-baseline-stereo-blog/posts/2020-09-16-local-descriptors-validation_files/att_00001.png" title="Patches, samples from PS dataset" class="img-fluid figure-img"></p>
<figcaption>image.png</figcaption>
</figure>
</div>
</section>
<section id="designing-the-evaluation-protocol" class="level2">
<h2 class="anchored" data-anchor-id="designing-the-evaluation-protocol">Designing the evaluation protocol</h2>
<p>Classical local descriptor matching consists of two parts: finding nearest neighbors and filtering unreliable ones based on some criterion. I have wrote a <a href="https://medium.com/@ducha.aiki/how-to-match-to-learn-or-not-to-learn-part-2-1ab52ede2022">blogpost, describing the matching strategies in details</a>.</p>
<p>The most used in practice criterion is the first to second nearest neighbor distance (Lowe’s) ratio threshold for filtering false positive matches. It is shown in the figure below.</p>
<p>The intuition is simple: if two candidates are too similar, then the match is unreliable and it is better to drop it.</p>
<div class="quarto-figure quarto-figure-center">
<figure class="figure">
<p><img src="https://ducha-aiki.github.io/wide-baseline-stereo-blog/posts/2020-09-16-local-descriptors-validation_files/att_00004.png" class="img-fluid figure-img"></p>
<figcaption>Second nearest ratio strategy. Features from img1 – blue circles – are matched to features from img2 – red squares. For each point in img1 we calculate two nearest neighbors and check their distance ratio . If both are too similar, i.e.&nbsp;&gt;0.8, bottom at Figure, then the match is discarded. Only confident matches are kept. Right graph is from SIFT paper, justification of such strategy.</figcaption>
</figure>
</div>
<p>Somehow, none of the local patch evaluation protocols does not take such filtering criterion in mind, although it greatly influences the overall performance.</p>
<p>So, let’s do the following:</p>
<ol type="1">
<li>Take the patches, which are extracted from only two images.</li>
<li>For the each patch, calculate the descriptor distance to the correct match and to the hardnest (closest) non-match. Calculate the Lowe’s ratio between this two.</li>
<li>Calculate accuracy for each of such triplets. If the correct match has smaller distance, score 1, if not - 0.</li>
<li>Sort the ratios from smallest to biggest and calculate <a href="https://en.wikipedia.org/wiki/Evaluation_measures_(information_retrieval)#Mean_average_precision">mean average precision</a> (mAP).</li>
</ol>
</section>
<section id="brown-phototour-revisied-implementation-details" class="level2">
<h2 class="anchored" data-anchor-id="brown-phototour-revisied-implementation-details">Brown PhotoTour Revisied: implementation details</h2>
<p>We have designed the protocol, now time for data. We could spend several month collecting and cleaning it…or we can just re-use great Brown PhotoTourism dataset. Re-visiting labeling and/or evaluation protocol of the time-tested dataset is a great idea.</p>
<p>Just couple of examples: <a href="https://github.com/fastai/imagenette">ImageNette</a> created by <a href="https://twitter.com/jeremyphoward">Jeremy Howard</a> from ImageNet, <a href="http://cmp.felk.cvut.cz/revisitop/">Revisited Oxford 5k</a> by <a href="https://filipradenovic.github.io/">Filip Radenovic</a> and so on.</p>
<p>For the protocol we designed above we need the information about the image id, where the patch was extracted from. Unfortunately, there is no such information in the Brown PhotoTourism, but there is suitable alternative – the image id, where the reference patch was detected. What does it mean?</p>
<p>Suppose, we have 4 images and 5 keypoints. All the keypoints present in all images, which gives us 20 patches. 3 keypoints were first detected in Image 1 and 2 in image 2.<br>
That means that we will have 12 patches labeled image 1 and 8 patches labeled image 2.</p>
<p>So, we will have results for image 1 and image 2. Let’s consider image 1. There are 12 patches, splitted in 3 “classes”, 4 patches in each class.</p>
<p>Then, for the each of those 12 patches we:</p>
<ul>
<li>pick each of the corresponding patched as positives, so 3 positives. <img src="https://latex.codecogs.com/png.latex?P_1">, <img src="https://latex.codecogs.com/png.latex?P_2">, <img src="https://latex.codecogs.com/png.latex?P_3"></li>
<li>find the closest negative N.</li>
<li>add triplets (A, <img src="https://latex.codecogs.com/png.latex?P_1">, N), (A, <img src="https://latex.codecogs.com/png.latex?P_2">, N), (A, <img src="https://latex.codecogs.com/png.latex?P_3">, N) to the evaluation.</li>
</ul>
<p>Repeat the same for the image 2. That mimics the two-view matching process as close, as possible, given the data available to us.</p>
</section>
<section id="installation" class="level2">
<h2 class="anchored" data-anchor-id="installation">Installation</h2>
<p><code>pip install brown_phototour_revisited</code></p>
</section>
<section id="how-to-use" class="level2">
<h2 class="anchored" data-anchor-id="how-to-use">How to use</h2>
<p>There is a single function, which does everything for you: <code>full_evaluation</code>. The original Brown benchmark consider evaluation, similar to cross-validation: train descriptor on one subset, evaluate on two others, repeat for all, so 6 evaluations are required. For the handcrafted descriptors, or those, that are trained on 3rd party datasets, only 3 evaluations are necessary. We are following it here as well.</p>
<p>However, if you need to run some tests separately, or reuse some functions – we will cover the usage below. In the following example we will show how to use <code>full_evaluation</code> to evaluate SIFT descriptor as implemented in kornia.</p>
<pre><code># !pip install kornia</code></pre>
<pre><code>import torch
import kornia
from IPython.display import clear_output
from brown_phototour_revisited.benchmarking import *
patch_size = 65 

model = kornia.feature.SIFTDescriptor(patch_size, rootsift=True).eval()

descs_out_dir = 'data/descriptors'
download_dataset_to = 'data/dataset'
results_dir = 'data/mAP'

results_dict = {}
results_dict['Kornia RootSIFT'] = full_evaluation(model,
                                'Kornia RootSIFT',
                                path_to_save_dataset = download_dataset_to,
                                path_to_save_descriptors = descs_out_dir,
                                path_to_save_mAP = results_dir,
                                patch_size = patch_size, 
                                device = torch.device('cuda:0'), 
                           distance='euclidean',
                           backend='pytorch-cuda')
clear_output()
print_results_table(results_dict)</code></pre>
<pre><code>------------------------------------------------------------------------------
Mean Average Precision wrt Lowe SNN ratio criterion on UBC Phototour Revisited
------------------------------------------------------------------------------
trained on       liberty notredame  liberty yosemite  notredame yosemite
tested  on           yosemite           notredame            liberty
------------------------------------------------------------------------------
Kornia RootSIFT        56.70              47.71               48.09 
------------------------------------------------------------------------------</code></pre>
</section>
<section id="results" class="level2">
<h2 class="anchored" data-anchor-id="results">Results</h2>
<p>So, let’s check how it goes. The latest results and implementation are in the following notebooks:</p>
<ul>
<li><a href="https://github.com/ducha-aiki/brown_phototour_revisited/blob/master/examples/evaluate_deep_descriptors.ipynb">Deep descriptors</a></li>
<li><a href="https://github.com/ducha-aiki/brown_phototour_revisited/blob/master/examples/evaluate_non_deep_descriptors.ipynb">Non-deep descriptors</a></li>
</ul>
<p>The results are the following:</p>
<pre><code>------------------------------------------------------------------------------
Mean Average Precision wrt Lowe SNN ratio criterion on UBC Phototour Revisited
------------------------------------------------------------------------------
trained on       liberty notredame  liberty yosemite  notredame yosemite
tested  on           yosemite           notredame            liberty
------------------------------------------------------------------------------
Kornia RootSIFT 32px   58.24              49.07               49.65 
HardNet 32px       70.64  70.31        61.93  59.56        63.06  61.64
SOSNet 32px        70.03  70.19        62.09  59.68        63.16  61.65
TFeat 32px         65.45  65.77        54.99  54.69        56.55  56.24
SoftMargin 32px    69.29  69.20        61.82  58.61        62.37  60.63
HardNetPS 32px         55.56              49.70               49.12 
R2D2_center_grayscal   61.47              53.18               54.98 
R2D2_MeanCenter_gray   62.73              54.10               56.17 
------------------------------------------------------------------------------

------------------------------------------------------------------------------
Mean Average Precision wrt Lowe SNN ratio criterion on UBC Phototour Revisited
------------------------------------------------------------------------------
trained on       liberty notredame  liberty yosemite  notredame yosemite
tested  on           yosemite           notredame            liberty
------------------------------------------------------------------------------
Kornia SIFT 32px       58.47              47.76               48.70 
OpenCV_SIFT 32px       53.16              45.93               46.00 
Kornia RootSIFT 32px   58.24              49.07               49.65 
OpenCV_RootSIFT 32px   53.50              47.16               47.37 
OpenCV_LATCH 65px  -----  -----        -----  37.26        -----  39.08
OpenCV_LUCID 32px      20.37              23.08               27.24 
skimage_BRIEF 65px     52.68              44.82               46.56 
Kornia RootSIFTPCA 3 60.73  60.64        50.80  50.24        52.46  52.02
MKD-concat-lw-32 32p 72.27  71.95        60.88  58.78        60.68  59.10
------------------------------------------------------------------------------</code></pre>
<p>So far - in agreement with IMC benchmark: SIFT and RootSIFT are good, but not the best, SOSNet and HardNet are the leaders, but within tens of percents, not by orders of magnitude.</p>
<div class="quarto-figure quarto-figure-center">
<figure class="figure">
<p><img src="https://ducha-aiki.github.io/wide-baseline-stereo-blog/posts/2020-09-16-local-descriptors-validation_files/att_00005.png" class="img-fluid figure-img"></p>
<figcaption>Image Matching Benchmark results, from https://arxiv.org/abs/2003.01587</figcaption>
</figure>
</div>
<section id="disclaimer-1-dont-trust-this-tables-fully" class="level3">
<h3 class="anchored" data-anchor-id="disclaimer-1-dont-trust-this-tables-fully">Disclaimer 1: don’t trust this tables fully</h3>
<p>I haven’t (yet!) checked if all the deep descriptors models, trained on Brown, were trained with flip-rotation 90 degrees augmentation. In the code below I assume that they were, however, it might not be true – and the comparison might not be completely fair. I will do my best to check it, but if you know that I have used wrong weights - please <a href="https://github.com/ducha-aiki/brown_phototour_revisited/issues">open an issue</a>. Thank you.</p>
</section>
<section id="disclaimer-2-it-is-not-benchmark." class="level3">
<h3 class="anchored" data-anchor-id="disclaimer-2-it-is-not-benchmark.">Disclaimer 2: it is not “benchmark”.</h3>
<p>The intended usage of the package is not to test and report the numbers in the paper. Instead think about is as cross-validation tool, helping the development. Thus, one CAN tune hyperparameters based on the benchmark results instead of doing so on <a href="https://github.com/hpatches/hpatches-benchmark">HPatches</a>. After you have finished tuning, please, evaluate your local descriptors on some downstream task like <a href="https://github.com/vcg-uvic/image-matching-benchmark">IMC image matching benchmark</a> or <a href="https://www.visuallocalization.net/">visual localization</a>.</p>
</section>
</section>
<section id="summary" class="level2">
<h2 class="anchored" data-anchor-id="summary">Summary</h2>
<p>It really pays off, to spend time designing a proper evaluation pipeline and gathering the data for it. If you can re-use existing work - great. But don’t blindly trust anything, even super-popular and widely adopted benchmarks. You need always check if the the protocol and data makes sense for your use-case personally.</p>
<p>Thanks for the reading, see you soon!</p>
</section>
<section id="citation" class="level2">
<h2 class="anchored" data-anchor-id="citation">Citation</h2>
<p>If you use the benchmark/development set in an academic work, please cite it.</p>
<pre><code>@misc{BrownRevisited2020,
  title={UBC PhotoTour Revisied},
  author={Mishkin, Dmytro},
  year={2020},
  url = {https://github.com/ducha-aiki/brown_phototour_revisited}
}</code></pre>
</section>
</section>
<section id="references" class="level1">
<h1>References</h1>
<p>[IMW2020] Jin Yuhe, Mishkin Dmytro, Mishchuk Anastasiia <em>et al.</em>, ``<em>Image Matching across Wide Baselines: From Paper to Practice</em>’’, arXiv preprint arXiv:2003.01587, vol.&nbsp;, number , pp.&nbsp;, 2020.</p>
<p>[pultar2020improving] Pultar Milan, ``<em>Improving the HardNet Descriptor</em>’’, arXiv ePrint:2007.09699, vol.&nbsp;, number , pp.&nbsp;, 2020.</p>


</section>

<p>Everything you (didn’t) want to know about image matching</p> ]]></description>
  <guid>https://ducha-aiki.github.io/wide-baseline-stereo-blog/posts/2020-09-23-local-descriptors-validation.html</guid>
  <pubDate>Tue, 22 Sep 2020 22:00:00 GMT</pubDate>
  <media:content url="https://ducha-aiki.github.io/wide-baseline-stereo-blog/posts/images/brown_phototour_revisited.jpg" medium="image" type="image/jpeg"/>
</item>
<item>
  <title>How to match images taken from really extreme viewpoints?</title>
  <dc:creator>Dmytro Mishkin</dc:creator>
  <link>https://ducha-aiki.github.io/wide-baseline-stereo-blog/posts/2020-08-06-affine-view-synthesis.html</link>
  <description><![CDATA[ 





<section id="in-this-post" class="level3">
<h3 class="anchored" data-anchor-id="in-this-post">In this post</h3>
<ol type="1">
<li>What to do, if you are in a desperate need of matching this particular image pair?</li>
<li>What are the limitations of the affine-covariant detectors like Hessian-Affine or HesAffNet?</li>
<li>ASIFT: brute-force affine view synthesis</li>
<li>Do as little as possible: MODS</li>
<li>What is the key factor of affine view synthesis? Ablation study</li>
</ol>
</section>
<section id="how-to-match-images-taken-from-really-extreme-viewpoints" class="level1">
<h1>How to match images taken from really extreme viewpoints?</h1>
<p>Standard wide-baseline stereo or 3d reconstruction pipelines work well in the many situations. Even if some image pair is not matched, it is usually not a problem. For example, one could match images from very different viewpoints, if there is a sequence of images in between, as shown in Figure below, from “From Single Image Query to Detailed 3D Reconstruction” paper.</p>
<div class="quarto-figure quarto-figure-center">
<figure class="figure">
<p><img src="https://ducha-aiki.github.io/wide-baseline-stereo-blog/posts/2020-08-04-affine-view-synthesis_files/att_00000.png" class="img-fluid figure-img"></p>
<figcaption>One could handle extreme viewpoint changes by proxy images. Figure from From Single Image Query to Detailed 3D Reconstruction</figcaption>
</figure>
</div>
<p>However, that might not always be possible. For example, the number of pictures is limited because they historical and there is no way how one could go and take more without inventing a time machine.</p>
<p>What to do? One way would be to use <a href="https://ducha-aiki.github.io/wide-baseline-stereo-blog/posts/2020-07-17-affine-correspondences.html">affine features</a> like Hessian-AffNet or MSER. However, they help only up to some extent and what if the view, we need to match are more extreme?</p>
<div class="quarto-figure quarto-figure-center">
<figure class="figure">
<p><img src="https://ducha-aiki.github.io/wide-baseline-stereo-blog/posts/2020-08-04-affine-view-synthesis_files/MODS-match-historically.png" class="img-fluid figure-img"></p>
<figcaption>Non-matchable by standard methods pair of historical photographies. Matched only with help of affine view synthesis in the MODS framework. Images from Location recognition over large time lags dataset</figcaption>
</figure>
</div>
<p>The image pair above is from “Location recognition over large time lags dataset” paper .</p>
<p>The solution is to simulate real viewpoint change by affine or perspective warps of the current image. This idea was first proposed by Lepetit and Fua in 2006. You can think about it as a special version of test-time augmentation, popular nowadays in deep learning. Later affine view synthesis for wide baseline stereo was extended and mathematically justified by Morel &amp; Yu in ASIFT paper. They proved that perspective image warps are can be approximated by synthetic affine views.</p>
<section id="what-is-wrong-with-affine-covariant-local-detectors" class="level2">
<h2 class="anchored" data-anchor-id="what-is-wrong-with-affine-covariant-local-detectors">What is wrong with affine-covariant local detectors?</h2>
<p>One could say that the goal of affine-covariant detectors like MSER, Hessian-Affine or Hessian-AffNet is to detect the same region on a planar surface, regardless the camera angle change. It is true to some extent, as we demostrate on toy example below with Hessian-Affine feature.</p>
<div class="quarto-figure quarto-figure-center">
<figure class="figure">
<p><img src="https://ducha-aiki.github.io/wide-baseline-stereo-blog/posts/2020-08-04-affine-view-synthesis_files/together_single.png" class="img-fluid figure-img"></p>
<figcaption>Hessian-Affine detector detects the same support region under synthetic tilt</figcaption>
</figure>
</div>
<p>The problem arises, when the image content, e.g.&nbsp;3 blobs on the figure below are situated close to each other, so under the tilt transform the merge into a single blob. So it is not the shape of region, which is detected incorrectly, but the center of the features themselves. For clarity, we omited affine shape estimation on the image below.</p>
<div class="quarto-figure quarto-figure-center">
<figure class="figure">
<p><img src="https://ducha-aiki.github.io/wide-baseline-stereo-blog/posts/2020-08-04-affine-view-synthesis_files/blobs_merging.png" class="img-fluid figure-img"></p>
<figcaption>When blobs are close to each other, they merge together when picture is taked from a side viewpoint. This results in failure of the blob detector.</figcaption>
</figure>
</div>
</section>
<section id="asift-brute-force-affine-view-synthesis" class="level2">
<h2 class="anchored" data-anchor-id="asift-brute-force-affine-view-synthesis">ASIFT: brute-force affine view synthesis</h2>
<p>So, to solve the problem explained above, Morel &amp; Yu proposed to do a lot affine warps of each image, as shown on the Figure below, as match each view against all others, which is <img src="https://latex.codecogs.com/png.latex?O(n%5E2)"> complexity, where <img src="https://latex.codecogs.com/png.latex?n"> is number of views generated.</p>
<div class="quarto-figure quarto-figure-center">
<figure class="figure">
<p><img src="https://ducha-aiki.github.io/wide-baseline-stereo-blog/posts/2020-08-04-affine-view-synthesis_files/att_00006.png" class="img-fluid figure-img"></p>
<figcaption>ASIFT algorithm: generate a lot of synthetic views, match all to all. Figure from Fast Affine Invariant Image Matching</figcaption>
</figure>
</div>
<p>The motivation do doing so it that assuming, original image to be a fronto-parallel one, to cover viewsphere really dense, as shown in the Figure below.</p>
<div class="quarto-figure quarto-figure-center">
<figure class="figure">
<p><img src="https://ducha-aiki.github.io/wide-baseline-stereo-blog/posts/2020-08-04-affine-view-synthesis_files/att_00007.png" class="img-fluid figure-img"></p>
<figcaption>Viewsphere covering by ASIFT. Figure from Fast Affine Invariant Image Matching</figcaption>
</figure>
</div>
<p>This leads to impressive performance on a very challenging image pairs, see an example below</p>
<div class="quarto-figure quarto-figure-center">
<figure class="figure">
<p><img src="https://ducha-aiki.github.io/wide-baseline-stereo-blog/posts/2020-08-04-affine-view-synthesis_files/att_00009.png" title="ASIFT find a lot of correspondences on challenging pair. Figure from ASIFT: An Algorithm for Fully Affine Invariant Comparison" class="img-fluid figure-img"></p>
<figcaption>image.png</figcaption>
</figure>
</div>
<p>In this section I have used great illustrations done by <a href="https://rdguez-mariano.github.io/">Mariano Rodríguez</a> for his paper “Fast Affine Invariant Image Matching” . Please, checkout his <a href="https://rdguez-mariano.github.io/">blog</a>.</p>
</section>
<section id="mods-do-as-little-as-possible" class="level2">
<h2 class="anchored" data-anchor-id="mods-do-as-little-as-possible">MODS: do as little as possible</h2>
<p>The main drawback of ASIFT algorithm is a huge computational cost: 82 views are generated regardless of the image pair difficulty. To overcome this, we proposed MODS algorithm: Matching with On-Demand Synthesis.</p>
<div class="quarto-figure quarto-figure-center">
<figure class="figure">
<p><img src="https://ducha-aiki.github.io/wide-baseline-stereo-blog/posts/2020-08-04-affine-view-synthesis_files/att_00004.png" class="img-fluid figure-img"></p>
<figcaption>MODS algorithm: synthetize more views until match</figcaption>
</figure>
</div>
<p>One starts with the fastest detector-descriptor without view synthesys and then uses more and more computationally expensive methods if needed. Moreover, by using affine-covariant detectors like MSER or Hessian-Affine, one could synthetise significantly less views, saving computations spent on local descriptor and matching.</p>
<p>This, together with <a href="https://medium.com/@ducha.aiki/how-to-match-to-learn-or-not-to-learn-part-2-1ab52ede2022">FGINN matching strategy</a>, specifically designed for the handling re-detections, MODS is able to match more challenging image pairs in less time than ASIFT.</p>
<div class="quarto-figure quarto-figure-center">
<figure class="figure">
<p><img src="https://ducha-aiki.github.io/wide-baseline-stereo-blog/posts/2020-08-04-affine-view-synthesis_files/att_00010.png" class="img-fluid figure-img"></p>
<figcaption>MODS outperforms ASIFT on a several dataset both in terms of speed and quality</figcaption>
</figure>
</div>
</section>
<section id="why-does-affine-synthesis-help" class="level2">
<h2 class="anchored" data-anchor-id="why-does-affine-synthesis-help">Why does affine synthesis help?</h2>
<p>Despite that ASIFT and other view-synthesis based approaches are know more than decade, we are not aware of a study, why does affine synthesis helps in practice. Could one get a similar performance without view synthesis? Specificallly:</p>
<ol start="0" type="1">
<li>May it be that the most of improvements come from the fact that we have much more features? That is why we fix the number of features for all approaches.</li>
<li>Some regions from ASIFT, when reprojected to the original image, are quite narrow. Could we get them just by removing edge-like feature filtering, which is done in SIFT, Hessian and other detectors. Denoted <strong>+edge</strong></li>
<li>Instead of doing affine view synthesis, one could directly use the same affine parameters to get the affine regions to describe, so the each keypoint would have several associated regions+descriptors. Denoted <strong>+MD</strong></li>
<li>Using AffNet to directly estimated local affine shape without multiple descriptors. Denoted <strong>+AffNet</strong></li>
<li>Combine (1), (2) and (3).</li>
</ol>
<p>So, we did the study on HPatches Sequences dataset, the hardest image pairs (1-6) of viewpoint subset. The metric is similar to one used in the “<a href="https://arxiv.org/abs/2003.01587">Image Matching across Wide Baselines: From Paper to Practice</a>” and CVPR 2020 <a href="http://cmp.felk.cvut.cz/cvpr2020-ransac-tutorial/">RANSAC in 2020</a> - mean average accuracy of the estimated homography.</p>
<div class="quarto-figure quarto-figure-center">
<figure class="figure">
<p><img src="https://ducha-aiki.github.io/wide-baseline-stereo-blog/posts/2020-08-04-affine-view-synthesis_files/att_00011.png" class="img-fluid figure-img"></p>
<figcaption>mean average accuracy of the estimated homography - used metric</figcaption>
</figure>
</div>
<p>We run Hessian detector with RootSIFT descriptor, FLANN matching and LO-RANSAC, as implemented in <a href="https://github.com/ducha-aiki/mods-light-zmq">MODS</a>. Features are sorted according the the detector response and their total number is clipped to 2048 or 8000 to ensure that the improvements do not come from just having more features.</p>
<p>Note, that we do not study, if view synthesis helps for the regular image pairs - it might actually hurt performance, similarly to <a href="https://ducha-aiki.github.io/wide-baseline-stereo-blog/posts/2020-07-17-affine-correspondences.html">affine features</a>. Instead we are focusing on the case, when view synthesis definitely helps: matching obscure views of the mostly planar scenes.</p>
<section id="feature-budget" class="level3">
<h3 class="anchored" data-anchor-id="feature-budget">8000 feature budget</h3>
<p>Results are in Figure below. Indeed, all of the factors: detecting more edge-like features, having multiple descriptors or better affine shape improve results over the plain Hessian detector, but even all of the combined are not good enough to match performance of the affine view synthesis + plain Hessian detector.</p>
<p>But the best setup is to use both Hessian-AffNet and view synthesis.</p>
<div class="quarto-figure quarto-figure-center">
<figure class="figure">
<p><img src="https://ducha-aiki.github.io/wide-baseline-stereo-blog/posts/2020-08-04-affine-view-synthesis_files/8k_budget.png" class="img-fluid figure-img"></p>
<figcaption>Ablation study of the different benefit from the affine view synthesis: ability to detect elongated blobs, multiple affine decriptors or everything together. Feature budget: 8k</figcaption>
</figure>
</div>
</section>
<section id="feature-budget-1" class="level3">
<h3 class="anchored" data-anchor-id="feature-budget-1">2048 feature budget</h3>
<p>The picture is a bit different in a small feature budget: neither multiple-(affine)-descriptors per keypoint, nor allowing edge-like feature help. From other hand, affine view synthesis still improves results of the Hessian. And, again, the best performance is achieved with combination of view synthesis and AffNet shape estimation.</p>
<div class="quarto-figure quarto-figure-center">
<figure class="figure">
<p><img src="https://ducha-aiki.github.io/wide-baseline-stereo-blog/posts/2020-08-04-affine-view-synthesis_files/2k_budget.png" class="img-fluid figure-img"></p>
<figcaption>Ablation study of the different benefit from the affine view synthesis: ability to detect elongated blobs, multiple affine decriptors or everything together. Feature budget: 8k</figcaption>
</figure>
</div>
</section>
</section>
<section id="summary" class="level2">
<h2 class="anchored" data-anchor-id="summary">Summary</h2>
<p>Affine view synthesis helps for matching challenging image pairs and its improvement are not just because of more local features used. It can be done effective and efficient – in the iterative MODS framework.</p>
</section>
</section>
<section id="references" class="level1">
<h1>References</h1>
<p>[SingleImage3dRec2015] J.L. Schonberger, F. Radenovic, O. Chum <em>et al.</em>, ``<em>From Single Image Query to Detailed 3D Reconstruction</em>’’, Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), June 2015.</p>
<p>[AffNet2018] D. Mishkin, F. Radenovic and J. Matas, ``<em>Repeatability is Not Enough: Learning Affine Regions via Discriminability</em>’’, ECCV, 2018.</p>
<p>[MSER2002] J. Matas, O. Chum, M. Urban <em>et al.</em>, ``<em>Robust Wide Baseline Stereo from Maximally Stable Extrema Regions</em>’’, BMVC, 2002.</p>
<p>[LostInPast2015] Fernando Basura, Tommasi Tatiana and Tuytelaars Tinne, ``<em>Location recognition over large time lags</em>’’, Computer Vision and Image Understanding, vol.&nbsp;139, number , pp.&nbsp;, 2015.</p>
<p>[AffineTree2006] Lepetit Vincent and Fua Pascal, ``<em>Keypoint Recognition Using Randomized Trees</em>’’, IEEE Trans. Pattern Anal. Mach. Intell., vol.&nbsp;28, number 9, pp.&nbsp;, sep 2006.</p>
<p>[ASIFT2009] Morel Jean-Michel and Yu Guoshen, ``<em>ASIFT: A New Framework for Fully Affine Invariant Image Comparison</em>’’, SIAM J. Img. Sci., vol.&nbsp;2, number 2, pp.&nbsp;, apr 2009.</p>
<p>[FastASIFT2018] Rodríguez Mariano, Delon Julie and Morel Jean-Michel, ``<em>Fast Affine Invariant Image Matching</em>’’, Image Processing On Line, vol.&nbsp;8, number , pp.&nbsp;, 2018.</p>
<p>[MODS2015] Mishkin Dmytro, Matas Jiri and Perdoch Michal, ``<em>MODS: Fast and robust method for two-view matching </em>’’, Computer Vision and Image Understanding , vol.&nbsp;, number , pp.&nbsp;, 2015.</p>


</section>

<p>Everything you (didn’t) want to know about image matching</p> ]]></description>
  <guid>https://ducha-aiki.github.io/wide-baseline-stereo-blog/posts/2020-08-06-affine-view-synthesis.html</guid>
  <pubDate>Wed, 05 Aug 2020 22:00:00 GMT</pubDate>
  <media:content url="https://ducha-aiki.github.io/wide-baseline-stereo-blog/posts/images/transition.png" medium="image" type="image/png" height="58" width="144"/>
</item>
<item>
  <title>Patch extraction: devil in details</title>
  <dc:creator>Dmytro Mishkin</dc:creator>
  <link>https://ducha-aiki.github.io/wide-baseline-stereo-blog/posts/2020-07-22-patch-extraction.html</link>
  <description><![CDATA[ 





<p>When working with local features one needs to pay attention to even a smallest details, or the whole process can be ruined. One of such details is how to extract the patch, which will be described by local descriptor such as SIFT or HardNet.</p>
<div class="quarto-figure quarto-figure-center">
<figure class="figure">
<p><img src="https://ducha-aiki.github.io/wide-baseline-stereo-blog/posts/2020-07-22-patch-extraction_files/att_00000.png" class="img-fluid figure-img"></p>
<figcaption>In order to describe the local patch one first has to extract it. Figure from A Few Things One Should Know About Feature Extraction, Description and Matching</figcaption>
</figure>
</div>
<p>Unfortunately, we cannot just extract patch from the image by cropping the patch and then resizing it. Or can we? Let’s check. We will use two versions of image: original and 4x smaller one and would like to extract same-looking fixed size patch from both of them. The patch we want to crop is showed by oriented red circle.</p>
<div class="quarto-figure quarto-figure-center">
<figure class="figure">
<p><img src="https://ducha-aiki.github.io/wide-baseline-stereo-blog/posts/2020-07-22-patch-extraction_files/both_imgs.png" class="img-fluid figure-img"></p>
<figcaption>Two versions of the same image: original and 4x smaller</figcaption>
</figure>
</div>
<section id="aliasing" class="level2">
<h2 class="anchored" data-anchor-id="aliasing">Aliasing</h2>
<p>And here what we get by doing a simple crop and resize to 32x32 pixels.</p>
<div class="quarto-figure quarto-figure-center">
<figure class="figure">
<p><img src="https://ducha-aiki.github.io/wide-baseline-stereo-blog/posts/2020-07-22-patch-extraction_files/naive_patches.png" class="img-fluid figure-img"></p>
<figcaption>Patches, which are extracted from images of different sizes look different, if extraction done in a naive way</figcaption>
</figure>
</div>
<p>Doesn’t look good. It is called “<a href="https://en.wikipedia.org/wiki/Aliasing">aliasing</a>” - a problem, which arise when we are trying to downscale big images into small resolution. Specifically: the original image contains finer details, than we could represent in thumbnail, which leads to artifacts. <img src="https://ducha-aiki.github.io/wide-baseline-stereo-blog/posts/2020-07-22-patch-extraction_files/ep.jpg" class="img-fluid" alt="One does not simply extract patch from the image"></p>
</section>
<section id="the-solution-anti-aliasing" class="level2">
<h2 class="anchored" data-anchor-id="the-solution-anti-aliasing">The solution: anti-aliasing</h2>
<p>The solution, which follows out of <a href="https://en.wikipedia.org/wiki/Nyquist%E2%80%93Shannon_sampling_theorem">sampling theorem</a> is known: remove the details, which cannot be seens in small image first, then resample image to small size.</p>
<p>The simplest way to remove the fine details is to blur image with the Gaussian kernel.</p>
<p>Lets do it and compare the results. <img src="https://ducha-aiki.github.io/wide-baseline-stereo-blog/posts/2020-07-22-patch-extraction_files/all_patches.png" class="img-fluid" alt="When extracted properly, patches look same"></p>
<p>By the way, you can try for yourself, all the required code is <a href="https://github.com/kornia/kornia-examples/blob/master/aliased-and-not-aliased-patch-extraction.ipynb">here, in kornia-examples</a></p>
</section>
<section id="performance" class="level2">
<h2 class="anchored" data-anchor-id="performance">Performance</h2>
<p>The problem is solved. Or is it?</p>
<p>The problem with properly antialiased patch extraction is that it is quite slow for two reasons. First, blurring a whole image is a costly operation. But, the worst part is that the required amount of blur depends on the patch size in the original image, or, in other words, keypoint scale. So for extracting, say 8000 patches, one needs to perform blurring 8000 times. Moreover, if one wants to extract elongated region and warp it to the square patch, the amount of blur in vertical and horizontal directions should be different!</p>
<p>What can be done? Well, instead of doing blurring 8000 times, one could create so called scale pyramid and then pick the level, which is the closest to optimal one, predicted by theorem.</p>
<div class="quarto-figure quarto-figure-center">
<figure class="figure">
<p><img src="https://ducha-aiki.github.io/wide-baseline-stereo-blog/posts/2020-07-22-patch-extraction_files/Image_pyramid.png" class="img-fluid figure-img"></p>
<figcaption>Image from Wikimedia by wiki-user Cmglee https://en.wikipedia.org/wiki/File:Image_pyramid.svg</figcaption>
</figure>
</div>
<p>This is exactly, what kornia function <a href="https://kornia.readthedocs.io/en/latest/feature.html#kornia.feature.extract_patches_from_pyramid">extract_patches_from_pyramid</a> does.</p>
<p>Also - I have a bit cheated with you above: the “anti-aliased” patches were actually extracted using the function above.</p>
</section>
<section id="how-it-impacts-local-descriptor-matching" class="level2">
<h2 class="anchored" data-anchor-id="how-it-impacts-local-descriptor-matching">How it impacts local descriptor matching?</h2>
<p>Let’s do the toy example first - describe four patches we have in the example above with HardNet descriptor and calculate the distance between them.</p>
<div class="quarto-figure quarto-figure-center">
<figure class="figure">
<p><img src="https://ducha-aiki.github.io/wide-baseline-stereo-blog/posts/2020-07-22-patch-extraction_files/att_00001.png" class="img-fluid figure-img"></p>
<figcaption>Descriptor difference between antialiased patches is 0.09 and between naively extracted – 0.44</figcaption>
</figure>
</div>
<p>So the descriptor difference between antialiased patches is 0.09 and between naively extracted – 0.44. 0.09 is not a big deal, but 0.44 is a lot, actually.</p>
<p>Let’s move to the non-toy example from the paper devoted to this topic: “<a href="http://cmp.felk.cvut.cz/~mishkdmy/lenc-2014-features-cvww.pdf">A Few Things One Should Know About Feature Extraction, Description and Matching</a>”.</p>
<p>The original data is lost I am too lazy to redo the experiments for the post, so I will just copy-past images with results. Here are abbrevations used in the paper:</p>
<ul>
<li><p><strong>OPE</strong> – Optimal Patch Extraction. The most correct and slow way of extracting, including different amount of bluring in different directions.</p></li>
<li><p><strong>NBPE</strong> – No-Blur Patch Extraction. The most naive way we started with</p></li>
<li><p><strong>PNBPE</strong> – Pyramid, No-Blur Patch Extraction. The one, we described above - sampling patches from scale pyramid.</p></li>
<li><p><strong>PSPE</strong> Pyramid-Smoothing Patch Extraction. Pick the matching pyramid level and then add anisotropic blur missing.</p></li>
</ul>
<p>As you can see, doing things optimally is quite slow. <img src="https://ducha-aiki.github.io/wide-baseline-stereo-blog/posts/2020-07-22-patch-extraction_files/att_00003.png" title="Time spent on a various stages of local feature extraction" class="img-fluid" alt="image.png"></p>
<p>Now let’s see how it influences performance.</p>
<div class="quarto-figure quarto-figure-center">
<figure class="figure">
<p><img src="https://ducha-aiki.github.io/wide-baseline-stereo-blog/posts/2020-07-22-patch-extraction_files/att_00005.png" class="img-fluid figure-img"></p>
<figcaption>Number of matches for MSER and Hessian-Affine + SIFT on Grffity sequence of Oxford-Affine dataset for different patch extraction methods</figcaption>
</figure>
</div>
<p>It looks like that influence is smaller than we thought. But recall that the experiment above is for SIFT descriptor only. Doing pyramid helps for the small viewpoint change almost as good, as going fully optimal, but with increasing the viewpoint difference, such approximation degrades. Moreover, it influnces MSER detector much more that than Hessian-Affine.</p>
<p>How does it work with deep descriptors like HardNet or SoSNet? That is the question which not answered yet. Drop me a message if you want to do it yourself and we can do the follow-up post together.</p>
</section>
<section id="references" class="level1">
<h1>References</h1>
<p>[PatchExtraction2014] K. Lenc, J. Matas and D. Mishkin, ``<em>A few things one should know about feature extraction, description and matching</em>’’, Proceedings of the Computer Vision Winter Workshop, 2014.</p>


</section>

<p>Everything you (didn’t) want to know about image matching</p> ]]></description>
  <guid>https://ducha-aiki.github.io/wide-baseline-stereo-blog/posts/2020-07-22-patch-extraction.html</guid>
  <pubDate>Tue, 21 Jul 2020 22:00:00 GMT</pubDate>
  <media:content url="https://ducha-aiki.github.io/wide-baseline-stereo-blog/posts/images/ep.jpg" medium="image" type="image/jpeg"/>
</item>
<item>
  <title>Local affine features: useful side product</title>
  <dc:creator>Dmytro Mishkin</dc:creator>
  <link>https://ducha-aiki.github.io/wide-baseline-stereo-blog/posts/2020-07-17-affine-correspondences.html</link>
  <description><![CDATA[ 





<section id="keypoints-are-not-just-points" class="level2">
<h2 class="anchored" data-anchor-id="keypoints-are-not-just-points">Keypoints are not just points</h2>
<div class="quarto-figure quarto-figure-center">
<figure class="figure">
<p><img src="https://ducha-aiki.github.io/wide-baseline-stereo-blog/posts/2020-07-17-affine-correspondences_files/att_00000.png" class="img-fluid figure-img"></p>
<figcaption>image.png</figcaption>
</figure>
</div>
<p>Wide baseline stereo matching often as perceived as establishing (key-)point correspondences between images. While this might be true for the some local features like SuperPoint , typically it is more than that.</p>
<p>Specifically, detectors like DoG, Harris, Hessian, KeyNet, ORB, and many others rate on scale-space provide at least 3 parameters: x, y, and scale.</p>
<p>Most of the local descriptors – SIFT, HardNet and so on – are not rotation invariant and those which are - mostly require complex matching function, , so the patch orientation has to be estimated anyway, in order to match reliably. This can be done by various methods: corners center of mass (ORB, dominant gradient orientation (SIFT) or by some learned estimator (OriNets,). Sometimes it is possible to rely on smartphone camera IMU or photographer and assume that images are upright.</p>
<p>Thus, we can assume that if local descriptors match, this means the local feature scale and orientation also match, at least approximately – see Figure below. Possible exceptions are cases, when the patch is symmetrical and orientation is ambiguous up to some symmetry group.</p>
<div class="quarto-figure quarto-figure-center">
<figure class="figure">
<p><img src="https://ducha-aiki.github.io/wide-baseline-stereo-blog/posts/2020-07-17-affine-correspondences_files/matches_patches.png" class="img-fluid figure-img"></p>
<figcaption>Selected matching SIFT keypoints and corresponding patches. One can see that not only patch centers correspond to each other, but also other pixels, although less precise. Image pair from Sacre Coeur IMW dataset.</figcaption>
</figure>
</div>
<p>In addition, one could assume that we observe the patch not from the fronto-parallel position and try to estimate local normal, or, more precisely, affine shape of the feature point, modeling it as an ellipse instead of circle. One could also think of affine shape estimation as finding the camera position, from where the patch is seen in some “canonical” view.</p>
<div class="quarto-figure quarto-figure-center">
<figure class="figure">
<p><img src="https://ducha-aiki.github.io/wide-baseline-stereo-blog/posts/2020-07-17-affine-correspondences_files/affinematches_patches.png" class="img-fluid figure-img"></p>
<figcaption>Selected matching SIFT-AffNet keypoints and corresponding patches. One can see that not only patch centers correspond to each other, but also other pixels, although less precise. Image pair from Sacre Coeur IMW dataset.</figcaption>
</figure>
</div>
<p>This gives us 3 points correspondences from a single local feature match, see an example in Figure below.</p>
<div class="quarto-figure quarto-figure-center">
<figure class="figure">
<p><img src="https://ducha-aiki.github.io/wide-baseline-stereo-blog/posts/2020-07-17-affine-correspondences_files/laf-check-illustration.png" class="img-fluid figure-img"></p>
<figcaption>Local affine correspondences. While centers of both regions A and B are correct point matches, only A is a correct affine correspondence.</figcaption>
</figure>
</div>
<p>Why is it important and how to use it – see in current post. How to esimate local affine features robustly – in the next post.</p>
</section>
<section id="benefits-of-local-affine-features" class="level1">
<h1>Benefits of local affine features</h1>
<section id="making-descriptor-job-easier" class="level2">
<h2 class="anchored" data-anchor-id="making-descriptor-job-easier">Making descriptor job easier</h2>
<p>The most straightforward benefit of using local affine features is that they increase the repeatability of the detector and potentially reduce appearance changes of a local patch caused by viewpoint difference. This makes possible matching more challenging image pairs.</p>
<div class="quarto-figure quarto-figure-center">
<figure class="figure">
<p><img src="https://ducha-aiki.github.io/wide-baseline-stereo-blog/posts/2020-07-17-affine-correspondences_files/test.png" class="img-fluid figure-img"></p>
<figcaption>Hessian features + HardNet matches + RANSAC inliers, Right: HessAffNet features + HardNet matches + RANSAC inliers. Image pair from Tanks &amp; Temples. Epipolar lines are shown in cyan.</figcaption>
</figure>
</div>
<div class="quarto-figure quarto-figure-center">
<figure class="figure">
<p><img src="https://ducha-aiki.github.io/wide-baseline-stereo-blog/posts/2020-07-17-affine-correspondences_files/att_00009.png" class="img-fluid figure-img"></p>
<figcaption>Repeatability and the number of correspondences. AffNet compared with the de facto standard Baumberg iteration according to the Mikolajczyk protocol. Left – images with illumination differences, right – with viewpoint and scale changes. SS – patch from the scale-space pyramid at the level of the detection, image – from the original image; 19 and 33 are patch sizes.</figcaption>
</figure>
</div>
<p>The practice is a little bit more complicated. Our recent benchmark, which measure the accuracy of the output fundamental matrix, shows that the difference in using affine and similarity-covariant features is quite minor. Specifically, the relative difference between SIFT vs SIFT-Affine features is 0.5% and between Hessian-SIFT and Hessian-AffNet SIFT is 5.1%, see Table below.</p>
<div class="quarto-figure quarto-figure-center">
<figure class="figure">
<p><img src="https://ducha-aiki.github.io/wide-baseline-stereo-blog/posts/2020-07-17-affine-correspondences_files/att_00011.png" class="img-fluid figure-img"></p>
<figcaption>Test – Stereo results with 8k features. We report: NF – Number of Features; NI – Number of Inliers produced by RANSAC; and mAA@10°.</figcaption>
</figure>
</div>
<p>Therefore, if the benefit of local features would be to only improve descriptor extraction stage, it would be arguably not worth it. Luckily, there are more benefits, which are more pronounced.</p>
</section>
<section id="making-ransac-job-easier" class="level2">
<h2 class="anchored" data-anchor-id="making-ransac-job-easier">Making RANSAC job easier</h2>
<p>Let’s recall how RANSAC works.</p>
<ol type="1">
<li>Randomly sample a minimally required number of tentative correspondences to fit the geometrical model of the scene: 4 for homography, 7 for epipolar geometry and estimate the model.</li>
<li>Calculate “support”: other correspondeces, which are consistent with the model.</li>
<li>Repeat steps (1), (2) and output the model which is supported with the most of correspondences. If you were lucky and have sampled all-inlier sample, meaning that all correspondences used to estimate the model were correct, you would have a correct model.</li>
</ol>
<p>Reality is more complicated than I have just described, but the principle is the same. The most important part is the sampling and it is sensitive to inlier ratio <img src="https://latex.codecogs.com/png.latex?%5Cnu"> - the percentage of the correct correspondences in the set. Lets denote the minimal number of correspondences required to estimate the model as <strong>m</strong>. To recover the correct model with the confidence <strong>p</strong> one needs to sample the number of correspondences, which is described by formula:</p>
<p><img src="https://latex.codecogs.com/png.latex?%5Cbegin%7Bequation%7D%0AN%20=%20%5Cfrac%7B%5Clog%7B(1%20-%20p)%7D%7D%7B%5Clog%7B(1%20-%20%5Cnu%5E%7Bm%7D)%7D%7D%0A%5Cend%7Bequation%7D"></p>
<p>Lets plot the how the number of required samples changes with inlier ratio for confidence equal 99%. Note the log scale on Y axis. Different lines are for different minimal sample size <strong>m</strong>.</p>
<div class="quarto-figure quarto-figure-center">
<figure class="figure">
<p><img src="https://ducha-aiki.github.io/wide-baseline-stereo-blog/posts/2020-07-17-affine-correspondences_files/att_00016.png" class="img-fluid figure-img"></p>
<figcaption>Number of samples to find correct model as a function of inlier ratio.</figcaption>
</figure>
</div>
<p>As you can see from the plot above, reducing the minimal sample size required for the model estimation even by 1 saves and order of magnitude of computation. In reality the benefit is a smaller, as modern RANSACs like GC-RANSAC and MAGSAC could estimate the correct model from the sample containing outliers, but it is still huge, especially for low inlier rate cases.</p>
<section id="image-retrieval" class="level3">
<h3 class="anchored" data-anchor-id="image-retrieval">Image retrieval</h3>
<p>The ideal case would be to estimate a model from just a single sample and that is exactly what is done in spatial reranking paper “<strong>Object retrieval with large vocabularies and fast spatial matching</strong>” by Philbin et.al .</p>
<p>Specifically, they are solving a particular object retrieval problem: given an image containing some object, return all the images from the database, which also containg the same object.</p>
<p>The inital list of images is formed by the descriptor distance and then is reranked. The authors propose to approximate a perspective change between two images as an affine image transformation, and count number of feature points, which are reprojected inside the second image. This number produces better ranking that the original short-list.</p>
<div class="quarto-figure quarto-figure-center">
<figure class="figure">
<p><img src="https://ducha-aiki.github.io/wide-baseline-stereo-blog/posts/2020-07-17-affine-correspondences_files/att_00012.png" class="img-fluid figure-img"></p>
<figcaption>Figure from Philbin et.al</figcaption>
</figure>
</div>
</section>
<section id="back-to-wide-baseline-stereo" class="level3">
<h3 class="anchored" data-anchor-id="back-to-wide-baseline-stereo">Back to wide baseline stereo</h3>
<p>While working for spatial re-ranking, 3-degrees of freedom camera model is too rough for the wide baseline stereo. Yet, going from 4 point correspondences (PC) to 2 affine correspondences (AC) for homogaphy and from 7 PC to 3 AC for the fundamental matrix would be huge benefit anyway for the robust model estimation.</p>
<p>Various variant of RANSAC working for local features were proposed in the last 15 years: Perdoch et.al, Pritts et.al., Barath and Kukelova , Rodríguez et.al.</p>
<p>Finally, the systematic study of using is presented by Barath et.al in “Making Affine Correspondences Work in Camera Geometry Computation” paper. Authors show that if used naively, affine correspondence lead to worse results, because they are more noisy than point correspondences. However, there is a bag of tricks presented in the paper, which allow to solve the noise issue and make the affine RANSAC working in practice, resulting in orders of magnitude faster computation.</p>
<div class="quarto-figure quarto-figure-center">
<figure class="figure">
<p><img src="https://ducha-aiki.github.io/wide-baseline-stereo-blog/posts/2020-07-17-affine-correspondences_files/att_00013.png" class="img-fluid figure-img"></p>
<figcaption>Figure from Making Affine Correspondences Work in Camera Geometry Computation</figcaption>
</figure>
</div>
<p>Moreover, for a special cases like autonomous driving, where the motion is mostly horizonal, one could even use 2 affine correspondes for both motion estimation and consistency check, significantly improving the efficiency of the outliers removal compared to the standard RANSAC loop.</p>
<div class="quarto-figure quarto-figure-center">
<figure class="figure">
<p><img src="https://ducha-aiki.github.io/wide-baseline-stereo-blog/posts/2020-07-17-affine-correspondences_files/att_00001.png" class="img-fluid figure-img"></p>
<figcaption>Empirical cumulative error distributions for KITTI sequence 00. Figure from Relative Pose Estimation for Multi-Camera Systems from Affine Correspondences.</figcaption>
</figure>
</div>
<p>Besides the special case considerations, additional contraints can also come from running other algorithms, like monocular depth estimation. Such a constraint could reduce the required number of matches from two affine correspondences to a single one for calibrated camera case.</p>
<div class="quarto-figure quarto-figure-center">
<figure class="figure">
<p><img src="https://ducha-aiki.github.io/wide-baseline-stereo-blog/posts/2020-07-17-affine-correspondences_files/att_00002.png" class="img-fluid figure-img"></p>
<figcaption>Figure from Relative Pose from Deep Learned Depth and a Single Affine Correspondence</figcaption>
</figure>
</div>
<div class="quarto-figure quarto-figure-center">
<figure class="figure">
<p><img src="https://ducha-aiki.github.io/wide-baseline-stereo-blog/posts/2020-07-17-affine-correspondences_files/att_00003.png" class="img-fluid figure-img"></p>
<figcaption>Figure from Relative Pose from Deep Learned Depth and a Single Affine Correspondence</figcaption>
</figure>
</div>
</section>
</section>
<section id="application-specific-benefits" class="level2">
<h2 class="anchored" data-anchor-id="application-specific-benefits">Application-specific benefits</h2>
<p>Besides the wide baseline stereo, local affine features and correspondences have other applications. I will briefly describe some of them here (to be updated).</p>
<section id="image-rectification" class="level3">
<h3 class="anchored" data-anchor-id="image-rectification">Image rectification</h3>
<p>Instead of matching local features between two images one might match them within a single image. Why would someone do it? This allows finding repeated pattern: think about windows, doors and so on. Typically they have the same physical size, therefore the diffrence in local features around them could tell us about the geometry of the scene and lens distortion.</p>
<div class="quarto-figure quarto-figure-center">
<figure class="figure">
<p><img src="https://ducha-aiki.github.io/wide-baseline-stereo-blog/posts/2020-07-17-affine-correspondences_files/att_00008.png" class="img-fluid figure-img"></p>
<figcaption>Repeated patterns detection with MSER andRootSIFT local features. Figure from the Minimal Solvers for Rectifying from Radially-Distorted Conjugate Translations, PAMI 2020 paper.</figcaption>
</figure>
</div>
<p>This is the idea of the series of works by Pritts and Chum.</p>
<div class="quarto-figure quarto-figure-center">
<figure class="figure">
<p><img src="https://ducha-aiki.github.io/wide-baseline-stereo-blog/posts/2020-07-17-affine-correspondences_files/att_00007.png" class="img-fluid figure-img"></p>
<figcaption>Figure from the ‘Minimal Solvers for Rectifying from Radially-Distorted Conjugate Translations’, PAMI 2020 paper.</figcaption>
</figure>
</div>
</section>
<section id="surface-normals-estimation" class="level3">
<h3 class="anchored" data-anchor-id="surface-normals-estimation">Surface normals estimation</h3>
<p>Ivan Eichhardt and Levente Hajder have a series of works, exploiting the local affine correspondences for surface normals estimation</p>
<div class="quarto-figure quarto-figure-center">
<figure class="figure">
<p><img src="https://ducha-aiki.github.io/wide-baseline-stereo-blog/posts/2020-07-17-affine-correspondences_files/att_00015.png" class="img-fluid figure-img"></p>
<figcaption>Estimated surface normals. Figure from Optimal Multi-View Surface Normal Estimation Using Affine Correspondences</figcaption>
</figure>
</div>
</section>
</section>
</section>
<section id="summary" class="level1">
<h1>Summary</h1>
<p>Despite not being popular right now, treating keypoints as local affine features has a lot of advantages over the traditional treatment the local correspondence as the point correspondences. In the next post I will describe a way of estimating the local feature affine shape and orientation.</p>
</section>
<section id="references" class="level1">
<h1>References</h1>
<p>[SuperPoint2017] Detone D., Malisiewicz T. and Rabinovich A., ``<em>Superpoint: Self-Supervised Interest Point Detection and Description</em>’’, CVPRW Deep Learning for Visual SLAM, vol.&nbsp;, number , pp.&nbsp;, 2018.</p>
<p>[Lowe99] D. Lowe, ``<em>Object Recognition from Local Scale-Invariant Features</em>’’, ICCV, 1999.</p>
<p>[Harris88] C. Harris and M. Stephens, ``<em>A Combined Corner and Edge Detector</em>’’, Fourth Alvey Vision Conference, 1988.</p>
<p>[Hessian78] P.R. Beaudet, ``<em>Rotationally invariant image operators</em>’’, Proceedings of the 4th International Joint Conference on Pattern Recognition, 1978.</p>
<p>[KeyNet2019] A. Barroso-Laguna, E. Riba, D. Ponsa <em>et al.</em>, ``<em>Key.Net: Keypoint Detection by Handcrafted and Learned CNN Filters</em>’’, ICCV, 2019.</p>
<p>[ORB2011] E. Rublee, V. Rabaud, K. Konolidge <em>et al.</em>, ``<em>ORB: An Efficient Alternative to SIFT or SURF</em>’’, ICCV, 2011.</p>
<p>[HardNet2017] A. Mishchuk, D. Mishkin, F. Radenovic <em>et al.</em>, ``<em>Working Hard to Know Your Neighbor’s Margins: Local Descriptor Learning Loss</em>’’, NeurIPS, 2017.</p>
<p>[RIFT2005] {Lazebnik} S., {Schmid} C. and {Ponce} J., ``<em>A sparse texture representation using local affine regions</em>’’, IEEE Transactions on Pattern Analysis and Machine Intelligence, vol.&nbsp;27, number 8, pp.&nbsp;1265-1278, 2005.</p>
<p>[sGLOH2] {Bellavia} F. and {Colombo} C., ``<em>Rethinking the sGLOH Descriptor</em>’’, IEEE Transactions on Pattern Analysis and Machine Intelligence, vol.&nbsp;40, number 4, pp.&nbsp;931-944, 2018.</p>
<p>[OriNet2016] K. M., Y. Verdie, P. Fua <em>et al.</em>, ``<em>Learning to Assign Orientations to Feature Points</em>’’, CVPR, 2016.</p>
<p>[AffNet2018] D. Mishkin, F. Radenovic and J. Matas, ``<em>Repeatability is Not Enough: Learning Affine Regions via Discriminability</em>’’, ECCV, 2018.</p>
<p>[PerdochRetrieval2009] M. {Perd’och}, O. {Chum} and J. {Matas}, ``<em>Efficient representation of local geometry for large scale object retrieval</em>’’, CVPR, 2009.</p>
<p>[IMW2020] Jin Yuhe, Mishkin Dmytro, Mishchuk Anastasiia <em>et al.</em>, ``<em>Image Matching across Wide Baselines: From Paper to Practice</em>’’, arXiv preprint arXiv:2003.01587, vol.&nbsp;, number , pp.&nbsp;, 2020.</p>
<p>[gcransac2018] D. Barath and J. Matas, ``<em>Graph-Cut RANSAC</em>’’, The IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2018.</p>
<p>[magsac2019] J.N. Daniel Barath, ``<em>MAGSAC: marginalizing sample consensus</em>’’, CVPR, 2019.</p>
<p>[Philbin07] J. Philbin, O. Chum, M. Isard <em>et al.</em>, ``<em>Object Retrieval with Large Vocabularies and Fast Spatial Matching</em>’’, CVPR, 2007.</p>
<p>[perd2006epipolar] M. Perd’och, J. Matas and O. Chum, ``<em>Epipolar geometry from two correspondences</em>’’, ICPR, 2006.</p>
<p>[PrittsRANSAC2013] J. {Pritts}, O. {Chum} and J. {Matas}, ``<em>Approximate models for fast and accurate epipolar geometry estimation</em>’’, 2013 28th International Conference on Image and Vision Computing New Zealand (IVCNZ 2013), 2013.</p>
<p>[Barath2019ICCV] D. Barath and Z. Kukelova, ``<em>Homography From Two Orientation- and Scale-Covariant Features</em>’’, ICCV, 2019.</p>
<p>[RANSACAffine2020] M. {Rodríguez}, G. {Facciolo}, R. G. <em>et al.</em>, ``<em>Robust estimation of local affine maps and its applications to image matching</em>’’, 2020 IEEE Winter Conference on Applications of Computer Vision (WACV), 2020.</p>
<p>[barath2020making] Barath Daniel, Polic Michal, Förstner Wolfgang <em>et al.</em>, ``<em>Making Affine Correspondences Work in Camera Geometry Computation</em>’’, arXiv preprint arXiv:2007.10032, vol.&nbsp;, number , pp.&nbsp;, 2020.</p>
<p>[guan2020relative] Guan Banglei, Zhao Ji, Barath Daniel <em>et al.</em>, ``<em>Relative Pose Estimation for Multi-Camera Systems from Affine Correspondences</em>’’, arXiv preprint arXiv:2007.10700, vol.&nbsp;, number , pp.&nbsp;, 2020.</p>
<p>[OneACMonoDepth2020] D.B. Ivan Eichhardt, ``<em>Relative Pose from Deep Learned Depth and a Single Affine Correspondence</em>’’, ECCV, 2020.</p>
<p>[SurfaceNormals2019] {Baráth} D., {Eichhardt} I. and {Hajder} L., ``<em>Optimal Multi-View Surface Normal Estimation Using Affine Correspondences</em>’’, IEEE Transactions on Image Processing, vol.&nbsp;28, number 7, pp.&nbsp;3301-3311, 2019.</p>


</section>

<p>Everything you (didn’t) want to know about image matching</p> ]]></description>
  <guid>https://ducha-aiki.github.io/wide-baseline-stereo-blog/posts/2020-07-17-affine-correspondences.html</guid>
  <pubDate>Thu, 16 Jul 2020 22:00:00 GMT</pubDate>
  <media:content url="https://ducha-aiki.github.io/wide-baseline-stereo-blog/posts/images/affine_matches.png" medium="image" type="image/png" height="48" width="144"/>
</item>
</channel>
</rss>
