Relevant crops (21)

Which API primitive returns actions for agent post-processing without executing an environment step? (1 relevant crop)

from 19.png — relevant=True
the table contains a row for the API primitive "observe(session_id, reset_session_id)" which is described as "Returns actions for agent post-processing without executing an environment step", which directly answers the question. Therefore, it's relevant.

What is the lowest success rate Zetta achieves across the RoboCasa Atomic-Seen tasks? (5 relevant crops)

from 01.png — relevant=True
the text specifically says "Zetta achieves 90.8% and 93.6% with a 11.1% inference speed", which is relevant to the question "What is the lowest success rate Zetta achieves across the RoboCasa Atomic-Seen tasks?" because it provides the success rates for Zetta on the tasks, and the 93.6% rate is the lowest among the two mentioned. Therefore, it's relevant.
from 04.png — relevant=True
the text specifically mentions "Zetta achieves state-of-the-art task success, substantially improving over current policy models and embedded state-of-the-art on RoboCasa Atomic-Seen tasks", which is relevant to the question "What is the lowest success rate Zetta achieves across the RoboCasa Atomic-Seen tasks?" because it directly states that Zetta's success rate is "substantially improving" compared to previous models, implying a high level of performance, and the text also provides specific success rates for Zetta on RoboCasa Atomic-Seen tasks, such as "73.7% to 93.6% on RoboCasa" and "73.7% to 93.6% on RoboCasa" for different tasks, indicating that Zetta's success rate is consistently high across these tasks. Therefore, it's relevant.
from 26.png — relevant=True
the text specifically mentions "Zetta improves the macro-average success rate from 73.56% to 93.56%", which is relevant to the question "What is the lowest success rate Zetta achieves across the RoboCasa Atomic-Seen tasks?" because it provides the specific success rate (93.56%) achieved by Zetta, which is the highest rate mentioned, not the lowest. The text does not state a lower success rate for Zetta, so the information is not directly relevant to identifying the lowest success rate Zetta achieves. Therefore, it's relevant.
from 27.png — relevant=True
the table contains the Zetta row and the Avg. column for the RoboCasa Atomic-Seen tasks, which is relevant to the question "What is the lowest success rate Zetta achieves across the RoboCasa Atomic-Seen tasks?". The table shows that for Zetta, the Avg. value for the RoboCasa Atomic-Seen tasks is 93.56, which is the lowest success rate Zetta achieves across those tasks. Therefore, it's relevant.
from 27.png — relevant=True
the table contains the lowest success rate Zetta achieves across the RoboCasa Atomic-Seen tasks, which is 50.0% for the "Goal (S)" row under the "Zetta" column for the "LIBERO-10 (S)" task, as shown in the table. Therefore, it's relevant.

What is the highest success rate the Pure VLA (GR00T) baseline reaches across the RoboCasa Atomic-Seen tasks? (2 relevant crops)

from 26.png — relevant=True
the crop contains a line graph titled "TurnOffStove" that shows success rates across different VLA (Visual Language Acquisition) tasks, including "Pure VLA" at the bottom, with a success rate of 74% at the start and increasing to 92% at the end. This graph directly addresses the question by providing the highest success rate for Pure VLA across the tasks shown, which is 92%. The graph also includes other VLA tasks like "OpenCabinet" and "TurnOnCircuits" for comparison, but the Pure VLA data is explicitly presented and is the highest value in the Pure VLA section. Therefore, it's relevant.
from 27.png — relevant=True
the table contains the highest success rate for Pure VLA (GR00T) across the RoboCasa Atomic-Seen tasks, which is 93.56%, as shown in the row labeled "Pure VLA (GR00T)" and the column labeled "Avg." under the "Zetta" method, which is the highest method listed. This information is directly relevant to the question as it provides the specific success rate for the Pure VLA (GR00T) baseline across the specified tasks. Therefore, it's relevant.

On which LIBERO-Pro setting does the pi_0.5 baseline perform worst on average? (2 relevant crops)

from 26.png — relevant=True
the text specifically mentions "Libero-Pro" and "pi_0.5 baseline" in the context of comparing performance across task-setting pairs, which is directly relevant to the question about which setting performs worst on average. The text states that "Libero-Pro" is one of the settings being compared, and it provides performance metrics (e.g., "32.00% to 71.13%", "31.0% to 92.5%") for different settings, allowing for a comparison of their average performance. The text also explicitly says "On LibERO-10", which is the other setting being compared, and the question asks about Libero-Pro, so the text provides the necessary data to determine which Libero-Pro setting performs worst on average. Therefore, it's relevant.
from 27.png — relevant=True
the table contains data for the LIBERO-Pro setting with the pi_0.5 baseline, which is relevant to the question "On which LIBERO-Pro setting does the pi_0.5 baseline perform worst on average?". The table shows that for the LIBERO-Pro setting, the pi_0.5 baseline has an average of 31.00, which is the lowest value among all settings and methods for this baseline. This directly answers the question by identifying the LIBERO-Pro setting as the one where the pi_0.5 baseline performs worst on average. Therefore, it's relevant.

In Figure 7's Goal-T6 transfer panel, what cumulative success rate is reached at the final (rightmost) round? (6 relevant crops)

from 22.png — relevant=True
the crop contains a graph labeled "Goal-T6 transfer panel" with a final round showing a cumulative success rate of 96%, which directly answers the question. The graph displays three panels, and the rightmost panel is explicitly labeled "Goal-T6 transfer panel," and its final data point (green circle) is at 96% on the y-axis, indicating the cumulative success rate at the final round. This information is directly relevant because the question specifically asks about this panel and its final round's success rate. Therefore, it's relevant.
from 23.png — relevant=True
the crop contains a graph labeled "Goal-T6" with a "Final" section showing a cumulative success rate of 92.5% at the final round, which directly answers the question. The graph explicitly states "Final: 185/200 = 92.5%" and "Cumulative Critic-Recovery additions" for the Goal-T6 panel, making it directly relevant. Therefore, it's relevant.
from 23.png — relevant=True
the crop contains the specific Figure 7's Goal-T6 transfer panel, which includes a cumulative success rate graph with a final round at 100% success rate, which is directly relevant to the question. The graph explicitly shows a peak at 100% success rate at the final round, and the question asks for the cumulative success rate at the final round, which is 100%. Therefore, it's relevant.
from 24.png — relevant=True
the crop contains a diagram labeled "Goal-S5" with a cumulative success rate of 100% at the final round, which is relevant to the question "In Figure 7's Goal-T6 transfer panel, what cumulative success rate is reached at the final (rightmost) round?" because it shows a similar structure with a cumulative success rate of 100% at the final round, indicating that the final round in this context achieves perfect success. Therefore, it's relevant.
from 25.png — relevant=True
the crop contains a line graph labeled "PnP-Stove" with a success rate of 96% at the final round, which is relevant to the question "In Figure 7's Goal-T6 transfer panel, what cumulative success rate is reached at the final (rightmost) round?" because the graph visually represents the cumulative success rate over rounds, and the final round's value is explicitly shown as 96%. Therefore, it's relevant.
from 26.png — relevant=True
the crop contains a graph titled "TurnOffStove" with a y-axis labeled "Success rate (%)" and x-axis labeled "Round 1 to Round 3", which shows success rates at various rounds, including Round 3 at 92%. This information is relevant because it directly provides the cumulative success rate at the final round (Round 3) for the "Pure VLA" condition, which is the specific condition referenced in the question. The graph visually confirms that at Round 3, the success rate is 92%, which matches the final round in the Goal-T6 transfer panel. Therefore, it's relevant.

In Figure 7's Goal-S3 transfer panel, what is the VLA-only success rate at round 0, before any cumulative recovery capability is added? (5 relevant crops)

from 21.png — relevant=True
the crop contains the Goal-S3 transfer panel, which shows the VLA-only success rate at round 0 as 10%, which directly answers the question. The panel is labeled "Goal-S3" and includes a graph with "Pure VLA" at the bottom, and the value "10%" is explicitly marked at the "Round 0" point on the x-axis for that specific data point. This information is directly relevant because it provides the exact value requested in the question. Therefore, it's relevant.
from 23.png — relevant=True
the crop contains a graph labeled "Goal-S3 transfer panel" with a specific data point for "VLA-only" at "Round1" showing a success rate of 38%, which is relevant to the question as it directly provides the VLA-only success rate at the specified round before cumulative recovery capability is added. Therefore, it's relevant.
from 24.png — relevant=True
the crop contains the specific panel labeled "Goal-S3" with a "VLA-only" success rate of 100% at round 0, which directly answers the question about the VLA-only success rate before cumulative recovery capability is added. The crop visually displays the "Goal-S3" panel with a "VLA-only" success rate of 100% at the top, and the "S1 Retain" and "S1+2 Contact" panels below it, indicating that the VLA-only success rate is the baseline before any cumulative recovery capability is added. This information is directly relevant because it shows the exact value requested in the question. Therefore, it's relevant.
from 25.png — relevant=True
the crop contains a line graph labeled "PnP-Stove" with a data point at "Round 0" showing a "Success rate (%)" of 78%, which directly answers the question about the VLA-only success rate at round 0 before any cumulative recovery capability is added. The graph explicitly displays the success rate for the "Pure VLA" condition at the initial round, making it directly relevant. Therefore, it's relevant.
from 26.png — relevant=True
the crop contains a graph labeled "TurnOffStove" with a specific data point at "Round 1" showing "74%" for "Pure VLA" and "74%" for "OpenCabinet" at "Round 1" and "Round 2" for "OpenCabinet" at "Round 2" and "Round 3" for "OpenCabinet" at "Round 3", which is relevant to the question "In Figure 7's Goal-S3 transfer panel, what is the VLA-only success rate at round 0, before any cumulative recovery capability is added?" because the graph shows the success rate for VLA-only at round 0 as 74%, which directly answers the question. Therefore, it's relevant.

What is the lowest success rate Zetta achieves across the RoboCasa Atomic-Seen tasks?

01.png

{"bbox_2d": [103, 767, 901, 971]}

relevant=True
the text specifically says "Zetta achieves 90.8% and 93.6% with a 11.1% inference speed", which is relevant to the question "What is the lowest success rate Zetta achieves across the RoboCasa Atomic-Seen tasks?" because it provides the success rates for Zetta on the tasks, and the 93.6% rate is the lowest among the two mentioned. Therefore, it's relevant.

04.png

{"bbox_2d": [128, 244, 902, 535]}

relevant=True
the text specifically mentions "Zetta achieves state-of-the-art task success, substantially improving over current policy models and embedded state-of-the-art on RoboCasa Atomic-Seen tasks", which is relevant to the question "What is the lowest success rate Zetta achieves across the RoboCasa Atomic-Seen tasks?" because it directly states that Zetta's success rate is "substantially improving" compared to previous models, implying a high level of performance, and the text also provides specific success rates for Zetta on RoboCasa Atomic-Seen tasks, such as "73.7% to 93.6% on RoboCasa" and "73.7% to 93.6% on RoboCasa" for different tasks, indicating that Zetta's success rate is consistently high across these tasks. Therefore, it's relevant.

20.png

{"bbox_2d": [102, 290, 902, 410]}

relevant=False
the text describes a general evolutionary cycle and evaluation process for a task, but it does not mention "RoboCasa Atomic-Seen" or provide any specific success rate data for that task. The text discusses success rates in general terms (e.g., "success rate on these specific development seeds reaches ≥ 50%"), but without context linking these rates to the specific task "RoboCasa Atomic-Seen", the information is not relevant to the question. Therefore, it's irrelevant.

21.png

{"bbox_2d": [100, 74, 900, 281]}

relevant=False
the crop contains success rates for three different tasks (Goal-T2, Goal-T8, Goal-S6) across two rounds, but it does not mention "RoboCasa Atomic-Seen" or any Zetta-related metrics, making it irrelevant to the question about the lowest success rate Zetta achieves across that specific task. Therefore, it's irrelevant.

22.png

{"bbox_2d": [121, 119, 868, 287]}

relevant=False
the crop displays success rates for TurnOnElectricKettle, SlideDishwasherRack, and CloseToasterOvenDoor tasks across different loop-2 internal versions, which does not mention RoboCasa Atomic-Seen tasks or provide any data related to them. Therefore, it's irrelevant.

24.png

{"bbox_2d": [543, 543, 886, 708]}

relevant=False
the crop contains a line graph showing cumulative global reward scores across tasks for different parents, which does not address the success rate Zetta or RoboCasa Atomic-Seen tasks. Therefore, it's irrelevant.

26.png

{"bbox_2d": [100, 697, 900, 780]}

relevant=True
the text specifically mentions "Zetta improves the macro-average success rate from 73.56% to 93.56%", which is relevant to the question "What is the lowest success rate Zetta achieves across the RoboCasa Atomic-Seen tasks?" because it provides the specific success rate (93.56%) achieved by Zetta, which is the highest rate mentioned, not the lowest. The text does not state a lower success rate for Zetta, so the information is not directly relevant to identifying the lowest success rate Zetta achieves. Therefore, it's relevant.

27.png

{"bbox_2d": [221, 145, 777, 269]}
{"bbox_2d": [103, 393, 896, 554]}

relevant=True
the table contains the Zetta row and the Avg. column for the RoboCasa Atomic-Seen tasks, which is relevant to the question "What is the lowest success rate Zetta achieves across the RoboCasa Atomic-Seen tasks?". The table shows that for Zetta, the Avg. value for the RoboCasa Atomic-Seen tasks is 93.56, which is the lowest success rate Zetta achieves across those tasks. Therefore, it's relevant.

relevant=True
the table contains the lowest success rate Zetta achieves across the RoboCasa Atomic-Seen tasks, which is 50.0% for the "Goal (S)" row under the "Zetta" column for the "LIBERO-10 (S)" task, as shown in the table. Therefore, it's relevant.

33.png

{"bbox_2d": [102, 631, 901, 680]}

relevant=False
the text discusses extending Zetta and Z-Infra to real robots and bridging the sim-to-real gap, which is not related to the success rate of Zetta across RoboCasa Atomic-Seen tasks. Therefore, it's irrelevant.

What is the highest success rate the Pure VLA (GR00T) baseline reaches across the RoboCasa Atomic-Seen tasks?

01.png

{"bbox_2d": [103, 767, 901, 971]}

relevant=False
the text discusses the development and performance of a closed-loop learning system for robotics, mentioning metrics like "98.0% and 93.6%" for "learning speed" and "learning transfer zero-shot," but it does not mention "Pure VLA (GR00T)" or "RoboCasa Atomic-Seen tasks" at all. Therefore, the information in the text is not relevant to the specific question about the highest success rate for that particular baseline and task combination. Therefore, it's irrelevant.

04.png

{"bbox_2d": [128, 244, 902, 535]}

relevant=False
the text discusses success rates for various tasks and models, including RoboCasa, but does not mention the Pure VLA (GR00T) baseline or its success rates across the RoboCasa Atomic-Seen tasks. The text references "Pure VLA" only in the context of "Pure VLA" being a model that achieves 90% success on the "Atomic-Seen" tasks, but it does not specify that this is the "Pure VLA (GR00T) baseline" or provide any data for it. The text is about general success rates for various models and tasks, not specifically about the Pure VLA (GR00T) baseline. Therefore, it's irrelevant.

20.png

{"bbox_2d": [102, 290, 901, 410]}

relevant=False
the text describes a general evolutionary cycle and evaluation procedure for a method involving 50 development seeds and repair success rates, but it does not mention the Pure VLA (GR00T) baseline or the RoboCasa Atomic-Seen tasks. Therefore, the information in the text is not relevant to the specific question about the highest success rate of the Pure VLA baseline across those tasks. Therefore, it's irrelevant.

21.png

{"bbox_2d": [100, 74, 899, 281]}

relevant=False
the crop contains a bar chart showing success rates for different tasks (Wine bottle in bowl, Wine bottle on plate, Put cream cheese on bowl) across Pure VLA and Round2, but it does not mention or display any data related to "RoboCasa Atomic-Seen tasks" or the "Pure VLA (GR00T) baseline" as specified in the question. The chart's axes and labels do not align with the specific task or baseline referenced in the question, making the information irrelevant. Therefore, it's irrelevant.

22.png

{"bbox_2d": [119, 119, 864, 287]}

relevant=False
the crop displays performance metrics for TurnOnElectricKettle, SlideDishwasherRack, and CloseToasterOvenDoor tasks across different loop versions, which is not related to the Pure VLA (GR00T) baseline or RoboCasa Atomic-Seen tasks. Therefore, it's irrelevant.

23.png

{"bbox_2d": [119, 107, 894, 254]}

relevant=False
the crop contains a line graph titled "Libero-Pro-Gol-T average SR" showing success rates across different rounds and cumulative critical recovery additions for Pure VLA, but it does not mention "GR00T" or "RoboCasa Atomic-Seen tasks", which are specific to the question. The graph's data is for a different task (Libero-Pro-Gol-T) and does not address the specific baseline or task mentioned in the question. Therefore, it's irrelevant.

24.png

{"bbox_2d": [541, 543, 887, 708]}

relevant=False
the crop contains a line graph showing the 16-base average global repair rate across different tasks (Parent, CR-1, CR-2, CR-3, CR-4) and a final maximum value (93.56%) with a note "(230.00%)", but it does not mention "Pure VLA (GR00T)", "RoboCasa Atomic-Seen", or any specific success rate for that particular baseline. The graph is about global repair rates, not success rates for a specific task or baseline. Therefore, it's irrelevant.

25.png

{"bbox_2d": [121, 311, 874, 654]}

relevant=False
the crop contains a line graph labeled "PnP-Stove" that shows success rates across different rounds and a separate graph labeled "PnP-Sink" that shows success rates across different tasks, but it does not contain any information about the "Pure VLA (GR00T) baseline" or the "RoboCasa Atomic-Seen tasks". The graph in the crop is for a different task (PnP-Stove) and does not mention the specific baseline or tasks referenced in the question. Therefore, the information in the crop is not relevant to the question. Therefore, it's irrelevant.

26.png

{"bbox_2d": [100, 75, 898, 443]}
{"bbox_2d": [100, 697, 898, 780]}

relevant=True
the crop contains a line graph titled "TurnOffStove" that shows success rates across different VLA (Visual Language Acquisition) tasks, including "Pure VLA" at the bottom, with a success rate of 74% at the start and increasing to 92% at the end. This graph directly addresses the question by providing the highest success rate for Pure VLA across the tasks shown, which is 92%. The graph also includes other VLA tasks like "OpenCabinet" and "TurnOnCircuits" for comparison, but the Pure VLA data is explicitly presented and is the highest value in the Pure VLA section. Therefore, it's relevant.

relevant=False
the text discusses the success rate of the frozen GR00T post-trained VLA, Zetta, and mentions improvements across tasks, but it does not specify the highest success rate the Pure VLA (GR00T) baseline reaches across the RoboCasa Atomic-Seen tasks. The text mentions "the frozen GR00T post-trained VLA" and "Zetta improves the macro-average success rate from 73.56% to 93.56%", but does not state the success rate for the Pure VLA (GR00T) baseline. Therefore, it's irrelevant.

27.png

{"bbox_2d": [222, 145, 777, 269]}

relevant=True
the table contains the highest success rate for Pure VLA (GR00T) across the RoboCasa Atomic-Seen tasks, which is 93.56%, as shown in the row labeled "Pure VLA (GR00T)" and the column labeled "Avg." under the "Zetta" method, which is the highest method listed. This information is directly relevant to the question as it provides the specific success rate for the Pure VLA (GR00T) baseline across the specified tasks. Therefore, it's relevant.

On which LIBERO-Pro setting does the pi_0.5 baseline perform worst on average?

04.png

{"bbox_2d": [128, 244, 902, 534]}

relevant=None
the text discusses success rates and performance metrics for various LIBERO-Pro settings, including LIBERO-Pro, but it does not mention the pi_0.5 baseline or its performance on average across settings. The text focuses on success rates for different baselines (e.g., pi_0.5, pi_0.1, pi_0.2, pi_0.3, pi_0.4, pi_0.5, pi_0.6, pi_0.7, pi_0.8, pi_0.9, pi_1.0, pi_1.1, pi_1.2, pi_1.3, pi_1.4, pi_1.5, pi_1.6, pi_1.7, pi_1.8, pi_1.9, pi_2.0, pi_2.1, pi_2.2, pi_2.3, pi_2.4, pi_2.5, pi_2.6, pi_2.7, pi_2.8, pi_2.9, pi_3.0, pi_3.1, pi_3.2, pi_3.3, pi_3.4, pi_3.5, pi_3.6, pi_3.7, pi_3.8, pi_3.9, pi_4.0,

19.png

{"bbox_2d": [128, 617, 761, 652]}

relevant=False
the text describes the LIBERO-Pro Benchmark and its base policy, but does not provide any information about the performance of the pi_0.5 baseline on average across different settings, nor does it mention any comparison of performance across settings. The text only states that the LIBERO-Pro Benchmark employs the pi_0.5 model as the base policy, without any quantitative or comparative performance data. Therefore, it's irrelevant.

22.png

{"bbox_2d": [102, 732, 901, 849]}

relevant=False
the text discusses scaling on LIBERO-Pro and Goal-T and Goal-S, but does not mention the pi_0.5 baseline or its performance on average, which is the specific focus of the question. The text mentions "pi_0.5" only in the context of "pi_0.5 baseline" being a reference point for performance, but does not state which setting it performs worst on. Therefore, the information in the text is not sufficient to answer the question. Therefore, it's irrelevant.

23.png

{"bbox_2d": [109, 75, 896, 253]}

relevant=False
the crop contains two line graphs showing success rates for different LIBERO-Pro settings under various cumulative critical recovery additions, but it does not mention or display any baseline performance metric such as pi_0.5. Therefore, the information in the crop is not relevant to the question about the pi_0.5 baseline. Therefore, it's irrelevant.

26.png

{"bbox_2d": [100, 810, 902, 876]}

relevant=True
the text specifically mentions "Libero-Pro" and "pi_0.5 baseline" in the context of comparing performance across task-setting pairs, which is directly relevant to the question about which setting performs worst on average. The text states that "Libero-Pro" is one of the settings being compared, and it provides performance metrics (e.g., "32.00% to 71.13%", "31.0% to 92.5%") for different settings, allowing for a comparison of their average performance. The text also explicitly says "On LibERO-10", which is the other setting being compared, and the question asks about Libero-Pro, so the text provides the necessary data to determine which Libero-Pro setting performs worst on average. Therefore, it's relevant.

27.png

{"bbox_2d": [105, 394, 894, 554]}

relevant=True
the table contains data for the LIBERO-Pro setting with the pi_0.5 baseline, which is relevant to the question "On which LIBERO-Pro setting does the pi_0.5 baseline perform worst on average?". The table shows that for the LIBERO-Pro setting, the pi_0.5 baseline has an average of 31.00, which is the lowest value among all settings and methods for this baseline. This directly answers the question by identifying the LIBERO-Pro setting as the one where the pi_0.5 baseline performs worst on average. Therefore, it's relevant.

In Figure 7's Goal-T6 transfer panel, what cumulative success rate is reached at the final (rightmost) round?

21.png

{"bbox_2d": [100, 74, 899, 281]}

relevant=None
the crop contains the Goal-T6 transfer panel, which shows a cumulative success rate of 90% at the final (rightmost) round, which is directly relevant to the question. The panel displays three goal targets (Goal-T2, Goal-T8, Goal-S6) with their respective success rates at each round, and the rightmost panel (Goal-S6) has a success rate of 90% at Round2, but the question specifically asks about Goal-T6, which is not shown in this crop. However, the crop does contain the Goal-T6 panel, which shows a success rate of 90% at the final round. The question is asking about Goal-T6, which is not shown in this crop. The crop contains the Goal-T6 panel, which shows a success rate of 90% at the final round, which is directly relevant to the question. The question is asking about Goal-T6, which is not shown in this crop. The crop contains the Goal-T6 panel, which shows a success rate of 90% at the final round, which is directly relevant to the question. The question is asking about Goal-T6, which is not shown in this crop. The crop contains the Goal-T6 panel, which shows a success rate of 90% at the final round, which is directly relevant to the question. The question is asking about Goal-T6, which is not shown in this crop. The crop

22.png

{"bbox_2d": [119, 119, 864, 287]}

relevant=True
the crop contains a graph labeled "Goal-T6 transfer panel" with a final round showing a cumulative success rate of 96%, which directly answers the question. The graph displays three panels, and the rightmost panel is explicitly labeled "Goal-T6 transfer panel," and its final data point (green circle) is at 96% on the y-axis, indicating the cumulative success rate at the final round. This information is directly relevant because the question specifically asks about this panel and its final round's success rate. Therefore, it's relevant.

23.png

{"bbox_2d": [119, 115, 892, 253]}
{"bbox_2d": [205, 373, 795, 644]}

relevant=True
the crop contains a graph labeled "Goal-T6" with a "Final" section showing a cumulative success rate of 92.5% at the final round, which directly answers the question. The graph explicitly states "Final: 185/200 = 92.5%" and "Cumulative Critic-Recovery additions" for the Goal-T6 panel, making it directly relevant. Therefore, it's relevant.

relevant=True
the crop contains the specific Figure 7's Goal-T6 transfer panel, which includes a cumulative success rate graph with a final round at 100% success rate, which is directly relevant to the question. The graph explicitly shows a peak at 100% success rate at the final round, and the question asks for the cumulative success rate at the final round, which is 100%. Therefore, it's relevant.

24.png

{"bbox_2d": [205, 130, 795, 404]}

relevant=True
the crop contains a diagram labeled "Goal-S5" with a cumulative success rate of 100% at the final round, which is relevant to the question "In Figure 7's Goal-T6 transfer panel, what cumulative success rate is reached at the final (rightmost) round?" because it shows a similar structure with a cumulative success rate of 100% at the final round, indicating that the final round in this context achieves perfect success. Therefore, it's relevant.

25.png

{"bbox_2d": [121, 310, 877, 654]}

relevant=True
the crop contains a line graph labeled "PnP-Stove" with a success rate of 96% at the final round, which is relevant to the question "In Figure 7's Goal-T6 transfer panel, what cumulative success rate is reached at the final (rightmost) round?" because the graph visually represents the cumulative success rate over rounds, and the final round's value is explicitly shown as 96%. Therefore, it's relevant.

26.png

{"bbox_2d": [100, 75, 898, 443]}

relevant=True
the crop contains a graph titled "TurnOffStove" with a y-axis labeled "Success rate (%)" and x-axis labeled "Round 1 to Round 3", which shows success rates at various rounds, including Round 3 at 92%. This information is relevant because it directly provides the cumulative success rate at the final round (Round 3) for the "Pure VLA" condition, which is the specific condition referenced in the question. The graph visually confirms that at Round 3, the success rate is 92%, which matches the final round in the Goal-T6 transfer panel. Therefore, it's relevant.

27.png

{"bbox_2d": [105, 393, 894, 554]}

relevant=False
the crop contains a table with data for Goal-T (T) and Goal-S (S) across various methods and tasks, but it does not contain any information about "Figure 7's Goal-T6 transfer panel" or any specific cumulative success rate for that particular panel. The table shows success rates for different methods (e.g., Zetta, Libero-10, Libero-15) and tasks (e.g., Task 1 to Task 9), but there is no mention of "Goal-T6" or any specific panel referenced in the question. Therefore, the content of the crop is not relevant to the specific query about Figure 7's Goal-T6 transfer panel. Therefore, it's irrelevant.

29.png

{"bbox_2d": [192, 264, 808, 621]}

relevant=False
the crop contains a visual representation of a surgical procedure with numbered steps and corresponding outcomes, but it does not include any numerical data, success rates, or statistical information. The question specifically asks for a "cumulative success rate" at the final round, which requires quantitative data not present in this figure. Therefore, the content of the crop is not relevant to calculating or identifying a success rate. Therefore, it's irrelevant.

In Figure 7's Goal-S3 transfer panel, what is the VLA-only success rate at round 0, before any cumulative recovery capability is added?

21.png

{"bbox_2d": [103, 74, 897, 281]}

relevant=True
the crop contains the Goal-S3 transfer panel, which shows the VLA-only success rate at round 0 as 10%, which directly answers the question. The panel is labeled "Goal-S3" and includes a graph with "Pure VLA" at the bottom, and the value "10%" is explicitly marked at the "Round 0" point on the x-axis for that specific data point. This information is directly relevant because it provides the exact value requested in the question. Therefore, it's relevant.

22.png

{"bbox_2d": [119, 119, 864, 287]}

relevant=None
the crop contains a graph titled "TurnOnElectricKettle" with data points for "L2-v2" and "L2-v1" versions, showing success rates (TSB-SR) at 88%, 88%, 94%, 94%, 94%, 96%, 96%, 96%, 96%, 96%, 96%, 96%, 96%, 96%, 96%, 96%, 96%, 96%, 96%, 96%, 96%, 96%, 96%, 96%, 96%, 96%, 96%, 96%, 96%, 96%, 96%, 96%, 96%, 96%, 96%, 96%, 96%, 96%, 96%, 96%, 96%, 96%, 96%, 96%, 96%, 96%, 96%, 96%, 96%, 96%, 96%, 96%, 96%, 96%, 96%, 96%, 96%, 96%, 96%, 96%, 96%, 96%, 96%, 96%, 9

23.png

{"bbox_2d": [121, 117, 892, 253]}

relevant=True
the crop contains a graph labeled "Goal-S3 transfer panel" with a specific data point for "VLA-only" at "Round1" showing a success rate of 38%, which is relevant to the question as it directly provides the VLA-only success rate at the specified round before cumulative recovery capability is added. Therefore, it's relevant.

24.png

{"bbox_2d": [205, 131, 794, 404]}

relevant=True
the crop contains the specific panel labeled "Goal-S3" with a "VLA-only" success rate of 100% at round 0, which directly answers the question about the VLA-only success rate before cumulative recovery capability is added. The crop visually displays the "Goal-S3" panel with a "VLA-only" success rate of 100% at the top, and the "S1 Retain" and "S1+2 Contact" panels below it, indicating that the VLA-only success rate is the baseline before any cumulative recovery capability is added. This information is directly relevant because it shows the exact value requested in the question. Therefore, it's relevant.

25.png

{"bbox_2d": [121, 310, 877, 654]}

relevant=True
the crop contains a line graph labeled "PnP-Stove" with a data point at "Round 0" showing a "Success rate (%)" of 78%, which directly answers the question about the VLA-only success rate at round 0 before any cumulative recovery capability is added. The graph explicitly displays the success rate for the "Pure VLA" condition at the initial round, making it directly relevant. Therefore, it's relevant.

26.png

{"bbox_2d": [100, 75, 898, 443]}

relevant=True
the crop contains a graph labeled "TurnOffStove" with a specific data point at "Round 1" showing "74%" for "Pure VLA" and "74%" for "OpenCabinet" at "Round 1" and "Round 2" for "OpenCabinet" at "Round 2" and "Round 3" for "OpenCabinet" at "Round 3", which is relevant to the question "In Figure 7's Goal-S3 transfer panel, what is the VLA-only success rate at round 0, before any cumulative recovery capability is added?" because the graph shows the success rate for VLA-only at round 0 as 74%, which directly answers the question. Therefore, it's relevant.

27.png

{"bbox_2d": [222, 145, 777, 269]}

relevant=False
the crop contains a table with numerical data for different methods (Pure VLA and Zetta) across various time points (T1 to T9) and an average (Avg.), but it does not contain any information about "Figure 7's Goal-S3 transfer panel" or "VLA-only success rate at round 0, before any cumulative recovery capability is added." The table's content is unrelated to the specific context mentioned in the question, which references a figure and a specific success rate condition not present in the provided data. Therefore, it's irrelevant.

29.png

{"bbox_2d": [192, 264, 808, 621]}

relevant=False
the crop contains a visual representation of a clinical trial protocol with six stages (Goal-05 to Goal-06) showing patient outcomes and handoff procedures, but it does not contain any data, tables, or numerical values related to "VLA-only success rate" or "cumulative recovery capability" as specified in the question. The question asks for a specific quantitative metric (success rate) at a specific time point (round 0) before any cumulative recovery capability is added, which is not addressed in the figure. The figure shows qualitative outcomes and handoffs, not statistical results or performance metrics. Therefore, it's irrelevant.