Relevant crops (27)

What is the identifier of the ARC-AGI-1 training task shown in Figure 1? (9 relevant crops)

from 02.png — relevant=True
the text specifically mentions "ARC-AGI-1" and "150M-parameter configuration", which is relevant to the question "What is the identifier of the ARC-AGI-1 training task shown in Figure 1?" because it directly references the task identifier and its associated parameters, which are key details for identifying the task. Therefore, it's relevant.
from 02.png — relevant=True
the text specifically mentions "ARC-AGI-1 cost-accuracy frontier" and "ARC-style data mixture to train the system", which is relevant to the question "What is the identifier of the ARC-AGI-1 training task shown in Figure 1?" because it directly references the ARC-AGI-1 task and its associated data mixture, indicating its identifier and context within the described system. Therefore, it's relevant.
from 03.png — relevant=True
the text specifically says "identifier 0520fde7", which directly answers the question "What is the identifier of the ARC-AGI-1 training task shown in Figure 1?" Therefore, it's relevant.
from 04.png — relevant=True
the text specifically mentions "ARC-AGI-1 training set" and "ARC-AGI-1 training set (Cholle, 2019)", which is relevant to the question "What is the identifier of the ARC-AGI-1 training task shown in Figure 1?" because it directly provides the identifier "ARC-AGI-1" as the name of the training task being discussed. Therefore, it's relevant.
from 05.png — relevant=True
the crop contains a chart titled "ARC-AGI-1 Training Task" with a legend indicating "The ARC-AGI-1 Training Task" and "The ARC-AGI-1 Training Task" is shown in the legend, which is relevant to the question "What is the identifier of the ARC-AGI-1 training task shown in Figure 1?" because the chart visually represents the ARC-AGI-1 training task and its associated cost and accuracy metrics, directly linking the identifier to the figure. Therefore, it's relevant.
from 10.png — relevant=True
the text specifically mentions "ARC-AGI-1" and "Figure 1" in the context of a statistical comparison, which is relevant to the question "What is the identifier of the ARC-AGI-1 training task shown in Figure 1?" The text states: "The full public ARC-AGI-1 evaluation, the MNI effort setting one third of STANDARD (0.00083839 versus 0.00262542 per task) and evaluated 111/400 or 118/400 pass2, a difference of −1.75 percentage point. The paired split was 105 tasks solved by both settings, 13 standard-only, 6 min-only, and 276 neither (two-sided exact McNemar p = 0.167). All four reported endpoints favor STANDARD in point estimate, but this comparison is statistically unresolvable." This indicates that ARC-AGI-1 is the identifier of the training task being evaluated, and the text provides details about its performance metrics, which are likely shown in Figure 1. Therefore, it's relevant.
from 13.png — relevant=True
the text specifically mentions "ARC-AGI-1" and "29.5% pass@", which is relevant to the question "What is the identifier of the ARC-AGI-1 training task shown in Figure 1?" because it directly references the task identifier and provides performance metrics associated with it. Therefore, it's relevant.
from 15.png — relevant=True
the text specifically mentions "ARC-AGI-1 evaluation set", which is relevant to the question "What is the identifier of the ARC-AGI-1 training task shown in Figure 1?" because it directly references the identifier "ARC-AGI-1" in the context of an evaluation set, which is a key component of the task's identification. The text also mentions "generated set calibrated to match its measured distribution" and "required operation", which are details about the task's structure and calibration, further supporting its relevance. Therefore, it's relevant.
from 16.png — relevant=True
the table contains the identifier "ARC-AGI-1" under the "Evaluation cohort" column, which directly corresponds to the training task mentioned in the question. The table also includes the number of tasks solved (118) and solve rate (29.5%) for this cohort, providing specific data relevant to the ARC-AGI-1 task. This information is directly relevant because it identifies the specific task (ARC-AGI-1) and its performance metrics, which are key to answering the question about its identifier. Therefore, it's relevant.

What pass@2 score does the 150M-parameter configuration reach on ARC-AGI-1, and at what per-task inference cost? (6 relevant crops)

from 01.png — relevant=True
the crop contains the specific information needed to answer the question. The text explicitly states: "A 150M-parameter configuration reaches 29.5% pass@2 at a computed inference cost of $0.007 per task—less than one-tenth of a cent." This directly provides both the pass@2 score (29.5%) and the per-task inference cost ($0.007), which are exactly the two metrics requested in the question. Therefore, it's relevant.
from 02.png — relevant=True
the crop contains the specific information needed to answer the question. The text explicitly states: "A 150M-parameter configuration reaches 29.5% pass@2 on ARC-AGI-1 at a computed $0.00070 per task, breaking through the previously reported cost-accuracy Pareto frontier." This directly provides both the pass@2 score (29.5%) and the per-task inference cost ($0.00070), which are exactly the two metrics requested in the question. Therefore, it's relevant.
from 04.png — relevant=True
the crop contains relevant information. The text states: "On the 150M-parameter system reaches 29.50% pass@2 in approximately 0.85 H200 GPU-seconds per task." This directly answers the question about the pass@2 score for the 150M-parameter configuration on ARC-AGI-1. Additionally, it mentions that this performance is "at least one less than the previously reported Pareto frontier," which provides context for the efficiency comparison. The text also references the ARC-AGI-1 cost efficiency and the BDD-CQ score, which are related metrics used to evaluate system performance. Therefore, it's relevant.
from 10.png — relevant=True
the crop contains relevant information. The text states: "The paired split was 105 tasks solved by both settings, 13 standard-only, 6 min-only, and 276 neither (two-sided exact McNemar p = 0.167). All four reported endpoints favor STANDARD in point estimate, but this comparison is statistically unresolvable." This indicates that for the 150M-parameter configuration (implied by the context of the full ARC-AGI-1 evaluation and the mention of "150M" in the original text), the pass@2 score is 111/1400, which is 0.8083393. The text also mentions "150M" as part of the full ARC-AGI-1 evaluation, confirming its relevance to the question. The per-task inference cost is not explicitly stated in this specific crop, but the context implies it is part of the same evaluation and is likely to be found in the full text or elsewhere in the document. Therefore, it's relevant.
from 12.png — relevant=True
the crop contains the specific information needed to answer the question. The text explicitly states: "a 150M-parameter BDH-CQ reaches 29.5% pass@2 on the public ARC-AGI-1 evaluation set at a computed $0.00070 per task." This directly provides both the pass@2 score (29.5%) and the per-task inference cost ($0.00070), which are exactly the two metrics requested in the question. Therefore, it's relevant.
from 13.png — relevant=True
the crop contains relevant information. The text states: "On ARC-AGI-1, this design reaches 29.5% pass@2 for $0.0070 per task, establishing a new state of an art benchmark cost efficiency." This directly answers the question by specifying the pass@2 score (29.5%) and the per-task inference cost ($0.0070) for the 150M-parameter configuration on ARC-AGI-1. Therefore, it's relevant.

In Table 1, what is the pass@1 percentage for ConceptARC with opaque IDs, tasks unit? (1 relevant crop)

from 06.png — relevant=True
the table contains the specific data point for ConceptARC with opaque IDs, tasks unit, which is 96 (60.00%), which directly answers the question. Therefore, it's relevant.

In Figure 4, what transformation does the 'Order eight bars by height' example illustrate? (2 relevant crops)

from 08.png — relevant=True
the crop contains the specific example labeled "Order eight bars by height" and its corresponding "Oracle output", which directly illustrates the transformation being asked about. The crop visually shows the input as a grid of eight bars of varying heights and the output as a grid of eight bars of varying heights, with the transformation applied to the data. This is precisely the information needed to answer the question about what transformation this example illustrates. Therefore, it's relevant.
from 08.png — relevant=True
the text specifically mentions "order bar colors from shortest to tallest" as part of the transformations shown in Figure 4, which directly addresses the question about what transformation the 'Order eight bars by height' example illustrates. The text states: "order bar colors from shortest to tallest", indicating that the transformation involves changing the color of bars from the shortest to the tallest, which is a form of height-based ordering. This is relevant because the question asks about the transformation illustrated in Figure 4, and the text explicitly describes this transformation. Therefore, it's relevant.

What does BDH-CQ do with the demonstrations instead of compressing them into a single task vector? (6 relevant crops)

from 02.png — relevant=True
the text describes BDH-CQ as combining BDI-in-context learning with iterative recurrent memory and evolving recurrent memory, which is relevant to the question about what BDH-CQ does with demonstrations instead of compressing them into a single task vector. The text states that BDH-CQ "combines BDI-in-context learning through evolving recurrent memory with iterative recurrent memory," indicating that it does not compress demonstrations into a single task vector but instead uses a more flexible, iterative, and context-aware approach. This is directly related to the question's focus on the method's behavior with demonstrations. Therefore, it's relevant.
from 02.png — relevant=True
the crop contains relevant information. The text states: "The task specification also makes the behavior of BDH-CQ accessible to controlled analysis. The demonstrations constitute the task format, while the small colored grids require little factual knowledge or linguistic fluency. They are nevertheless express object relations, counting, symmetry, topology, spatial transformations, and compositions." This indicates that BDH-CQ does not compress demonstrations into a single task vector; instead, it allows for controlled analysis of multiple demonstrations (which are composed of small grids) to understand complex behaviors. The phrase “accessible to controlled analysis” directly supports the idea that BDH-CQ preserves the demonstrations as a set of tasks rather than reducing them to a single vector. Therefore, it's relevant.
from 03.png — relevant=True
the text specifically says "Rather than compressing the demonstrations into a single task vector, BDH-CQ processes their elements sequentially", which directly answers the question about what BDH-CQ does with the demonstrations instead of compressing them into a single task vector. Therefore, it's relevant.
from 11.png — relevant=True
the text discusses BDH-CQ combining with an autoregressive computational scratchpad and a continuous recurrent workspace, which is relevant to the question about what BDH-CQ does with demonstrations instead of compressing them into a single task vector. The text states that BDH-CQ "combines the same contextual flexibility with a continuous recurrent workspace," implying it does not compress demonstrations into a single vector but instead integrates them into a more flexible, continuous model. This directly addresses the question by describing BDH-CQ's approach to handling demonstrations rather than reducing them to a single task vector. Therefore, it's relevant.
from 11.png — relevant=True
the crop contains relevant information. The text states: "BDH-CQ explores a different operating regime: demonstrations update task memory, and a structured multi-vector workspace previously unseen in transformational models." This directly addresses the question by explaining that BDH-CQ does not compress demonstrations into a single task vector; instead, it uses a structured multi-vector workspace for demonstrations, which is a distinct operational approach. The phrase “demonstrations update task memory” further supports this by indicating that demonstrations are processed as separate, structured entities rather than compressed into a single vector. Therefore, it's relevant.
from 12.png — relevant=True
the crop contains relevant information. The text states: "BDH-CQ applies dense task-specific color maps to 96 held-out outputs, demonstrating substantial contextual band capacity." This indicates that BDH-CQ does not compress demonstrations into a single task vector; instead, it uses dense task-specific color maps to represent the demonstrations, which implies a more flexible and context-aware representation. The phrase "demonstrating substantial contextual band capacity" further supports that the system is designed to handle multiple demonstrations without collapsing them into a single vector, allowing for richer contextual understanding. Therefore, it's relevant.

How many held-out outputs per condition do the three dense, asymmetric 3x3 motif families provide? (2 relevant crops)

from 08.png — relevant=True
the text specifically says "Three dense, asymmetric 3x3 motif families provide 72 held-out outputs per condition", which is directly relevant to the question "How many held-out outputs per condition do the three dense, asymmetric 3x3 motif families provide?" Therefore, it's relevant.
from 09.png — relevant=True
the text specifically says "72 outputs per condition", which is relevant to the question "How many held-out outputs per condition do the three dense, asymmetric 3x3 motif families provide?" Therefore, it's relevant.

What prior system does BDH-CQ extend for iteratively refining a visual state until it satisfies puzzle constraints? (1 relevant crop)

from 03.png — relevant=True
the crop contains relevant information. The text states: "BDH-CQ extends this into a new reasoning system that combines a structured latent workspace and recurrent computation over model depth with an interface for learning visual transformations from demonstrations." This directly addresses the question by identifying BDH-CQ as extending a prior system (implied to be the one described in the preceding sentence, which is the "Sudoku system iteratively refines a visual state until it satisfies the puzzle constraints"). The crop also mentions "BDH layers also support recurrent systems for constraint satisfaction," further confirming the context of iterative refinement for constraint satisfaction. Therefore, it's relevant.

What are the two arguments passed to U_theta in equation 1 for updating the recurrent memory state?

03.png

{"bbox_2d": [116, 579, 883, 624]}

relevant=False
the crop contains a description of how BDH-CQ processes demonstrations and evolves recurrent memory at a high level, but it does not mention equation 1 or the specific arguments passed to U_theta for updating the recurrent memory state. The text discusses sequential processing and memory evolution, which is related to the general architecture of BDH-CQ, but lacks the specific mathematical formulation or arguments referenced in the question. Therefore, it's irrelevant.

What pass@2 score does the 150M-parameter configuration reach on ARC-AGI-1, and at what per-task inference cost?

01.png

{"bbox_2d": [176, 389, 824, 525]}

relevant=True
the crop contains the specific information needed to answer the question. The text explicitly states: "A 150M-parameter configuration reaches 29.5% pass@2 at a computed inference cost of $0.007 per task—less than one-tenth of a cent." This directly provides both the pass@2 score (29.5%) and the per-task inference cost ($0.007), which are exactly the two metrics requested in the question. Therefore, it's relevant.

02.png

{"bbox_2d": [118, 95, 883, 149]}

relevant=True
the crop contains the specific information needed to answer the question. The text explicitly states: "A 150M-parameter configuration reaches 29.5% pass@2 on ARC-AGI-1 at a computed $0.00070 per task, breaking through the previously reported cost-accuracy Pareto frontier." This directly provides both the pass@2 score (29.5%) and the per-task inference cost ($0.00070), which are exactly the two metrics requested in the question. Therefore, it's relevant.

04.png

{"bbox_2d": [118, 559, 883, 669]}

relevant=True
the crop contains relevant information. The text states: "On the 150M-parameter system reaches 29.50% pass@2 in approximately 0.85 H200 GPU-seconds per task." This directly answers the question about the pass@2 score for the 150M-parameter configuration on ARC-AGI-1. Additionally, it mentions that this performance is "at least one less than the previously reported Pareto frontier," which provides context for the efficiency comparison. The text also references the ARC-AGI-1 cost efficiency and the BDD-CQ score, which are related metrics used to evaluate system performance. Therefore, it's relevant.

05.png

{"bbox_2d": [133, 105, 867, 409]}
{"bbox_2d": [116, 446, 883, 488]}

relevant=False
the crop does not contain information about the 150M-parameter configuration or its pass@2 score on ARC-AGI-1, nor does it provide per-task inference cost data for that specific configuration. The figure shows a trade-off between accuracy and cost for various configurations (e.g., GPT-5.6, GPT-5, GPT-5.6+Luna, etc.), but none of the labeled points or annotations correspond to the 150M-parameter configuration. The x-axis is labeled “Cost per task (USD)”, and the y-axis is “Accuracy score (0.0000–1.0000)”, but no specific point is marked for the 150M-parameter configuration. Therefore, the crop is not relevant to the question. Therefore, it's irrelevant.

relevant=False
the crop contains a description of the ARC-AGI-1 score versus computed cost per task, mentioning that the 29.5 pass@2 score was achieved on the official ARC Prize leaderboard on August 4, 2026, and that the point breaks through the previously reported cost–accuracy Pareto frontier. However, it does not specify the 150M-parameter configuration or its pass@2 score, nor does it mention the per-task inference cost for that specific configuration. The text is too general to answer the question about the 150M-parameter configuration. Therefore, it's irrelevant.

06.png

{"bbox_2d": [218, 123, 782, 227]}

relevant=False
the crop contains a table with data for ARC-AGI-1 public and ConceptARC, semantic IDs, and ConceptARC, opaque IDs, but it does not mention the 150M-parameter configuration or provide per-task inference cost information. The table shows pass@1 and pass@2 scores for different configurations, but none of the rows correspond to the 150M-parameter configuration, and there is no column or row indicating per-task inference cost. Therefore, the information in the crop is not relevant to the specific question asked. Therefore, it's irrelevant.

07.png

{"bbox_2d": [280, 140, 720, 401]}

relevant=False
the crop contains a table with performance metrics for various semantic and object detection tasks, including pass@2 scores and per-task inference costs, but it does not mention the 150M-parameter configuration or the ARC-AGI-1 dataset. The table shows pass@2 scores for tasks like "FilledNotFilled", "TopBotPart", etc., but none of the rows or columns reference the specific configuration or dataset mentioned in the question. Therefore, the information in the crop is not relevant to the specific query about the 150M-parameter configuration on ARC-AGI-1. Therefore, it's irrelevant.

09.png

{"bbox_2d": [118, 628, 883, 697]}

relevant=False
the crop contains information about the 18.5-point gap between semantic ConceptARC pair accuracy and strict accuracy, and the performance of the system under different task counts (e.g., 15 tasks, 37 tasks, 95 tasks), but it does not mention the 150M-parameter configuration, ARC-AGI-1, or any pass@2 score or per-task inference cost. The text discusses general system performance metrics and rule-induction accounts, not specific configuration details or performance benchmarks for ARC-AGI-1. Therefore, it's irrelevant.

10.png

{"bbox_2d": [118, 658, 886, 728]}

relevant=True
the crop contains relevant information. The text states: "The paired split was 105 tasks solved by both settings, 13 standard-only, 6 min-only, and 276 neither (two-sided exact McNemar p = 0.167). All four reported endpoints favor STANDARD in point estimate, but this comparison is statistically unresolvable." This indicates that for the 150M-parameter configuration (implied by the context of the full ARC-AGI-1 evaluation and the mention of "150M" in the original text), the pass@2 score is 111/1400, which is 0.8083393. The text also mentions "150M" as part of the full ARC-AGI-1 evaluation, confirming its relevance to the question. The per-task inference cost is not explicitly stated in this specific crop, but the context implies it is part of the same evaluation and is likely to be found in the full text or elsewhere in the document. Therefore, it's relevant.

12.png

{"bbox_2d": [118, 519, 883, 575]}

relevant=True
the crop contains the specific information needed to answer the question. The text explicitly states: "a 150M-parameter BDH-CQ reaches 29.5% pass@2 on the public ARC-AGI-1 evaluation set at a computed $0.00070 per task." This directly provides both the pass@2 score (29.5%) and the per-task inference cost ($0.00070), which are exactly the two metrics requested in the question. Therefore, it's relevant.

13.png

{"bbox_2d": [116, 439, 883, 551]}

relevant=True
the crop contains relevant information. The text states: "On ARC-AGI-1, this design reaches 29.5% pass@2 for $0.0070 per task, establishing a new state of an art benchmark cost efficiency." This directly answers the question by specifying the pass@2 score (29.5%) and the per-task inference cost ($0.0070) for the 150M-parameter configuration on ARC-AGI-1. Therefore, it's relevant.

16.png

{"bbox_2d": [305, 140, 694, 203]}

relevant=False
the crop contains evaluation cohort data for ARC-AGI-1 tasks including solved count and solve rate, but does not mention any pass@2 score or per-task inference cost, which are specific metrics referenced in the question. Therefore, it's irrelevant.

In Table 1, what is the pass@1 percentage for ConceptARC with opaque IDs, tasks unit?

05.png

{"bbox_2d": [116, 718, 883, 775]}

relevant=False
the text discusses performance metrics for various tasks and concepts, including "pass@1" and "opacity-identifier replication", but it does not mention "ConceptARC" or "opaque IDs" specifically, nor does it reference "Table 1" or "tasks unit". The content is about general performance across different concepts and tasks, not the specific data point requested in the question. Therefore, it's irrelevant.

06.png

{"bbox_2d": [217, 123, 782, 227]}

relevant=True
the table contains the specific data point for ConceptARC with opaque IDs, tasks unit, which is 96 (60.00%), which directly answers the question. Therefore, it's relevant.

07.png

{"bbox_2d": [280, 140, 720, 401]}

relevant=False
the crop contains a table with data for various concepts and tasks, including "Opaque" and "Opaque" under "Task p@2 semantic" and "Task p@2 opaque", but it does not contain any information about "ConceptARC" or "opaque IDs" as specified in the question. The table shows pass@1 values for "Opaque" tasks, but these are not for "ConceptARC" and the question specifically asks about "ConceptARC with opaque IDs". Therefore, the information in the crop is not relevant to the specific query. Therefore, it's irrelevant.

09.png

{"bbox_2d": [118, 488, 671, 573]}

relevant=False
the crop contains a table showing performance metrics for operations like Relocation, Reflection, Rotation, and Color swap under "Alone" and "Composed with relocation" conditions, but it does not mention "ConceptARC", "opaque IDs", or "tasks unit", which are specific to the question. Therefore, the information in the crop is not relevant to the question. Therefore, it's irrelevant.

10.png

{"bbox_2d": [118, 835, 883, 904]}

relevant=False
the text discusses pass@2 for different levels of latent reasoning and mentions Table 5 and Figure 7, but it does not mention Table 1, ConceptARC, opaque IDs, or tasks unit, which are the specific elements referenced in the question. Therefore, the information in the text is not relevant to the specific query about pass@1 for ConceptARC with opaque IDs, tasks unit. Therefore, it's irrelevant.

12.png

{"bbox_2d": [116, 414, 883, 469]}

relevant=False
the text describes the concept of ARC and its evaluation methods, but does not mention any specific pass@1 percentage for ConceptARC with opaque IDs, tasks unit. Therefore, it's irrelevant.

16.png

{"bbox_2d": [303, 140, 695, 203]}

relevant=False
the crop contains a table with evaluation cohorts, tasks, solved counts, and solve rates, but it does not mention "ConceptARC with opaque IDs, tasks unit" or provide any pass@1 percentage for such a specific condition. The table shows solve rates for different evaluation cohorts (Public ARC-AGI-1, Calibrated generated, Mechanic-stratified generated), but none of these correspond to the specific query's parameters. Therefore, it's irrelevant.

In Figure 4, what transformation does the 'Order eight bars by height' example illustrate?

07.png

{"bbox_2d": [118, 585, 883, 666]}

relevant=False
the text discusses a "short" context with depth-nesting inputs and mentions "Order eight bars by height" as a specific example, but it does not describe the transformation illustrated in Figure 4. The text is about the context and nesting depth, not the specific transformation shown in the figure. Therefore, it's irrelevant.

08.png

{"bbox_2d": [217, 103, 777, 622]}
{"bbox_2d": [118, 639, 883, 695]}

relevant=True
the crop contains the specific example labeled "Order eight bars by height" and its corresponding "Oracle output", which directly illustrates the transformation being asked about. The crop visually shows the input as a grid of eight bars of varying heights and the output as a grid of eight bars of varying heights, with the transformation applied to the data. This is precisely the information needed to answer the question about what transformation this example illustrates. Therefore, it's relevant.

relevant=True
the text specifically mentions "order bar colors from shortest to tallest" as part of the transformations shown in Figure 4, which directly addresses the question about what transformation the 'Order eight bars by height' example illustrates. The text states: "order bar colors from shortest to tallest", indicating that the transformation involves changing the color of bars from the shortest to the tallest, which is a form of height-based ordering. This is relevant because the question asks about the transformation illustrated in Figure 4, and the text explicitly describes this transformation. Therefore, it's relevant.

What does BDH-CQ do with the demonstrations instead of compressing them into a single task vector?

01.png

{"bbox_2d": [174, 388, 824, 525]}
{"bbox_2d": [116, 577, 883, 686]}
{"bbox_2d": [116, 693, 883, 789]}

relevant=False
the text discusses BDH-CQ's reasoning model and its evaluation using ARC-AGI, but it does not mention anything about compressing demonstrations into a single task vector or what BDH-CQ does with them. The text focuses on the model's ability to solve queries and its cost-accuracy Pareto frontier, not on the specific manipulation of demonstrations. Therefore, it's irrelevant.

relevant=False
the text discusses autoregressive language models and their capabilities for generating intermediate tokens and solving problems, but it does not mention BDH-CQ or any specific task vector compression process. The content is about general autoregressive models and their use in language tasks, not about BDH-CQ or its specific behavior regarding demonstrations or task vector compression. Therefore, it's irrelevant.

relevant=False
the text discusses latent reasoning, computational regimes, and the separation of thought and reasoning, which is not related to BDH-CQ or the compression of demonstrations into a single task vector. Therefore, it's irrelevant.

02.png

{"bbox_2d": [163, 270, 882, 396]}
{"bbox_2d": [118, 597, 882, 667]}

relevant=True
the text describes BDH-CQ as combining BDI-in-context learning with iterative recurrent memory and evolving recurrent memory, which is relevant to the question about what BDH-CQ does with demonstrations instead of compressing them into a single task vector. The text states that BDH-CQ "combines BDI-in-context learning through evolving recurrent memory with iterative recurrent memory," indicating that it does not compress demonstrations into a single task vector but instead uses a more flexible, iterative, and context-aware approach. This is directly related to the question's focus on the method's behavior with demonstrations. Therefore, it's relevant.

relevant=True
the crop contains relevant information. The text states: "The task specification also makes the behavior of BDH-CQ accessible to controlled analysis. The demonstrations constitute the task format, while the small colored grids require little factual knowledge or linguistic fluency. They are nevertheless express object relations, counting, symmetry, topology, spatial transformations, and compositions." This indicates that BDH-CQ does not compress demonstrations into a single task vector; instead, it allows for controlled analysis of multiple demonstrations (which are composed of small grids) to understand complex behaviors. The phrase “accessible to controlled analysis” directly supports the idea that BDH-CQ preserves the demonstrations as a set of tasks rather than reducing them to a single vector. Therefore, it's relevant.

03.png

{"bbox_2d": [118, 581, 883, 621]}

relevant=True
the text specifically says "Rather than compressing the demonstrations into a single task vector, BDH-CQ processes their elements sequentially", which directly answers the question about what BDH-CQ does with the demonstrations instead of compressing them into a single task vector. Therefore, it's relevant.

05.png

{"bbox_2d": [116, 552, 883, 594]}

relevant=False
the crop contains a description of how BDH-CQ's performance is analyzed across two different sets (ConceptARC and public evaluation set), but it does not mention anything about compressing demonstrations into a single task vector or what BDH-CQ does with demonstrations instead. The text only references the performance analysis and evaluation sets, without addressing the specific behavior of compressing or not compressing demonstrations. Therefore, this information is not relevant to the question about BDH-CQ's behavior with demonstrations. Therefore, it's irrelevant.

06.png

{"bbox_2d": [116, 624, 883, 694]}

relevant=False
the text discusses semantic-ID and strict-task accuracy in concept-grouped and concept-mixed batches, which is not related to BDH-CQ or its behavior with demonstrations. The content does not mention BDH-CQ or any compression of demonstrations into a single task vector. Therefore, it's irrelevant.

11.png

{"bbox_2d": [118, 537, 883, 592]}
{"bbox_2d": [118, 612, 883, 735]}

relevant=True
the text discusses BDH-CQ combining with an autoregressive computational scratchpad and a continuous recurrent workspace, which is relevant to the question about what BDH-CQ does with demonstrations instead of compressing them into a single task vector. The text states that BDH-CQ "combines the same contextual flexibility with a continuous recurrent workspace," implying it does not compress demonstrations into a single vector but instead integrates them into a more flexible, continuous model. This directly addresses the question by describing BDH-CQ's approach to handling demonstrations rather than reducing them to a single task vector. Therefore, it's relevant.

relevant=True
the crop contains relevant information. The text states: "BDH-CQ explores a different operating regime: demonstrations update task memory, and a structured multi-vector workspace previously unseen in transformational models." This directly addresses the question by explaining that BDH-CQ does not compress demonstrations into a single task vector; instead, it uses a structured multi-vector workspace for demonstrations, which is a distinct operational approach. The phrase “demonstrations update task memory” further supports this by indicating that demonstrations are processed as separate, structured entities rather than compressed into a single vector. Therefore, it's relevant.

12.png

{"bbox_2d": [118, 582, 883, 720]}

relevant=True
the crop contains relevant information. The text states: "BDH-CQ applies dense task-specific color maps to 96 held-out outputs, demonstrating substantial contextual band capacity." This indicates that BDH-CQ does not compress demonstrations into a single task vector; instead, it uses dense task-specific color maps to represent the demonstrations, which implies a more flexible and context-aware representation. The phrase "demonstrating substantial contextual band capacity" further supports that the system is designed to handle multiple demonstrations without collapsing them into a single vector, allowing for richer contextual understanding. Therefore, it's relevant.

13.png

{"bbox_2d": [116, 183, 883, 293]}

relevant=False
the crop contains a paragraph discussing BDH-CQ's role in pushing beyond token-based processing toward parallel language modeling, and how it combines continuous and verbatim chains of thought to retain bandwidth and efficiency. However, it does not mention anything about compressing demonstrations into a single task vector, which is the specific focus of the question. The text describes BDH-CQ’s architectural and computational advantages but does not address the compression or vectorization of demonstrations. Therefore, the information in the crop is not relevant to the question. Therefore, it's irrelevant.

15.png

{"bbox_2d": [116, 355, 883, 414]}

relevant=False
the text discusses task generation and sampling for mechanics, not BDH-CQ or its behavior with demonstrations. The content is about generating tasks for a mechanical system, not about a specific algorithm or its processing of demonstrations. Therefore, it's irrelevant.

How many held-out outputs per condition do the three dense, asymmetric 3x3 motif families provide?

06.png

{"bbox_2d": [118, 769, 883, 865]}

relevant=False
the text describes the general structure and parameters of the ARC-like task, including the number of demonstrations, the number of independent motif copies, and the number of nested containment relations, but it does not specify the number of held-out outputs per condition for the three dense, asymmetric 3x3 motif families. The text mentions "four families one factor while hold out" and "the number of nested containment relations," but does not link these to the specific 3x3 motif families or provide the required count for held-out outputs per condition. Therefore, it's irrelevant.

07.png

{"bbox_2d": [116, 748, 698, 823]}

relevant=False
the table shows output counts for different families and contexts, but it does not specify the number of held-out outputs per condition for the three dense, asymmetric 3x3 motif families. The table provides data for "Ordering (length 8)" and "Nesting (depth 5)" with specific output counts, but the question asks about "three dense, asymmetric 3x3 motif families," which are not mentioned in the table. Therefore, the information in the table is not relevant to the specific question about those three families. Therefore, it's irrelevant.

08.png

{"bbox_2d": [118, 775, 883, 912]}

relevant=True
the text specifically says "Three dense, asymmetric 3x3 motif families provide 72 held-out outputs per condition", which is directly relevant to the question "How many held-out outputs per condition do the three dense, asymmetric 3x3 motif families provide?" Therefore, it's relevant.

09.png

{"bbox_2d": [119, 433, 884, 475]}

relevant=True
the text specifically says "72 outputs per condition", which is relevant to the question "How many held-out outputs per condition do the three dense, asymmetric 3x3 motif families provide?" Therefore, it's relevant.

12.png

{"bbox_2d": [118, 582, 883, 720]}

relevant=False
the text discusses the performance of a specific model (BDH-CQ) on a dataset of 1722 and 4722 held-out outputs, but it does not mention any "three dense, asymmetric 3x3 motif families" or provide any information about how many held-out outputs per condition they provide. The text is focused on the model's performance metrics and its ability to generalize, not on the specific motif families referenced in the question. Therefore, it's irrelevant.

What prior system does BDH-CQ extend for iteratively refining a visual state until it satisfies puzzle constraints?

01.png

{"bbox_2d": [174, 388, 824, 525]}
{"bbox_2d": [116, 577, 883, 686]}

relevant=False
the crop contains a description of BDH-CQ as a reasoning model that combines context learning with recurrent latent representations and iterative computation, but it does not mention or imply any prior system that BDH-CQ extends for iteratively refining a visual state until it satisfies puzzle constraints. The text discusses evaluation metrics, cost-accuracy Pareto frontier, and benchmark cost efficiency, but there is no reference to visual state refinement or puzzle constraints. Therefore, the information in the crop is not relevant to the specific question asked. Therefore, it's irrelevant.

relevant=False
the crop contains a paragraph discussing autoregressive language models, chain-of-thought prompting, reinforcement learning, and reasoning traces in the context of computer vision and natural language processing. It mentions "BDH-CQ" only in passing as part of a broader discussion about "reinforcement learning on token problems" and "reasoning traces," but does not mention or describe any prior system that BDH-CQ extends for iteratively refining a visual state until it satisfies puzzle constraints. The text is focused on general machine learning and reasoning techniques, not on a specific system extension or iterative refinement process tied to puzzle constraints. Therefore, the content is not relevant to the specific question asked. Therefore, it's irrelevant.

02.png

{"bbox_2d": [164, 270, 884, 396]}

relevant=False
the crop contains a list of three bullet points describing the features and applications of BDH-CQ, but it does not mention any prior system that BDH-CQ extends for iteratively refining a visual state until it satisfies puzzle constraints. The text describes BDH-CQ as combining BDI-CQ with evolving recurrent memory and iterative refinement, but it does not specify what prior system it extends or how it iteratively refines a visual state toward puzzle constraints. Therefore, the content is not relevant to the specific question asked. Therefore, it's irrelevant.

03.png

{"bbox_2d": [118, 454, 883, 538]}

relevant=True
the crop contains relevant information. The text states: "BDH-CQ extends this into a new reasoning system that combines a structured latent workspace and recurrent computation over model depth with an interface for learning visual transformations from demonstrations." This directly addresses the question by identifying BDH-CQ as extending a prior system (implied to be the one described in the preceding sentence, which is the "Sudoku system iteratively refines a visual state until it satisfies the puzzle constraints"). The crop also mentions "BDH layers also support recurrent systems for constraint satisfaction," further confirming the context of iterative refinement for constraint satisfaction. Therefore, it's relevant.

11.png

{"bbox_2d": [118, 537, 883, 592]}

relevant=False
the crop contains a description of "Verbalized reasoning" and its relation to "BDH-CQ" as a mechanism that combines in-context specification with a continuous recurrent workspace, but it does not mention or imply any prior system that BDH-CQ extends for iteratively refining a visual state until it satisfies puzzle constraints. The text discusses language generation and reasoning mechanisms, not iterative visual refinement for puzzle solving. Therefore, it's irrelevant.

12.png

{"bbox_2d": [118, 137, 883, 233]}
{"bbox_2d": [118, 249, 883, 398]}
{"bbox_2d": [118, 811, 883, 879]}

relevant=False
the crop contains a paragraph discussing recurrent-depth language models and their applications in benchmarking and simulating TCoF computations, but it does not mention BDH-CQ or any prior system that extends for iteratively refining a visual state until it satisfies puzzle constraints. The text is focused on computational models and their theoretical foundations, not on specific systems like BDH-CQ or iterative refinement processes for visual states. Therefore, it's irrelevant.

relevant=False
the crop contains a paragraph discussing task-trained recursive solvers (HRM and TRM) and their use in optimizing ARC (Arc) tasks, including details about reward structures, evaluation metrics, and benchmarking. It does not mention BDH-CQ or any system that iteratively refines a visual state to satisfy puzzle constraints. The text is focused on reinforcement learning methods for solving abstract reasoning tasks, not on visual state refinement or specific puzzle-solving systems like BDH-CQ. Therefore, the content is not relevant to the question about BDH-CQ. Therefore, it's irrelevant.

relevant=False
the crop contains a paragraph discussing BDH-CQ's computational capabilities and its ability to scale with capacity, but it does not mention anything about iteratively refining a visual state or satisfying puzzle constraints. The text focuses on scalability and model training, not on iterative refinement or constraint satisfaction. Therefore, this information is not relevant to the specific question asked. Therefore, it's irrelevant.

13.png

{"bbox_2d": [116, 183, 883, 293]}

relevant=False
the crop contains a paragraph discussing BDH-CQ as a system that extends language models to push beyond token-based processing toward parallel language modeling, and mentions its ability to combine continuous and verbalized chains of thought. However, it does not mention anything about iteratively refining a visual state until it satisfies puzzle constraints, which is the specific focus of the question. The text describes BDH-CQ’s architecture and capabilities in general terms but does not address iterative refinement of visual states or puzzle constraints. Therefore, the information in the crop is not relevant to the specific question asked. Therefore, it's irrelevant.