Free Databricks Context Engineer Associate Practice Exam
Try 45 original Context Engineer Associate questions with retrieval, memory, tool and agent-workflow evidence, answer explanations and a topic review worksheet.
Use this fixed 45-question set to practise context-engineering decisions across all seven domains. It contains 24 single-answer questions and 21 Select TWO questions; this mix is editorial, not a claimed Databricks exam ratio.
These are original IT Mastery practice questions, not official Databricks questions, copied live-exam content or exam dumps.
Start Question 1 · Review your attempt
How to use this free exam
- Allow 90 minutes. Record one answer unless the question says Select TWO.
- Award one point only for the exact correct answer or pair. Do not give partial credit for selecting one correct option in a Select TWO question.
- Mark guesses and compare every option’s explanation with the supplied evidence after finishing.
Use your own timer and the page’s question navigation. This page does not run a timer, record choices or calculate a score. Code blocks offer Wrap lines; wide tables scroll horizontally.
The official guide describes approximately 45 scored items, permits multiple selection and allows additional unscored items. This set uses the seven domain weights with whole-question rounding. It does not reproduce the official examination or predict a passing result.
Practice-set coverage
| Domain | Official range | Questions in this set |
|---|---|---|
| Foundations of Context Engineering | 16% | 7 |
| System Prompt and Instruction Design | 9% | 4 |
| Knowledge Retrieval and Genie Configuration | 20% | 9 |
| Memory Architecture with Lakebase and MLflow | 18% | 8 |
| Tool Design, MCP, and Agent Context | 13% | 6 |
| Context Compression and Compaction | 11% | 5 |
| Multi-Agent and Long-Horizon Task Design | 13% | 6 |
Practice questions
Questions 1-25
Question 1
Topic: Knowledge Retrieval and Genie Configuration
An engineer prepares policy passages for a Databricks AI Search index using a served BAAI/bge-large-en-v1.5 embedding model. The answer-generating model has a 32,768-token context window.
Deployment contract:
- The embedding input limit is 512 tokens, including two special tokens.
- Every input requires a 24-token source prefix identifying the document and section.
- The endpoint silently truncates over-limit inputs from the end before returning a successful embedding response.
Measured inputs: Counts use the embedding model’s tokenizer and include the required prefix and special tokens. Each heading-defined subsection covers a self-contained topic.
| Passage | Final input tokens |
|---|---|
| Whole section | 926 |
| Eligibility subsection | 286 |
| Approval subsection | 336 |
| Exceptions subsection | 356 |
Current ingestion submits the whole section and logs a successful embedding response.
Which conclusions are supported? Select TWO.
Options:
A. The successful embedding response establishes that the full section text contributed to the stored vector.
B. The three heading-defined subsections can be embedded separately while retaining the required source prefix.
C. Each 512-token body window fits the embedding limit after the required source prefix is added.
D. The generation model’s context window permits the intact section to be embedded without input truncation.
E. The current whole-section ingestion produces a vector from truncated input rather than from the complete section.
Correct answers: B and E
Explanation: Embedding limits apply to the embedding model’s input, independently of the model used to generate answers. The complete section requires 926 tokens, but this embedding model processes at most 512. Because the endpoint truncates from the end, a successful response can conceal omitted content. A larger generation context window does not prevent this loss.
Subdivision at the Eligibility, Approval, and Exceptions headings produces coherent passages of 286, 336, and 356 tokens. Each count already includes the required source prefix and special tokens, so all three fit while retaining source identification. By contrast, a 512-token body would require 538 tokens after that overhead is added. Chunk sizing must therefore account for the complete formatted embedding input, not just the passage body.
Why each option fits or fails:
A. The endpoint returns success after truncating over-limit inputs, so the status does not establish that all 926 tokens were processed.
B. The subsection inputs contain 286, 336, and 356 tokens including required overhead, all below the embedding model’s 512-token limit.
C. A 512-token body plus the 24-token source prefix and two special tokens requires 538 tokens, exceeding the embedding limit.
D. The generation model’s 32,768-token window does not change the embedding model’s separate 512-token input limit.
E. The 926-token whole-section input exceeds the 512-token limit, and the deployment contract specifies truncation from the end.
Question 2
Topic: Context Compression and Compaction
A team is evaluating a replacement context compactor for a Databricks-based purchasing agent. Workflows typically span six compaction cycles. The release goal is to preserve active purchasing requirements throughout the workflow.
Runtime contract: Each cycle receives the previous summary plus new messages. Discarded transcript content is not reloaded. The complete source transcript is available to the offline evaluation harness.
Observed MLflow 3 trace: These excerpts show tracked fields from one run. The approved price ceiling remained active, and the required compliance review was never completed or waived.
source: order=PO-714; ceiling=EUR 48/unit; compliance_review=pending
cycle_1: order=PO-714; ceiling=EUR 48/unit; compliance_review=pending
cycle_2: order=PO-714; ceiling=EUR 48/unit
cycle_3: order=PO-714; awaiting supplier quote
Which evaluation design best supports the release goal under this runtime contract?
Options:
A. Replay six independent first-cycle trials and check every resulting summary for active requirements recorded in the source transcript.
B. Replay six chained compaction cycles and check every resulting summary for high semantic similarity to that cycle’s input context.
C. Replay six history-restored compaction cycles and check every resulting summary for active requirements recorded in the source transcript.
D. Replay six chained compaction cycles and check every resulting summary for active requirements recorded in the source transcript.
Best answer: D
Explanation: Iterative compaction must be tested as a stateful process: later cycles consume earlier summaries, so an omission can persist or compound. Here, the first summary preserves both requirements. The second loses the pending compliance review, and the third also loses the price ceiling, despite neither requirement changing.
Replay representative six-cycle workflows through the production context path. After every cycle, compare retained task state with requirements established by the source transcript, accounting for legitimate updates. Check exact identifiers, values, units, and unresolved obligations rather than overall textual similarity. Cycle-level results expose where fidelity first degrades and whether the replacement compactor preserves required information through the full workflow.
Why each option fits or fails:
A. Independent first-cycle trials never reuse a compressed output, so they measure one-pass retention rather than losses accumulating across later cycles.
B. High semantic similarity can coexist with losing a short price limit or pending obligation, so it does not establish task-critical fidelity.
C. Restoring full history reintroduces previously discarded information, masking the cumulative loss that the production context path cannot repair.
D. Chained replay matches production, and source-based checks reveal whether still-active constraints and obligations disappear at any cycle.
Question 3
Topic: System Prompt and Instruction Design
A production Genie space receives competing reporting instructions.
Global instruction:
Round all monetary metrics to the nearest whole dollar.
Supplied description for cost_per_ticket_usd:
Cost per ticket is total cost divided by ticket count, displayed with two decimal places.
An attached, verified parameterized trusted query calculates per-ticket amounts from unrounded totals and returns:
| Query result | Unrounded value |
|---|---|
total_cost_usd | 1,004.40 |
ticket_count | 8 |
cost_per_ticket_usd | 125.55 |
refund_per_ticket_usd | 14.94 |
Genie currently displays cost per ticket as $126. The reporting owner confirms that cost per ticket is the only intended exception to the whole-dollar convention.
Which instruction revision and expected displays best resolve the conflict when validating a request for total cost, cost per ticket, and refunds per ticket together, in that order?
Options:
A. Apply global whole-dollar rounding to totals before deriving metrics, with two decimals in the description for cost per ticket only. Expect displays of $1,004, $125.50, and $15.
B. Make whole-dollar rounding the global final-display default, with two decimals in the description for cost per ticket only. Expect displays of $1,004, $125.55, and $15.
C. Make whole-dollar rounding the global final-display default, with two decimals in the descriptions for all per-ticket monetary metrics. Expect displays of $1,004, $125.55, and $14.94.
D. Keep global whole-dollar rounding as an initial formatting step, with two decimals in the description for cost per ticket only. Expect displays of $1,004, $126.00, and $15.
Best answer: B
Explanation: Presentation rounding should not change a trusted metric’s calculation. The global Genie instruction should establish whole-dollar formatting as the default for final monetary results and explicitly recognize metric-specific exceptions. The two-decimal exception belongs in the cost-per-ticket description, not in a broader rule covering every per-ticket measure.
Using the unrounded query results, total cost displays as $1,004, cost per ticket as $125.55, and refunds per ticket as $15. Validating all three together checks that Genie both honors the exception and preserves the global convention for neighboring metrics. Testing only cost per ticket could miss an exception that unintentionally changes other monetary displays.
Why each option fits or fails:
A. Dividing the rounded $1,004 total by 8 produces $125.50, changing the metric from the trusted query’s unrounded result of $125.55.
B. The narrowly scoped final-display exception preserves cost per ticket at $125.55 while total cost and refunds per ticket retain the whole-dollar convention.
C. Extending the exception to all per-ticket metrics changes refunds per ticket to $14.94 even though only cost per ticket is exempt from whole-dollar display.
D. Rounding $125.55 to $126 before applying two-decimal formatting destroys precision; displaying $126.00 does not satisfy the metric-specific reporting requirement.
Question 4
Topic: Foundations of Context Engineering
An assistant must retain users’ formatting preferences across new conversations and application restarts. A verified user’s earlier request for concise bullets was committed to a Lakebase-backed user_preferences table:
user_id: user_42
response_format: concise_bullets
Every trial below reads this same record successfully, including after a full application process restart. Each trial uses a fresh conversation with empty history and the same request, which contains no formatting instructions. Model weights and system instructions remain unchanged. Earlier interactions are retained as MLflow traces.
| Trial | Stored preference in inference context | Response format |
|---|---|---|
| Before restart | No | Paragraphs |
| After restart | No | Paragraphs |
| Replay after restart | Yes | Concise bullets |
Which conclusions about preference placement and initialization are supported? Select TWO.
Options:
A. New-conversation initialization needs to include the relevant stored preference in the context supplied to the model.
B. The preference must be encoded in model weights to remain available after an application process restart.
C. The Lakebase record provides durable user-preference state across new conversations and application process restarts.
D. The complete prior MLflow trace is required context for applying the preference in a new conversation.
E. The preference has session-only persistence and becomes unavailable when the original conversation’s message history is discarded.
Correct answers: A and C
Explanation: Durable preferences belong in application state that outlives individual conversations and processes. Here, the Lakebase-backed record survives both boundaries. The paragraph responses therefore indicate missing context integration, not lost storage.
Persistence and inference-time use are separate responsibilities. At new-conversation initialization, the application should retrieve the verified caller’s relevant preferences and include them in the model’s context. The controlled replay demonstrates this distinction: the same stored value changes the response format when included, with no model-weight change. Retaining MLflow traces supports diagnosis, but trace retention does not itself supply preference context to a new inference request.
Why each option fits or fails:
A. With persistence verified, changing only whether the stored preference enters inference context changes the observed response format.
B. The preference survives in Lakebase and affects the replay’s response even though the model weights remain unchanged.
C. Successful readback in fresh conversations and after restart shows that the preference remains available beyond the original session and process.
D. The replay applies the preference using only the stored value, demonstrating that the complete prior trace is unnecessary.
E. Fresh conversations have empty history, yet Lakebase still returns the preference, so its persistence is independent of session messages.
Question 5
Topic: Knowledge Retrieval and Genie Configuration
An engineer is preparing a shared Genie space for deployment. Embedded author credentials provide warehouse access, while Unity Catalog enforces data permissions using the signed-in end user’s identity. The author can read both attached, validated sources: sales_summary and margin_detail.
Representative requests:
- “What was total revenue for Q2 2026?” uses
sales_summary. - “What was gross margin by product for Q2 2026?” uses
margin_detail.
Intended end-user permissions:
| Role | Revenue request | Margin request |
|---|---|---|
| Sales analyst | Allowed | Denied |
| Finance analyst | Allowed | Allowed |
So far, only the author has tested these requests. Both answers matched approved reference results from frozen source snapshots. Those references and data-access traces are available for further testing.
Which release-validation plan best establishes both answer correctness and permission enforcement for the intended users?
Options:
A. Run Genie as each intended end user, compare allowed results with references, and inspect denied-query final responses for withheld restricted values.
B. Run Genie as the author with each role supplied in instructions, compare allowed results with references, and inspect denied-query responses for refusals.
C. Run Genie as each intended end user, compare allowed results with references, and inspect denied-query data-access evidence for enforced source restrictions.
D. Run correctness checks in Genie as the author, compare allowed results with references, and test source-access denials through separate SQL sessions for each role.
Best answer: C
Explanation: Warehouse connectivity and data authorization are separate boundaries in shared Genie. Embedded author credentials enable warehouse access; they do not grant end users the author’s Unity Catalog permissions.
Validation should exercise both representative requests under each intended end-user identity. Compare permitted answers with the approved references: sales analysts may retrieve revenue, while finance analysts may retrieve both revenue and margin. For the sales analyst’s margin request, inspect data-access evidence to verify the effective identity and that restricted source data cannot be returned. A refusal is acceptable behavior, but its wording alone does not prove enforcement.
Author-only demonstrations establish behavior for a privileged identity. Direct SQL permission tests are useful supporting evidence, but they do not replace end-to-end testing of the deployed Genie retrieval path.
Why each option fits or fails:
A. A final refusal can conceal restricted data already returned by a tool, so response checks alone do not establish data-access enforcement.
B. Naming a role in instructions does not change the author’s effective data identity, so these runs cannot exercise the sales analyst’s access restriction.
C. Actual end-user runs exercise the deployed retrieval path, while reference comparisons and data-access evidence establish correctness and enforced restrictions separately.
D. Separate SQL sessions can confirm grants, but author-only Genie runs do not validate propagation of end-user identity through the deployed retrieval path.
Question 6
Topic: Foundations of Context Engineering
A support agent must answer customers using the currently approved standard return window. An engineer compares two isolated runs that produced the same incorrect answer.
Both runs use the same Databricks AI Search index snapshot, which contains these accessible sources:
| Source | Policy status | Return window |
|---|---|---|
returns_v1 | Retired | 30 days |
returns_v2 | Current approved | 45 days |
Neither run has session history or memory. The displayed policy_context is the complete policy evidence supplied to the model; policy status is not included in that context.
Application trace:
Run A
retrieved_sources: []
policy_context: []
final_answer: "The return window is 30 days."
Run B
retrieved_sources: ["returns_v1"]
policy_context: "Standard returns are accepted within 30 days."
final_answer: "The return window is 30 days."
Which interpretation of the evidence should guide the engineer’s investigation?
Options:
A. Run A received no policy evidence; Run B received no policy evidence.
B. Run A received no policy evidence; Run B received conflicting policy evidence.
C. Run A received no policy evidence; Run B received outdated policy evidence.
D. Run A received outdated policy evidence; Run B received outdated policy evidence.
Best answer: C
Explanation: Diagnose evidence failures from what actually reached the model, not just its final answer or the contents of the source system. Identical incorrect outputs do not establish identical failure mechanisms.
Run A received no retrieved policy evidence. Its unsupported 30-day answer does not reveal where that value originated. Run B received a relevant-looking statement from a retired source, making its supplied evidence misleading for a current-policy request. Neither run received the approved 45-day statement, even though it was available in the index.
Investigate why relevant evidence failed to reach Run A and why a retired source was selected for Run B. Keep retrieval coverage and source freshness separate when assessing repairs.
Why each option fits or fails:
A. Run B received a relevant but outdated policy statement; missing the current revision is not the same as receiving no policy evidence.
B. Run B received only the retired statement; the current source’s presence in the index does not create conflicting statements in the model input.
C. Run A had an empty evidence context, while Run B received the retired 30-day statement rather than the currently approved 45-day policy.
D. Run A’s retrieved context was empty, so its 30-day answer alone does not establish that outdated policy evidence entered the model input.
Question 7
Topic: Knowledge Retrieval and Genie Configuration
At 10:05, an order-support agent must answer this customer request:
Tell me whether order O-482 is still on hold and what is required to release it.
Available evidence:
- The agent previously retrieved an AI Search order note written at 09:10. It recorded
ON_HOLD,ADDRESS_REVIEW, andpolicy_ref=AR-4. The index refreshed at 10:00, but the source note did not change. - A Genie space has an attached, Unity Catalog-governed order system-of-record table. A verified parameterized trusted query returns current committed
status,hold_code, andpolicy_ref. - AI Search contains the complete approved policy text for
AR-4andAR-5. All approved revisions were verified as indexed at 10:00, with no policy changes since then.
Applicability: O-482’s policy assignment is fixed at creation as AR-4. The newer AR-5 covers the same hold category but applies only to newer orders.
Which evidence-selection plan should the engineer use?
Options:
A. Read status from the indexed order note and retrieve the approved policy identified by the note’s
policy_ref.B. Read status from the current Genie query and retrieve the approved policy identified by the returned
policy_ref.C. Read status from the current Genie query and retrieve the newest approved policy for the returned hold category.
D. Read status from the indexed order note and retrieve the newest approved policy for the note’s hold category.
Best answer: B
Explanation: Current transactional state and descriptive policy evidence answer different parts of the request. Use the verified Genie query to establish the order’s current status, hold code, and assigned policy reference. Then retrieve the complete approved policy text matching that reference from AI Search to explain the release requirements.
Freshness depends on the underlying source, not merely the index synchronization time. Refreshing the index at 10:00 does not update a note written at 09:10, so that note cannot establish the status at 10:05. Policy applicability also differs from recency: AR-4 still governs O-482 even though AR-5 is newer. The validated policy index is sufficiently current here because no approved policy changes occurred after synchronization.
Why each option fits or fails:
A. The referenced policy is applicable, but refreshing an index does not make the unchanged 09:10 note evidence of the order’s 10:05 status.
B. The trusted query supplies current transactional state, while the matching approved policy supplies the release requirements that govern this order.
C. The current query provides valid status evidence, but AR-5 does not govern O-482 merely because it is the newest approved policy.
D. The unchanged order note may be stale, and the newest policy revision applies only to newer orders rather than O-482.
Question 8
Topic: Knowledge Retrieval and Genie Configuration
A support assistant uses Databricks AI Search to retrieve runbook RB-17. Its application tool queries ops.support.runbooks_index using the verified caller’s Unity Catalog identity. Endpoint and tool-level permissions are satisfied.
Business entitlement: The data owner has approved support_triage to query this index, but not its backing ingestion table, ops.support.runbooks_source. Caller nina@example.com belongs to support_triage and has no other applicable grants.
Normalized access check and trace:
effective_principal: nina@example.com
USE CATALOG ops: granted
USE SCHEMA ops.support: granted
SELECT ops.support.runbooks_index: absent
target: ops.support.runbooks_index
outcome: PERMISSION_DENIED
candidate_retrieval: not_started
retrieved_chunks: []
Validation: An authorized steward’s query retrieves RB-17 from the same index version.
Which conclusions are supported by this evidence? Select TWO.
Options:
A. The missing privilege is USE SCHEMA on
ops.support, so granting it tosupport_triageis the least-privilege correction.B. The missing privilege is SELECT on
ops.support.runbooks_source, so granting it tosupport_triageis the least-privilege correction.C. The missing privilege is SELECT on
ops.support.runbooks_index, so granting it tosupport_triageis the least-privilege correction.D. The empty chunk list reflects an authorization failure before candidate retrieval, not absence of
RB-17from the index.E. The empty chunk list reflects an indexing failure for
RB-17, so re-ingesting the runbook is required before retrying retrieval.
Correct answers: C and D
Explanation: Denied access and missing indexed content are different failure mechanisms. A permission check can reject a query before candidate retrieval begins, leaving an empty chunk list even when the requested content exists. Here, retrieval never starts, and the steward’s successful lookup confirms that RB-17 is present in the same index version.
Unity Catalog privileges on the parent catalog and schema do not themselves grant SELECT on the index. Because the support group’s business entitlement is approved and its namespace privileges are already effective, the least-privilege correction is to grant SELECT on ops.support.runbooks_index to support_triage. Granting access to the backing ingestion table would broaden access beyond that entitlement. After the grant, retry under the same verified caller identity to confirm retrieval succeeds.
Why each option fits or fails:
A. USE SCHEMA is already effective for the caller, and granting it again does not supply the missing index SELECT privilege.
B. Direct source-table access is outside the approved entitlement and does not replace the required SELECT privilege on the index.
C. The caller already has the required catalog and schema privileges; SELECT on the approved index is the missing Unity Catalog grant.
D. Candidate retrieval never started, and the steward retrieved RB-17 from the same index version, establishing that the requested content is indexed.
E. The successful steward lookup shows RB-17 is indexed; the caller’s denied query provides no basis for re-ingestion.
Question 9
Topic: Memory Architecture with Lakebase and MLflow
An engineering team is revising persistence for a Databricks agent with two durable workloads:
- Interaction state: Small updates to the active goal and unresolved actions are saved before acknowledging each turn and fetched by session ID on the next request. The state-write budget is p95 <= 100 ms.
- Analytical task records: Completed results and source references are batch-appended and retained across sessions. Analysts run daily joins and 90-day scans across millions of records while conversations are active. The current Delta reporting path meets its SLA.
Instrumented MLflow traces from production-sized tests show:
| State-write test | p95 latency |
|---|---|
| Delta | 620 ms |
| Lakebase, no report scans | 24 ms |
| Lakebase, concurrent report scans | 230 ms |
The team will retain the tested capacity and use a single Lakebase database.
Which durable-storage placement is best supported by these access patterns and measurements?
Options:
A. Store interaction state in Delta and analytical task records in Delta.
B. Store interaction state in Lakebase and analytical task records in Delta.
C. Store interaction state in Delta and analytical task records in Lakebase.
D. Store interaction state in Lakebase and analytical task records in Lakebase.
Best answer: B
Explanation: Durability and access pattern are separate concerns. Lakebase suits frequently updated, session-keyed operational state, while Delta-backed tables suit batch-appended results used for shared analytical queries.
Here, Lakebase state writes without report scans meet the 100 ms p95 budget; Delta state writes do not. Running analytical scans in the same Lakebase database also breaks that budget at the tested capacity. Keeping completed task records in the existing Delta reporting path preserves analytics while avoiding the measured operational contention.
These observations support separating the workloads, not a universal prohibition on analytical SQL in Lakebase. MLflow traces provide measurement evidence; they do not replace either durable state store.
Why each option fits or fails:
A. Delta supports the existing reporting workload, but its measured 620 ms state-write latency exceeds the 100 ms per-turn budget.
B. Lakebase state writes meet the latency budget when analytical scans remain in Delta, whose existing reporting path already meets its SLA.
C. Placing interaction state in Delta retains the measured state-write bottleneck while moving analytical records away from an already adequate reporting path.
D. Concurrent analytical scans in the shared Lakebase database raise state-write latency to 230 ms, exceeding the budget at the tested capacity.
Question 10
Topic: Knowledge Retrieval and Genie Configuration
A Genie space has both tables below attached, but their business context omits table-grain and relationship descriptions. Stores assign order numbers independently. All reporting-period rows are shown, and amounts are in USD.
An authorized user asks:
What is the total booked order revenue for each store?
Data snapshot:
orders
store_id order_id order_total
A 101 100
A 102 100
B 101 60
order_lines
store_id order_id line_id
A 101 1
A 101 2
A 102 1
B 101 1
B 101 2
Observed SQL:
SELECT o.store_id, SUM(o.order_total) AS revenue
FROM orders AS o
JOIN order_lines AS l
ON o.order_id = l.order_id
GROUP BY o.store_id;
Genie reports $500 for store A and $240 for store B.
Which facts should be added to the business context to address the inflated totals? Select TWO.
Options:
A.
order_totalmust contribute once per distinct monetary amount within each store when calculating booked revenue.B. The join must match both
store_idandorder_idto associate each line with its own order.C.
order_totalmust contribute once per order when calculating booked revenue, regardless of the number of matching lines.D. Matching both
store_idandorder_idcreates a one-to-one relationship between the order and line tables.E. Grouping by
store_idremoves repeated order contributions beforeSUM(order_total)computes revenue for each store.
Correct answers: B and C
Explanation: Table grain defines what one row represents. Here, orders contains one row per (store_id, order_id), while order_lines can contain several rows for that order. The header measure order_total must therefore contribute once per order.
The current join has two defects. Order number 101 appears in both stores, so matching only order_id links each header to four lines across the stores. Adding store_id removes those cross-store matches, but still leaves two matching lines for each store’s order 101. Summing header totals over that corrected join would still duplicate revenue.
Document both the composite relationship and the order-grain measure semantics in the Genie space. For revenue without a line-level condition, aggregate orders directly. If line data is required, preserve one contribution per order before summing. The correct totals are $200 for store A and $60 for store B.
Why each option fits or fails:
A. Store A has two distinct orders totaling $100 each; distinct monetary amounts would discard one valid order contribution.
B. Order numbers are local to a store, and 101 occurs in both stores, so joining on order_id alone mixes their lines.
C. An order’s header total repeats on each matching line row, so summing those rows counts the same revenue multiple times.
D. Both stores have two lines for order 101, so the correctly keyed order-to-line relationship remains one-to-many.
E. Grouping gathers all joined rows for each store; it does not remove repeated header totals contributed by multiple matching lines.
Question 11
Topic: Context Compression and Compaction
A rollout agent is verifying a Databricks AI Search index. It may mark the rollout ready only after the retrieved policy matches the validated source revision. Before its next verification pass, it must compact repetitive health-check history.
Task header already preserved:
- Index:
support_policies_idx - Policy ID:
R42 - Probe query:
refund eligibility
Availability health checks test whether the retrieval endpoint responds. The content probe compares the retrieved policy’s revision with the validated source revision.
Application check trace (UTC):
09:00 availability PASS
09:01 availability PASS
09:02 content_probe FAIL expected_revision=18 returned_revision=17
09:03 availability PASS
09:04 availability PASS
09:05 availability PASS
The full trace is retained for audit but will not be automatically loaded into the next model turn.
Which compaction best preserves the state needed for the next verification pass?
Options:
A. Summarize availability as healthy; retain the unresolved probe with expected revision 18 and returned revision 17 for the next content check.
B. Summarize availability as healthy; retain the expected and returned revisions, and mark the probe failure resolved by the subsequent successful health checks.
C. Summarize availability as healthy; retain the unresolved probe’s failure status, and leave the expected and returned revision values only in the audit trace.
D. Summarize availability as healthy; retain the probe failure as unresolved, and set the next probe’s expected revision to the observed revision 17.
Best answer: A
Explanation: Compaction should remove repetition without changing the meaning or resolution status of observations. Successful availability checks concern endpoint reachability; they cannot close a failed content-freshness probe.
The compact state should retain the index, policy ID and query, one availability summary, and the open exception: the validated source requires revision 18, but retrieval returned revision 17. Those values preserve the comparison needed for the next content check and prevent stale retrieved content from becoming the new expected baseline.
The audit trace preserves detailed history outside active context. It does not automatically supply that history to the next inference. Repeated successful checks can remain in the trace, while the unresolved exception and decision-relevant evidence remain in task context.
Why each option fits or fails:
A. This preserves the unresolved mismatch and its comparison values while removing repeated successes that add no new content-freshness evidence.
B. Later availability successes establish endpoint reachability, not current policy content, so they cannot resolve the earlier revision mismatch.
C. Keeping only the failure status removes the comparison values from the next turn, forcing the agent to recover evidence already available during compaction.
D. Revision 17 is the stale retrieved value; making it the expected value would erase the validated requirement to retrieve revision 18.
Question 12
Topic: Tool Design, MCP, and Agent Context
An inventory agent must select one MCP tool for a pilot. All four candidates have compatible read-only contracts, equally authoritative sources, and independently verified access for the caller.
Selection rule: Require end-to-end p95 latency <= 850 ms and maximum observed source age <= 3.0 seconds at final-response time. Among eligible tools, minimize end-to-end p95 latency.
Benchmark conditions: Setup includes discovery and authentication and is a fixed per-request cost before lookup. Response assembly takes another fixed 250 ms after result receipt. No stages overlap. Source ages below are maximum observed ages at result receipt under the same workload.
| MCP tool | Setup (ms) | Lookup p95 (ms) | Source age (s) |
|---|---|---|---|
managed_stock | 180 | 220 | 1.4 |
external_stock | 40 | 140 | 2.8 |
custom_cache | 60 | 300 | 2.6 |
custom_live | 80 | 460 | 0.4 |
Based on these measurements, which tool should the agent select?
Options:
A. Route inventory lookups to
custom_cache.B. Route inventory lookups to
external_stock.C. Route inventory lookups to
managed_stock.D. Route inventory lookups to
custom_live.
Best answer: A
Explanation: Tool selection should use the complete measured execution path and freshness at the point of use, not server category or lookup duration alone. Fixed serial setup and assembly costs shift the lookup’s p95 by their combined duration. Assembly also increases source age after receipt.
For custom_cache, end-to-end p95 is \(60 + 300 + 250 = 610\) ms, and maximum source age at final-response time is \(2.6 + 0.25 = 2.85\) seconds. It meets both pilot limits and has the lowest latency among eligible tools. Managed, external, and custom MCP hosting do not establish a universal performance or freshness ranking.
Why each option fits or fails:
A. Its 610 ms end-to-end p95 and 2.85-second source age meet both limits, and it has the lowest latency among eligible tools.
B. Its 430 ms end-to-end p95 is fastest, but response assembly increases its maximum source age to 3.05 seconds, exceeding the freshness limit.
C. Its 650 ms end-to-end p95 and 1.65-second source age meet both limits, but its latency exceeds the feasible minimum of 610 ms.
D. Its 0.65-second source age meets the freshness limit, but its 790 ms end-to-end p95 does not minimize latency among eligible tools.
Question 13
Topic: Multi-Agent and Long-Horizon Task Design
An orchestrator is building a procurement recommendation and stores task checkpoints in Lakebase. Each worker receives only its listed inputs at dispatch. Completed outputs are not automatically recomputed when a source changes. Restarting a worker supersedes its previous dispatch and rejects late results from it.
Current checkpoint:
| Step | State | Inputs |
|---|---|---|
| Demand snapshot | Complete | Forecast snapshot F1 |
| Supplier eligibility | Complete | Supplier snapshot S1; policy v4 |
| Purchase allocation | Complete | Demand snapshot; supplier eligibility |
| Recommendation | Running | Purchase allocation |
Before the recommendation finishes, the policy owner withdraws v4 and makes v5 effective. Policy v5 disallows a supplier selected in the completed purchase allocation. Forecast snapshot F1 and supplier snapshot S1 remain unchanged and valid.
Which recovery action should the orchestrator take to produce a recommendation valid under policy v5?
Options:
A. Update the checkpoint to policy v5, rerun supplier eligibility followed by purchase allocation, and restart the recommendation worker while retaining the completed demand snapshot.
B. Update the checkpoint to policy v5, rerun supplier eligibility, and restart the recommendation worker while retaining the completed purchase allocation and demand snapshot.
C. Update the checkpoint to policy v5, restart the recommendation worker, and retain the completed supplier eligibility, purchase allocation, and demand snapshot.
D. Update the checkpoint to policy v5, rerun purchase allocation, and restart the recommendation worker while retaining the completed supplier eligibility and demand snapshot.
Best answer: A
Explanation: A material source change can invalidate both its direct consumers and downstream work derived from their outputs. Policy v5 directly affects supplier eligibility. The purchase allocation depends on that eligibility, and the running recommendation depends on the allocation. Recompute this chain in dependency order and supersede the recommendation’s old dispatch so its late result cannot replace the refreshed result.
The demand snapshot depends only on unchanged forecast F1, so it can be retained. Supplier snapshot S1 can also be reused when eligibility is recalculated.
Refresh the durable checkpoint with the effective policy version, affected output status, replacement outputs, and current dispatch status. Updating a source reference alone does not repair stale derived context. Resume from the earliest affected dependency rather than continuing the original plan blindly or replaying unaffected work.
Why each option fits or fails:
A. The policy change invalidates supplier eligibility and its downstream allocation and recommendation, while the demand snapshot has no dependency on that policy.
B. The retained purchase allocation still selects a supplier prohibited by v5; recalculating eligibility does not retroactively change that completed allocation.
C. Changing the checkpoint’s policy reference leaves the completed outputs unchanged, so the restarted recommendation still receives an allocation containing a prohibited supplier.
D. Purchase allocation still receives eligibility computed under withdrawn policy v4, so the changed policy has not reached the allocation worker’s inputs.
Question 14
Topic: Foundations of Context Engineering
A self-service agent must report the verified caller’s on-call team for the week beginning October 5, 2026.
Execution contract: The owner filter is applied before candidate retrieval. An independent Unity Catalog check authorizes index access using the verified caller, who may read both schedules shown below.
Application trace: This simplified trace includes the binding rule and current records from the queried AI Search index.
verified_caller.user_id = u42
task.subject_user_id = u42
body.user_id = u17 (unverified)
session.context_user_id = u17
context_user_id = session.context_user_id or verified_caller.user_id
uc.effective_caller = u42
uc.index_access = ALLOW
ai_search.filter = owner_user_id:u17, week:2026-10-05
ai_search.result = owner_user_id:u17, team:Atlas
answer = "Your on-call team is Atlas."
index_record = owner_user_id:u17, week:2026-10-05, team:Atlas
index_record = owner_user_id:u42, week:2026-10-05, team:Birch
Which conclusions about the cause and repair are supported? Select TWO.
Options:
A. Deriving
context_user_idfrom the request body’s user identifier fixes the binding defect while retaining the independent Unity Catalog check.B. Reranking the returned candidates can recover the verified caller’s schedule while retaining the existing
context_user_idbinding.C. The successful Unity Catalog access check confirms that the retrieval filter selected the verified caller’s schedule.
D. The restored session identity is the earliest supported cause of both the wrong retrieval and the wrong on-call conclusion.
E. Deriving
context_user_idfrom the verified caller on each request fixes the binding defect while retaining the independent Unity Catalog check.
Correct answers: D and E
Explanation: The identity-binding step gives a restored session value precedence over the current authenticated identity. Because session.context_user_id is u17, retrieval targets u17 even though the verified caller and validated task identify u42. The Atlas answer follows the returned record, but that record concerns the wrong person. Both symptoms originate in the earlier binding defect.
For this self-service task, derive context_user_id from the verified caller on every request. Restored session values are historical context, not identity authority. Keep the independent Unity Catalog check using the verified caller.
Authorization and retrieval relevance are separate: permission to read a schedule does not establish that it belongs to the requested subject. Reranking cannot recover the Birch record after the owner filter excludes it from candidate retrieval.
Why each option fits or fails:
A. The body identifier is unverified and still equals u17, so this replacement preserves the wrong subject and does not establish trusted caller identity.
B. The owner_user_id = u17 filter excludes u42 before retrieval, so reranking the returned candidates cannot recover the Birch schedule.
C. The access check allows u42 to read both schedules; it does not certify that the owner filter matches the requested subject.
D. The rule selects u17 from the restored session, so retrieval and the final answer concern that user instead of verified caller u42.
E. The verified caller and validated task both identify u42; using that trusted identity removes the stale binding without changing access enforcement.
Question 15
Topic: Knowledge Retrieval and Genie Configuration
An engineer is compacting passages returned by Databricks AI Search for a support agent. The goal is to remove repeated packaging while retaining the policy evidence needed for this request:
Summarize storage, retention, and approval requirements for US and EU support exports, including how EU approvals must be linked.
All three passages come from different sections of the same approved, current policy:
[P1] Support Export Controls | revision 6
All support exports must be encrypted at rest.
US exports: retain 30 days; data-owner approval before execution.
[P2] Support Export Controls | revision 6
All support exports must be encrypted at rest.
EU exports: retain 7 days; data-owner and privacy-officer approval before execution.
[P3] Support Export Controls | revision 6
All support exports must be encrypted at rest.
EU data-owner and privacy-officer approvals must use the same request ID.
Which consolidated context best satisfies this goal?
Options:
A. Encrypt support exports at rest and obtain approvals before execution. US: 7-day retention, data-owner approval. EU: 7-day retention, data-owner and privacy-officer approval. EU approvals share a request ID. Sources: P1-P3.
B. Encrypt support exports at rest and obtain approvals before execution. US: 30-day retention, data-owner approval. EU: 7-day retention, data-owner and privacy-officer approval. EU approvals share a request ID. Sources: P1-P3.
C. Encrypt support exports at rest and obtain approvals before execution. US: 30-day retention, data-owner and privacy-officer approval. EU: 7-day retention, data-owner and privacy-officer approval. EU approvals share a request ID. Sources: P1-P3.
D. Encrypt support exports at rest and obtain approvals before execution. US: 30-day retention, data-owner approval. EU: 30-day retention, data-owner and privacy-officer approval. EU approvals share a request ID. Sources: P1-P3.
Best answer: B
Explanation: Context compaction removes redundant representation, not meaningful distinctions. A shared header and revision identify related material, but do not make the passages interchangeable. Here, the repeated header can be omitted and the encryption requirement stated once.
The consolidated context must preserve the regional qualifiers: US exports require 30-day retention and data-owner approval, while EU exports require seven-day retention and both approval roles. Approvals precede execution, and the EU approvals must share a request ID. Retaining the source references supports verification of these claims.
Different scoped requirements are valid source variations, not conflicts to resolve by selecting one passage or applying a single rule globally.
Why each option fits or fails:
A. Applying the EU retention period to US exports replaces a scoped policy fact rather than removing duplicated wording.
B. The shared encryption rule appears once, while regional retention periods, pre-execution approval requirements, and EU request-ID linkage remain intact.
C. The passages assign privacy-officer approval to EU exports; adding it to the US rule introduces an unsupported requirement.
D. The 30-day retention period applies to US exports; extending it to EU exports discards the distinct seven-day requirement.
Question 16
Topic: Context Compression and Compaction
An agent compacts a contract review into a handoff for another agent. The handoff must remain small while allowing later verification of the evidence supporting the original conclusion.
Evidence before compaction:
renewal_policy, revisionr7, section4.2: Allows cancellation within 30 days after renewal.orion_addendum, revisionr2, section2.1: Waives Orion’s cancellation fee.
Compacted handoff:
Orion may cancel within 30 days after renewal with no cancellation fee.
A reviewer disputes this conclusion after both documents receive newer revisions. The application retains immutable document revisions and can fetch individual sections. Its Databricks AI Search index contains only current revisions. The reviewer is authorized to access both original and current sources.
The goal is to verify the original evidence, not reassess current contract terms. Which change to the compacted handoff best supports this goal?
Options:
A. Retain each clause’s quoted supporting text, document title, and review timestamp; locate matching current sections when verification is requested.
B. Retain each clause’s document ID, section ID, and review timestamp; retrieve the corresponding current sections when verification is requested.
C. Retain each clause’s document ID, original revision ID, and section ID; retrieve the corresponding archived sections when verification is requested.
D. Retain each clause’s retrieval query, search filters, and result ranks; rerun the recorded search when verification is requested.
Best answer: C
Explanation: Compaction should preserve a claim-to-evidence map, not just the conclusion. Separate links are needed for Orion’s cancellation window and fee waiver because they came from different documents. Each link should identify the document, the immutable revision actually consulted, and the supporting section.
Those archived sections can then be fetched only when a reviewer challenges a clause, keeping the handoff small while preserving auditability. The compacted state does not need every original passage if its evidence references remain resolvable. Verifying historical support is distinct from checking whether the conclusion remains valid under newer contract terms.
Why each option fits or fails:
A. Quotes and titles preserve wording but do not identify the archived revisions; searching current sections cannot reliably recover the historical source linkage.
B. A review timestamp does not pin a document revision, and current sections can contain terms different from those originally consulted.
C. These identifiers link both clauses to the exact archived sections used, enabling targeted verification without reloading all original passages.
D. The current-only index cannot recreate evidence from superseded revisions, even when the original search query, filters, and ranks are retained.
Question 17
Topic: Foundations of Context Engineering
An agent is asked:
What is the maximum operating temperature of SKU
TS-410-B?
A custom retrieval service reads a current, approved Unity Catalog table with one record per SKU. It supports exact-SKU filters, product-family filters, semantic ranking, and field projection. Every route enforces the verified end user’s Unity Catalog permissions.
Observed tool response excerpt:
Returned: 2,400 products; 31,600 tokens
TS-410-A: 60 C; family=ThermalSense; source=catalog:r18:TS-410-A
TS-410-B: 85 C; family=ThermalSense; source=catalog:r18:TS-410-B
The full response fits the model’s context window. Across repeated tests, the agent sometimes reports 60 C and cites TS-410-A.
The engineer wants to minimize irrelevant context while preserving the requested product’s evidence. Every proposed change retains SKU, temperature, unit, and source reference. Which retrieval change best meets this objective?
Options:
A. Apply a product-family filter and add every matching record, projected to the retained fields.
B. Project the full catalog to the retained fields and add all resulting records to context.
C. Apply an exact-SKU filter and add the matching record, projected to the retained fields.
D. Semantically rank the catalog and add the highest-ranked record, projected to the retained fields.
Best answer: C
Explanation: A narrow product lookup should retrieve the smallest authoritative slice that answers the request. The catalog already contains the required evidence, but the full response introduces unrelated products and easily confused variants. Fitting within the context window does not guarantee reliable use of that evidence.
Filter on the exact SKU TS-410-B, then project its identifier, maximum operating temperature, unit, and source reference. This supplies 85 C with provenance catalog:r18:TS-410-B, preserving the connection between the requested product and its specification.
Filtering reduces irrelevant records; projection reduces irrelevant fields. Both help here. Semantic similarity can support discovery, but it does not replace an exact identifier constraint when the requested SKU is already known.
Why each option fits or fails:
A. Family filtering retains both TS-410-A and TS-410-B, leaving unnecessary variants and the ambiguity already observed in the answer.
B. Projection removes unnecessary attributes but still leaves thousands of unrelated products, including the variant that supplied the incorrect temperature.
C. Exact filtering selects TS-410-B, while projection preserves its temperature, unit, identifier, and source reference without including unrelated products.
D. Semantic ranking selects by similarity rather than enforcing SKU equality, so a closely related variant can still supply the temperature.
Question 18
Topic: System Prompt and Instruction Design
An engineer is reducing few-shot examples for a policy-answering agent. The baseline contains examples E1 through E4. Examples E3 and E4 use nearly identical policy-question wording.
Evaluation contract:
- Each behavioral slice contains 40 held-out cases. Acceptance requires at least 39 passing cases in every slice in every run.
- The slices test abstention when supporting evidence is absent, clarification when the region is missing, and selection of the authoritative current source.
- Cases, source snapshots, system instructions, model settings, and the relative order of retained examples remain fixed.
- Each table cell reports the minimum passing-case count across three repeated runs.
Example lengths: E1: 220 tokens; E2: 90 tokens; E3: 140 tokens; E4: 140 tokens.
| Examples omitted | Abstention | Clarification | Source selection |
|---|---|---|---|
| None | 40 | 40 | 40 |
| E1 | 40 | 40 | 40 |
| E2 | 40 | 40 | 40 |
| E3 | 40 | 29 | 40 |
| E4 | 40 | 40 | 28 |
| E1 and E2 | 31 | 40 | 40 |
Which conclusions are supported by the measured behavior and token costs? Select TWO.
Options:
A. Both E1 and E2 are individually necessary for the abstention slice because jointly omitting them lowers that slice’s score.
B. The set E3, E4 meets the acceptance rule because separately omitting E1 or E2 leaves all measured scores unchanged.
C. The set E2, E3, E4 is the lowest-token tested configuration that meets the acceptance rule for all three slices.
D. The set E1, E2, E4 meets the acceptance rule because perfect abstention and source-selection results compensate for weaker clarification.
E. E3 and E4 have distinct measured contributions despite similar wording; their omissions reduce clarification and source-selection performance, respectively.
Correct answers: C and E
Explanation: Ablation measures an example’s contribution within a particular configuration. Surface similarity alone does not establish redundant behavioral coverage.
Removing either E1 or E2 individually preserves the measured behaviors, but removing both reduces abstention below the acceptance threshold. Individual omission results therefore cannot be combined as though examples contributed independently.
The tested retained set E2, E3, E4 passes every slice and uses 370 example tokens. Retaining E1, E3, E4 also passes but uses 500 tokens; the baseline uses 590. Thus, removing E1 produces the smallest tested configuration that qualifies.
E3 and E4 demonstrate different contributions: their omissions harm clarification and source selection, respectively. Their similar wording is not sufficient grounds for treating them as interchangeable. These conclusions apply to the evaluated cases and configurations, not to every possible deployment.
Why each option fits or fails:
A. Each individual omission preserves 40 passing abstention cases, so neither example is individually necessary while the other remains.
B. Jointly omitting E1 and E2 reduces abstention to 31 passing cases; unchanged individual results do not justify their simultaneous removal.
C. Removing E1 preserves 40 passing cases in every slice and leaves 370 example tokens, fewer than any other passing tested configuration.
D. Omitting E3 leaves only 29 passing clarification cases; acceptance requires at least 39 in each slice, so other scores cannot compensate.
E. Omitting E3 lowers clarification to 29 passing cases, whereas omitting E4 lowers source selection to 28, demonstrating different behavioral contributions.
Question 19
Topic: Multi-Agent and Long-Horizon Task Design
A coordinator is assembling a procurement review from three independent worker subtasks. It may finalize the review only after all three subtasks complete successfully.
Application return contract:
statusis the validated execution state;summaryis descriptive narrative.- A non-null
result_refidentifies validated final evidence for that subtask. - A
checkpoint_refidentifies validated progress that can be reused when resuming.
Latest worker returns:
[
{"task":"spend_totals","status":"completed","result_ref":"totals_v4",
"summary":"Totals ready."},
{"task":"policy_review","status":"partial","checkpoint_ref":"policy_v4",
"remaining":["review addendum"],"summary":"Purchases appear compliant."},
{"task":"exception_scan","status":"failed","result_ref":null,
"error":"query_timeout","summary":"No exceptions expected."}
]
Which conclusions are supported by these returns? Select TWO.
Options:
A. The combined worker summaries are sufficient to mark the procurement review as completed.
B. The unfinished addendum step makes the
policy_reviewcheckpoint unsuitable for further continuation.C. The
spend_totalsresult can satisfy its assigned prerequisite for the procurement review.D. The
exception_scansummary establishes a completed search with a verified no-exceptions finding.E. The
policy_reviewcheckpoint is reusable progress toward a still-unfinished prerequisite for the review.
Correct answers: C and E
Explanation: A coordinator should track completion separately from narrative content. Here, status supplies execution state, while result and checkpoint references identify different kinds of evidence. spend_totals has completed and supplies validated final evidence, satisfying its prerequisite. policy_review still requires an addendum review; its checkpoint preserves reusable progress but does not close the prerequisite. exception_scan timed out without a result, so its optimistic summary cannot establish a successful search or a verified absence of exceptions.
The coordinator should retain these states, evidence references, and outstanding obligations independently. The overall review remains unfinished until the remaining work completes successfully. A plausible narrative cannot promote partial or failed execution to completion.
Why each option fits or fails:
A. Finalization requires three successful completions, but the policy review is partial and the exception scan has failed.
B. The checkpoint is explicitly validated and reusable; an outstanding step makes the subtask incomplete rather than invalidating its saved progress.
C. Its completed status and validated result reference satisfy the totals prerequisite, even though the accompanying narrative is brief.
D. A timed-out scan without a result provides no verified finding; an expectation of no exceptions is not a successful empty search.
E. The partial status and outstanding addendum leave the prerequisite open, while the contract permits resumption from validated checkpoint progress.
Question 20
Topic: Multi-Agent and Long-Horizon Task Design
An AI-agent coordinator is preparing a validated month-end close report. Finalization depends on extracting 12 statements, reconciling all 12, and validating them against the current policy registry.
Worker reports:
- Extraction: All 12 statements were extracted, and the complete manifest was saved.
- Reconciliation: Eight statements were reconciled with no discrepancies found. A partial comparison table was saved; four statements remain unprocessed.
- Policy validation: The registry lookup timed out before any checks ran. A diagnostic note describing the intended checks was saved.
All worker runs have terminated, all response messages were received, and all returned artifacts are accessible. The coordinator currently starts finalization whenever those conditions hold.
Which change to the worker-return contract and finalization gate best prevents unfinished work from being treated as complete?
Options:
A. Return findings status separately from narrative detail; finalize when every prerequisite worker reports
no_confirmed_issue.B. Return worker-execution status separately from narrative detail; finalize when every prerequisite worker reports
terminated.C. Return artifact-availability status separately from narrative detail; finalize when every prerequisite worker reports
available.D. Return subtask-completion status separately from narrative detail; finalize when every prerequisite subtask reports
completed.
Best answer: D
Explanation: Subtask completion describes the assigned work, not whether a worker stopped running or returned a useful message. A worker-return contract should expose a machine-readable status such as completed, partial, or failed, with progress, errors, and artifact references kept separately.
Here, extraction is completed, reconciliation is partial, and current-policy validation is failed. Finalization remains blocked because two prerequisites are unfinished. The coordinator can use the narrative to schedule reconciliation of the remaining four statements and recovery of the registry lookup. Those details support recovery but do not establish completion. Once the unfinished work succeeds, its status can be updated and the dependency gate reevaluated.
Why each option fits or fails:
A. The absence of confirmed issues does not establish that the remaining four statements or the current-policy requirements were checked.
B. Termination means the worker stopped running, not that reconciliation and policy validation finished successfully.
C. The partial comparison table and diagnostic note are available artifacts, but neither demonstrates completion of its assigned prerequisite subtask.
D. Extraction is completed, reconciliation is partial, and policy validation failed, so explicit subtask statuses keep the dependent finalization step blocked.
Question 21
Topic: Tool Design, MCP, and Agent Context
An agent discovers Agent Skills using only their names and one-line descriptions. It loads a skill’s detailed procedure only after selecting that skill. Five synchronization skills currently share the description “Repair synchronization data.”
Verified procedure contracts:
| Skill | Intended operation | Prerequisites |
|---|---|---|
sync_backfill | Fill historical range; preserve live watermark | Bounded dates and retained source records |
sync_resume | Continue interrupted live synchronization | Valid persisted checkpoint for that run |
sync_retry | Retry failed operations only | Stored per-operation outcomes from failed run |
sync_rebuild | Replace target with complete snapshot | Complete source snapshot and replacement approval |
sync_preview | Preview checkpoint-based changes; no writes | Valid checkpoint and readable change feed |
Which proposed descriptions accurately distinguish their skills’ intent, scope, and prerequisites for discovery? Select TWO.
Options:
A.
sync_resume: “Continue an interrupted live synchronization from that run’s valid persisted checkpoint.”B.
sync_preview: “Preview changes without writes using the available change feed when the checkpoint has expired.”C.
sync_backfill: “Backfill a bounded historical range from retained source records while preserving the live watermark.”D.
sync_rebuild: “Replace the target with retained incremental changes after obtaining replacement approval.”E.
sync_retry: “Retry a failed synchronization run by replaying its recorded successful and failed operations.”
Correct answers: A and C
Explanation: Reliable skill discovery requires concise metadata that identifies the intended task, applicability conditions, and important scope boundaries. A shared description such as “Repair synchronization data” hides distinctions that the selector needs before loading a procedure.
Historical backfill applies to a bounded range with retained source records and preserves the live watermark. Resumption instead continues an interrupted live run using that run’s valid checkpoint. Those distinctions belong in discovery descriptions, while detailed steps and implementation guidance can remain in the conditionally loaded procedure bodies.
This is progressive disclosure: provide enough information to select the relevant skill, then load its instructions. Descriptions guide selection; they do not replace execution-time prerequisite validation or authorization.
Why each option fits or fails:
A. This description identifies interrupted-run recovery and its required checkpoint, distinguishing resumption from historical backfill or full target replacement.
B. The preview contract requires a valid checkpoint, so an expired checkpoint cannot support the advertised checkpoint-based change preview.
C. This description captures the historical-range intent, retained-source prerequisite, and watermark boundary needed to distinguish backfill from live synchronization recovery.
D. The rebuild procedure requires a complete source snapshot; replacement approval does not make an incremental-only source sufficient.
E. The retry contract covers failed operations only; replaying successful operations expands the procedure beyond its verified scope.
Question 22
Topic: Knowledge Retrieval and Genie Configuration
An engineer is curating business context for a Genie space with orders and order_lines attached. Existing descriptions define columns but omit table grain and aggregation guidance.
A user requests:
Return total revenue for orders containing at least one Hardware line.
Revenue means the full orders.order_total for each qualifying order.
Complete source snapshot:
orders:
| order_id | order_total (USD) |
|---|---|
| 4101 | 120.00 |
| 4102 | 120.00 |
order_lines:
| order_id | line_no | product_family |
|---|---|---|
| 4101 | 1 | Hardware |
| 4101 | 2 | Hardware |
| 4102 | 1 | Hardware |
| 4102 | 2 | Service |
Observed SQL:
SELECT SUM(o.order_total) AS revenue
FROM orders AS o
JOIN order_lines AS l ON o.order_id = l.order_id
WHERE l.product_family = 'Hardware';
Genie reports $360. The engineer must preserve the eligibility rule and revenue definition. Which update to the table and relationship guidance best addresses the error?
Options:
A. Grouping joined rows by order restores the revenue grain; sum matching
order_totalvalues per order, then add those subtotals.B.
order_totalvalues identify duplicate revenue after a join; sum each distinct amount once across the qualifying orders.C. The joined order-line pair is the revenue grain; sum
order_totalonce per distinct (order_id,line_no) pair among qualifying lines.D. The order remains the revenue grain in a one-to-many join; use qualifying lines to select orders, then sum
order_totalonce per order.
Best answer: D
Explanation: Table grain states what one row represents. orders has one row per order_id, while order_lines has one row per (order_id, line_no). Their one-to-many join repeats order-level measures for every matching line.
Order 4101 has two Hardware lines and order 4102 has one, creating three $120 contributions and a reported total of $360. The requested metric includes each qualifying order once, producing $240. Use matching lines only to determine eligible order IDs, then aggregate from the order-grained table. A semi-join using EXISTS, or joining distinct eligible IDs back to orders, preserves this behavior. Curated guidance should state both table grains, join cardinality, and the aggregation grain of order_total.
Why each option fits or fails:
A. Grouping does not remove repeated contributions: order 4101 still contributes $240 across its two matching lines, leaving the overall result at $360.
B. Both qualifying orders have the same $120 total, so deduplicating monetary values removes a valid order and returns $120 instead of $240.
C. Every order-line pair is already unique; retaining both Hardware lines for order 4101 still counts its order-level total twice.
D. Hardware lines identify qualifying orders, but each order-level total must contribute once; the two qualifying orders therefore contribute $240.
Question 23
Topic: Foundations of Context Engineering
An engineer uses MLflow 3 evaluations to determine whether a larger retrieved-context budget improves an agent’s grounded correctness. The model version, instructions, retrieval method, and custom scorer are unchanged. The scorer checks reference-answer correctness and citation support.
Observed results:
| Evaluation detail | Run A | Run B |
|---|---|---|
| Retrieved-context cap | 6,000 tokens | 18,000 tokens |
| Question set | support_v1 | support_v2 |
| Source snapshot | kb_sep18 | kb_oct02 |
| Grounded-correctness pass rate | 72% | 85% |
The newer question set replaces 30 of the 100 questions. The newer source snapshot contains revised policy documents. All artifacts remain available for reruns. Each proposed comparison will use equal numbers of repeated trials with the model, retrieval method, and scorer unchanged.
Which next comparison would best isolate the effect of the retrieved-context budget?
Options:
A. Compare both caps on their original question sets and source snapshots, reporting confidence intervals.
B. Compare both caps on their original question sets, fixing the source snapshot to
kb_oct02for both.C. Compare both caps on
support_v2, keeping the original source snapshot for each cap.D. Compare both caps on
support_v2, fixing the source snapshot tokb_oct02for both.
Best answer: D
Explanation: A controlled comparison changes the context budget while holding the evaluation workload and available evidence constant. Here, the larger budget was also tested on different questions and revised documents. Its higher pass rate could therefore reflect question difficulty, source quality, or context size.
Evaluate both caps on support_v2 against the frozen kb_oct02 snapshot, using the same reference answers, model, retrieval method, and scorer. Comparing results for the same questions across repeated trials provides stronger evidence about the budget’s effect and helps distinguish consistent gains from run-to-run variation.
More repetitions improve precision; they do not remove systematic differences between experimental conditions. Any observed benefit applies to the controlled workload and source snapshot, not necessarily to every agent task.
Why each option fits or fails:
A. Confidence intervals quantify uncertainty, but retaining different question sets and source snapshots leaves the systematic confounds intact.
B. The source snapshot is controlled, but differences in question difficulty and answerability still prevent attributing the pass-rate difference to the context cap.
C. The question set is controlled, but source revisions still differ between caps, so changes in source content can remain responsible for the gain.
D. Holding the question set and source snapshot fixed leaves the context cap as the experimental difference under the unchanged model, retrieval, and scoring settings.
Question 24
Topic: Multi-Agent and Long-Horizon Task Design
A coordinator must explain apparently inconsistent cancellation totals for tenant T17 in May 2026. Two read-only workers investigate separate frozen source snapshots that the workflow is authorized to access.
Synthesis contract: Each worker returns at most 250 tokens. The coordinator receives only the shared tenant/month scope and the worker returns, not source rows or full worker transcripts. It does not query the snapshots during synthesis.
Trial-run record:
billing:
inputs: shared_scope, BILL_MAY
worker_dependencies: none
result: complete; count=74; definition=cancellations effective in May
source_ref: BILL_MAY; uncertainty=request-to-event linkage not checked
support:
inputs: shared_scope, CASE_MAY
worker_dependencies: none
result: complete; count=92; definition=cancellation requests opened in May
source_ref: CASE_MAY; uncertainty=request-to-event linkage not checked
Which workflow decisions are supported by this record? Select TWO.
Options:
A. Dispatch the two workers concurrently because each investigation has its required inputs before the other worker returns.
B. Add the returned totals during synthesis because independent source investigations produce non-overlapping cancellation evidence for the common reporting period.
C. Schedule billing before support because investigations of the same tenant’s cancellations require a shared intermediate result.
D. Retain counting definitions and source references in bounded worker returns so the coordinator can reconcile the differently scoped totals.
E. Reduce worker returns to counts and completion status because the shared tenant and reporting month already establish a common metric.
Correct answers: A and D
Explanation: Independent evidence-collection tasks can run concurrently when each has its required inputs and neither depends on the other’s findings. Both investigations meet those conditions: they perform read-only work against frozen snapshots using the shared scope. Different returned totals do not create an execution dependency. Final synthesis follows both returns.
Bounding results should remove unnecessary detail, not the information needed to interpret evidence. Here, counting definitions, source references, completion status, and unresolved linkage uncertainty remain useful. Billing counts cancellations effective in May; support counts requests opened in May. These are different measures, not established disjoint populations. The coordinator should reconcile their meanings rather than add the totals or interpret their difference as a verified number of pending cancellations.
Why each option fits or fails:
A. Both workers already have the shared scope and their assigned frozen snapshots, and neither consumes the other’s output.
B. Independent execution does not establish disjoint event populations, and request-open dates and cancellation-effective dates define different measures.
C. Sharing a tenant and reporting month does not create a data dependency; neither recorded investigation requires another worker’s result.
D. The coordinator cannot query the snapshots, so returned definitions distinguish effective cancellations from opened requests, while references identify each finding’s source.
E. Tenant and month identify the investigation scope, but they do not reveal whether a worker counted opened requests or effective cancellations.
Question 25
Topic: Tool Design, MCP, and Agent Context
A Databricks agent is preparing a service-credit claim for shipment S-204. A read-only MCP tool returned 3,600 delivery-event rows, mostly unrelated to this shipment.
Verified extraction:
shipment_id: S-204
observed_delay_hours: 18
contract_revision_at_booking: CR-17
current_contract_revision: CR-22
snapshot_row_ref: audit-62/delivery-events/E-91
live_row_ref: ops.delivery.events/E-91
The saved snapshot is an immutable, durable artifact that the agent is authorized to read. The live reference points to a mutable table.
The remaining workflow must load the contract terms that applied at booking, calculate the shipment’s credit, and cite the delivery evidence used in the audit. The application’s compaction protocol permits replacing the result body while leaving its matching tool-call/result envelope unchanged.
Which replacement best reduces active context while preserving what the remaining work needs?
Options:
A. Keep S-204’s 18-hour delay, contract revision CR-17, and the snapshot row reference; remove the raw rows.
B. Keep S-204’s 18-hour delay, contract revision CR-22, and the live row reference; remove the raw rows.
C. Keep S-204’s 18-hour delay, contract revision CR-22, and the snapshot row reference; remove the raw rows.
D. Keep S-204’s 18-hour delay, contract revision CR-17, and the live row reference; remove the raw rows.
Best answer: A
Explanation: Tool-result compaction should preserve a task-specific working record rather than carry forward the entire source payload. For S-204, that record needs the observed 18-hour delay, contract revision CR-17, and the immutable snapshot reference for row E-91.
CR-17 remains a dependency for the next lookup even though CR-22 is now current. The snapshot reference preserves the evidence used during the audit; a reference to the live table could later resolve to changed data.
Once these facts and their provenance are retained, the bulky raw result can leave active context. The durable snapshot remains available for verification, and the matching tool-call/result envelope remains intact under the application’s protocol.
Why each option fits or fails:
A. This retains the shipment-specific fact, the historical revision needed by the next lookup, and reproducible evidence while removing unrelated rows.
B. The current revision does not govern the booked shipment, and the mutable reference does not preserve the audit’s original evidence.
C. CR-22 is current, but CR-17 governed S-204 at booking; retaining CR-22 would direct the later lookup to the wrong contract terms.
D. The live reference can resolve to changed data, so it does not preserve the evidence snapshot underlying the observed 18-hour delay.
Questions 26-45
Question 26
Topic: Multi-Agent and Long-Horizon Task Design
A coordinator uses rollout and validation workers to prepare a maintenance plan. Both must schedule work within the window in the authoritative Unity Catalog table maintenance_policy. The application requires the merged plan to use the latest approved source snapshot available at merge time.
The application trace includes all source updates during this planning attempt. All times are UTC on the same day. Neither worker has executed its plan.
08:55 source approved version=17 window=14:00-15:00
09:00 shared task state source_version=17
09:00 dispatch rollout snapshot=17
09:01 source approved version=18 window=16:00-17:00
09:02 dispatch validation snapshot=latest
09:03 rollout returned read_version=17 plan_window=14:00-15:00
09:04 validation returned read_version=18 plan_window=16:00-17:00
09:05 merge attempted result=window_conflict
Which conclusions are supported by this evidence? Select TWO.
Options:
A. Reading the same Unity Catalog table establishes a common source snapshot for both workers’ existing plans.
B. The rollout plan requires reconciliation against version 18 before inclusion in the merged maintenance plan.
C. The rollout worker’s return time establishes version 18 as the source snapshot underlying its existing plan.
D. The conflicting plan windows arise from source-snapshot skew between the rollout and validation workers.
E. Updating shared task state’s
source_versionto 18 is sufficient to reconcile the existing worker plans.
Correct answers: B and D
Explanation: Shared source authority does not ensure that independently dispatched workers use the same snapshot. Here, the source changed between dispatches. The rollout plan reflects version 17’s 14:00-15:00 window, while validation reflects version 18’s 16:00-17:00 window. Their disagreement is explained by snapshot skew.
At 09:05, the application’s merge rule makes version 18 applicable. The coordinator should propagate that explicit snapshot to the rollout worker and regenerate its plan, or explicitly reconcile the affected plan content against version 18. Source-version provenance should remain attached to each result and be checked before merging. A shared-state pointer records the intended baseline; changing it does not refresh plans already derived from an earlier snapshot.
Why each option fits or fails:
A. A common governed table provides shared source authority, not snapshot consistency; the recorded reads used different versions.
B. Version 18 is the latest approved snapshot at 09:05, so the rollout plan derived from version 17 must be re-evaluated against it.
C. Return time does not change source provenance; the rollout result explicitly records version 17 despite returning after version 18 became available.
D. The rollout window matches version 17, while the validation window matches version 18, showing that the workers planned against different snapshots.
E. Changing the source-version pointer does not reconcile already derived plan content; the rollout result still uses version 17’s maintenance window.
Question 27
Topic: System Prompt and Instruction Design
A team evaluates a longer system prompt that adds verified few-shot examples for contract-exception handling in a subscription-support agent.
Experiment conditions:
- Paired, repeated evaluations use identical model settings, tools, and source snapshots. Treat the reported task-success rates as stable estimates.
- The evaluation contains equal numbers of cases per route; production traffic has the mix shown below.
- Requests are independent. An existing accurate router identifies the route before prompt assembly, with no additional routing overhead.
- The shorter system prompt uses 1,200 tokens; the longer one uses 3,200 tokens. Other input tokens are unchanged.
| Route | Production traffic | Shorter success | Longer success |
|---|---|---|---|
| Routine account questions | 75% | 96% | 96% |
| Standard renewals | 20% | 85% | 85% |
| Contract exceptions | 5% | 50% | 80% |
The deployment objective is to maximize estimated production task success, then minimize mean system-prompt tokens among deployments tied on success.
Which conclusions are supported? Select TWO.
Options:
A. Using the longer prompt on every route best meets the stated deployment objective.
B. Using the shorter prompt on every route best meets the stated deployment objective.
C. The all-route deployment improves production-weighted task success by 1.5 percentage points.
D. Using the longer prompt only for contract exceptions best meets the stated deployment objective.
E. The all-route deployment improves production-weighted task success by 10 percentage points.
Correct answers: C and D
Explanation: Aggregate quality must reflect production traffic rather than the number of evaluation cases in each route. Only contract exceptions improve, from 50% to 80%, and they represent 5% of requests. The production-weighted gain is \(0.05 \times (80 - 50) = 1.5\) percentage points, increasing estimated success from 91.5% to 93%.
Both all-route and exception-only use of the longer prompt achieve that 93% success rate. Exception-only deployment adds 2,000 tokens to just 5% of requests, giving a mean system-prompt length of \(1,200 + 0.05 \times 2,000 = 1,300\) tokens. All-route deployment averages 3,200 tokens. Because existing routing adds no overhead, selective deployment captures the demonstrated benefit with fewer prompt tokens. The evidence supports a route-specific improvement, not a universal need for the longer prompt.
Why each option fits or fails:
A. All-route use achieves the same estimated success as exception-only use while consuming more system-prompt tokens on routes with unchanged success.
B. Keeping the shorter prompt everywhere minimizes tokens but yields 91.5% estimated success rather than the achievable 93%, violating the quality-first priority.
C. The 30-percentage-point improvement affects 5% of production traffic, producing an overall gain of 1.5 percentage points.
D. Exception-only use retains the full measured quality gain while averaging 1,300 system-prompt tokens, compared with 3,200 for all-route use.
E. Ten percentage points is the equally weighted route-average gain, not the gain weighted by the supplied production traffic.
Question 28
Topic: Tool Design, MCP, and Agent Context
A Databricks agent invokes read-only tools on a custom MCP server through an application adapter. Live discovery from that same server reports these contracts:
| Tool name | Required argument | Description |
|---|---|---|
policy_search | query: string | Search current policy text by topic. |
policy_revision | policy_id: string | Return the current revision for an exact policy identifier. |
The server validates tool names and required arguments before starting a handler. This simplified application diagnostic record is not a Databricks SDK payload:
Search:
agent choice: policy_search {"query":"expense reimbursement"}
adapter sent: search_policies {"query":"expense reimbursement"}
server: unknown_tool; handler_started=false
Revision:
agent choice: policy_revision {"policy_id":"0042"}
adapter sent: policy_revision {"id":"0042"}
server: invalid_arguments; handler_started=false
Which conclusions are supported by the record? Select TWO.
Options:
A. The revision call fails because the selected tool finds no matching revision for the policy identifier.
B. The search call fails because the adapter substitutes a tool name absent from the server’s live catalog.
C. The revision call fails because the agent supplies a policy identifier with the wrong JSON type.
D. The revision call fails because the adapter replaces the required
policy_idfield withid.E. The search call fails because its description does not distinguish topical search from exact identifier lookup.
Correct answers: B and D
Explanation: Tool descriptions help an agent choose a capability; invocation still depends on exact tool names and input contracts. The agent selects policy_search with a string query and policy_revision with a string policy_id, matching the published contracts. The adapter then changes policy_search to the unregistered search_policies and replaces policy_id with id.
Both failures occur before a handler starts, so the record does not show a retrieval or policy-data failure. Repair the adapter’s name and argument mappings, and validate dispatched requests against live discovery. Rewriting already accurate descriptions does not correct these transformations.
Why each option fits or fails:
A. Argument validation fails with handler_started=false, so no revision lookup occurs and the record provides no evidence of a missing policy.
B. The server exposes policy_search, but the adapter sends search_policies, producing unknown_tool before execution.
C. Both requests contain a string identifier, matching the declared type; the transmitted request instead lacks the required field name.
D. The adapter transmits id instead of policy_id, so the required argument is absent even though its value is retained.
E. The description explicitly specifies topic search, and the agent selects policy_search for the topic query before the adapter changes its name.
Question 29
Topic: System Prompt and Instruction Design
An engineer maintains a Databricks Genie space for finance. The team wants to represent its fiscal calendar without rewriting every reusable SQL example.
Reporting rules:
- Ordinary annual requests use an April 1-March 31 fiscal year, labeled by its ending year. An unqualified request for “revenue for 2027” means April 1, 2026-March 31, 2027.
- Two named reports, “Statutory sales report” and “Annual supplier statement,” use January 1-December 31 instead.
Current configuration:
The global instructions contain:
Interpret a year as January 1-December 31 unless the request explicitly says fiscal.
- Reusable SQL examples illustrate joins and metric calculations without fixed year boundaries.
- Two report-specific SQL examples correctly implement the calendar-year exceptions.
- An unqualified request for “revenue for 2027” returns calendar-year totals.
Which configuration change best establishes the required default while preserving the exceptional reports?
Options:
A. Retain the calendar-year default globally for unqualified annual requests; add a general fiscal-year SQL example alongside the two report-specific SQL examples.
B. State the fiscal-calendar default globally, recognizing the two named exceptions; keep their calendar-year logic in the report-specific SQL examples.
C. State the fiscal-calendar rule globally for all annual requests; keep the two report-specific SQL examples with their existing calendar-year logic.
D. Remove the calendar default from global instructions; add a general fiscal-year SQL example and use example matching to select each reporting convention.
Best answer: B
Explanation: Global instructions should define behavior shared across the Genie space. Here, they should establish April-to-March fiscal years as the default and explain that the fiscal-year label identifies the ending year. Ordinary requests for 2027 therefore cover April 1, 2026-March 31, 2027.
The global rule should acknowledge the two named exceptions, while their detailed calendar-year query logic remains in the relevant SQL examples. This keeps the default and exceptional behavior consistent without promoting narrow report logic into a universal rule. Since the reusable examples contain no fixed year boundaries, their joins and metric calculations do not need rewriting. Examples should support the governing instructions, not compete with a contradictory calendar default.
Why each option fits or fails:
A. The retained global instruction still assigns calendar years to unqualified requests; adding a fiscal SQL example does not resolve that conflict.
B. The global default governs ordinary annual requests, while the explicitly scoped exceptions preserve the calendar-year behavior of the two reports.
C. An all-request fiscal rule conflicts with the two valid calendar-year reports, so retaining their examples would leave contradictory guidance.
D. Example matching provides query patterns rather than an explicit organization-wide default, leaving ordinary annual interpretation dependent on which example is selected.
Question 30
Topic: Tool Design, MCP, and Agent Context
A policy-answering agent loads an Agent Skill that references a retired custom MCP tool. The procedure was valid under the tool’s version 1 contract:
Loaded procedure:
tool: policy_search_v1
arguments: query, min_similarity=0.80
result: passages with cosine_similarity >= 0.80
next: use returned passages as grounding evidence
Current tool discovery exposes only the replacement, with no compatibility alias:
Current supported contract:
tool: policy_search_v2
arguments: query, max_results
result: list of candidate passages
passage fields: text, source_id, cosine_distance
cosine_distance = 1 - cosine_similarity
filtering: no relevance cutoff is applied
Both versions use the same embeddings and policy corpus. The application will supply max_results=20 when calling the replacement.
Which skill update should the engineer validate before restoring execution while preserving the original relevance cutoff?
Options:
A. Rebind to
policy_search_v2and retain candidates satisfyingcosine_distance <= 0.80.B. Rebind to
policy_search_v2and retain candidates satisfyingcosine_distance >= 0.20.C. Rebind to
policy_search_v2and retain candidates satisfyingcosine_distance >= 0.80.D. Rebind to
policy_search_v2and retain candidates satisfyingcosine_distance <= 0.20.
Best answer: D
Explanation: A loaded skill supplies procedural guidance; it does not override an MCP server’s supported tool names, arguments, or response semantics. Updating an outdated reference requires validating both the replacement call and downstream interpretation of its results.
Version 1 enforced the similarity cutoff during retrieval. Version 2 returns candidates without that cutoff and reports cosine distance instead. Under the supplied contract, a similarity of at least 0.80 corresponds to a distance of at most 0.20.
Rebind the skill to policy_search_v2, supply max_results=20, and apply the converted filter before adding passages to the agent’s grounding context. Contract tests should cover distances below, at, and above 0.20 to verify the comparison direction and inclusive boundary.
Why each option fits or fails:
A. This permits cosine similarities as low as 0.20, weakening the skill’s original minimum similarity of 0.80.
B. This retains cosine similarities of 0.80 or lower, reversing the comparison and excluding the strongest matches.
C. This selects cosine similarities of 0.20 or lower, reusing the legacy threshold on the wrong metric and reversing relevance.
D. The replacement returns distance rather than similarity, so a minimum similarity of 0.80 becomes a maximum distance of 0.20.
Question 31
Topic: Memory Architecture with Lakebase and MLflow
An analyst-facing agent prepares reports across multiple sessions. After a reconnect or application restart, it must resume from the last acknowledged workflow checkpoint, including selected source versions, completed-step results, and pending work.
Lakebase is available for durable application state. MLflow records experiment runs and traces for analysis. The application’s MLflow export path must remain asynchronous to avoid adding diagnostic logging latency to requests.
Observed behavior:
- Structured workflow state exists only in process memory; chat messages contain progress summaries rather than intermediate results.
validate_inputscompletes, and the agent acknowledges completion.- The process crashes before exporting that trace. The latest available MLflow trace still ends at
load_sources. - Restart logic reconstructs state from that trace and treats
validate_inputsas pending.
Which persistence design best supports reliable resumption while retaining offline quality evaluation?
Options:
A. Snapshot in-memory state periodically to Lakebase, recover intervening completed steps from MLflow traces, and use MLflow for offline quality evaluation.
B. Persist chat messages in Lakebase, derive workflow progress from MLflow traces, and use MLflow for offline quality evaluation.
C. Persist checkpoints in Lakebase before acknowledging progress, resume directly from them, and use MLflow for offline quality evaluation.
D. Persist checkpoints as MLflow run artifacts, store session-to-run mappings in Lakebase, and use MLflow for offline quality evaluation.
Best answer: C
Explanation: Operational checkpoints and experiment records serve different purposes. The application should define a structured checkpoint and commit it to Lakebase before acknowledging progress. Resumption then reads that checkpoint to recover source versions, completed results, and pending work.
MLflow preserves runs, traces, and evaluation results for analysis, but asynchronous diagnostic export can lag execution. Here, validation completed even though the latest exported trace stops at source loading. Reconstructing operational state from that trace therefore loses acknowledged progress.
MLflow can continue supporting failure analysis and evaluation of restart tests without becoming the application’s state authority. Lakebase provides durable storage; checkpoint structure, commit timing, and recovery logic remain application responsibilities.
Why each option fits or fails:
A. Progress between snapshots can be acknowledged but absent from delayed traces, leaving the same recovery gap after a crash.
B. Progress summaries omit required intermediate results, and the delayed trace does not contain the completed validation needed for resumption.
C. Committing application state before acknowledgment makes completed progress recoverable independently of delayed MLflow export.
D. Asynchronously exported artifacts can lag acknowledged progress; storing a session-to-run mapping does not preserve a checkpoint missing from MLflow.
Question 32
Topic: Context Compression and Compaction
An engineer is considering a single compaction step before an agent resumes. A separate billed model call produces a summary that replaces the original history block.
Projection assumptions:
- Two future agent turns are expected; compare horizons of one through four turns.
- Each turn includes the original history block or its summary once. The displayed token counts remain constant.
- Required task information is preserved. All other input tokens, agent output tokens, and charges are identical between the two paths.
- No caching discounts or additional compaction apply.
Billed usage:
| Component | Tokens | Rate per million tokens |
|---|---|---|
| Summary-generation input | 12,000 | $2 |
| Summary-generation output | 3,000 | $8 |
| Original history input | 12,000 per turn | $2 |
| Replacement summary input | 3,000 per turn | $2 |
Which conclusions about total cost relative to retaining the original history are supported? Select TWO.
Options:
A. The first horizon at which compaction lowers total cost is four future turns.
B. Across the two expected future turns, compaction raises total cost by $0.012.
C. Across the two expected future turns, compaction lowers total cost by $0.036.
D. Across the two expected future turns, compaction lowers total cost by $0.012.
E. The first horizon at which compaction lowers total cost is three future turns.
Correct answers: B and E
Explanation: Compaction exchanges a one-time generation cost for lower recurring input charges. Its financial value depends on how many future turns reuse the shorter context.
- Summary generation costs $0.024 for 12,000 input tokens and $0.024 for 3,000 output tokens, totaling $0.048.
- Each future turn saves 9,000 input tokens at $2 per million tokens, or $0.018.
- Two turns therefore save $0.036, which is $0.012 less than the generation cost.
At the expected two-turn horizon, history-related costs are $0.048 without compaction versus $0.060 with compaction. Shared charges cancel from the comparison.
Three turns save $0.054 and are the first integer horizon where savings exceed the $0.048 generation cost. Preserving required information makes the comparison meaningful, but it does not guarantee that compaction reduces total cost.
Why each option fits or fails:
A. Three turns already save $0.054 against a $0.048 generation cost, so four is not the first cheaper horizon.
B. Generating the summary costs $0.048, while two uses save $0.036 in input charges, leaving a net cost increase of $0.012.
C. The $0.036 is the gross reduction in history-input charges; including the $0.048 summary-generation charge makes compaction more expensive.
D. A $0.012 saving results from counting only the summarizer’s $0.024 input charge and omitting its $0.024 output charge.
E. Three turns save $0.054, exceeding the $0.048 summary cost; one or two turns save less than that cost.
Question 33
Topic: Context Compression and Compaction
An engineer reviews compaction for a Databricks agent that drafts warehouse transfer requests. The original record remains unchanged in Lakebase, but the drafting step receives only the compacted context and cannot reread that record.
Drafting contract:
transfer_idmust match the source string exactly.ship_onmust match the source date. The drafting step interprets slash-form dates asDD/MM/YYYYand emits ISO dates.- Mass may use grams or kilograms if the physical quantity remains unchanged.
Compacted context:
Transfer TR-731 should ship on 11/04/2027 with a mass of 2.4 kg.
Observed trace:
| Field | Source record | Draft after compaction |
|---|---|---|
transfer_id | TR-000731 | TR-731 |
ship_on | 2027-11-04 | 2027-04-11 |
| Mass | 2,400 g | 2.4 kg |
MLflow recorded a custom summary-fluency score of 4.9/5. The draft also passed type-and-format validation.
Which conclusions about this compaction are supported? Select TWO.
Options:
A. The unchanged Lakebase record compensates for altered exact fields during the current drafting step.
B. The identifier rewrite changes the exact string key needed to address the original transfer record.
C. The mass rewrite changes the physical quantity needed to fulfill the original transfer request.
D. The date rewrite changes the scheduled date under the drafting step’s stated parsing convention.
E. The fluency score and type-and-format validation confirm that required transfer values survived compaction.
Correct answers: B and D
Explanation: Compaction must preserve task-critical information, not merely readable prose. Here, removing zeros changes an exact identifier, while replacing an ISO date with a localized representation changes its interpretation. The mass conversion is valid because both the quantity and its associated unit preserve the same physical mass.
A safer strategy retains required identifiers and unambiguous dates in structured active state, while compressing surrounding narrative. Evaluation should compare identifiers and dates directly with the source record and compare quantities using unit-aware normalization. Fluency scoring and type-and-format validation measure different qualities and cannot replace these checks.
Lakebase remains the durable source of truth, but its values help only when retained in context or explicitly reloaded. Repeat field-level checks across successive compaction cycles to detect accumulating errors.
Why each option fits or fails:
A. The drafting step cannot reread Lakebase, so durable storage does not restore information altered in its active context.
B. The contract requires exact string equality, so TR-731 cannot substitute for TR-000731; compaction must preserve the full identifier.
C. 2.4 kg equals 2,400 g, so the mass remains correct under the contract despite the changed numeric representation.
D. Under DD/MM/YYYY, 11/04/2027 means 11 April 2027, not the source date of 4 November 2027.
E. Fluency and structural validity do not establish exact identifier or date agreement; both fields fail despite these checks.
Question 34
Topic: Multi-Agent and Long-Horizon Task Design
A coordinator must produce a policy comparison using the source versions listed in an approved manifest. The application dispatches tasks only after their predecessors are COMPLETE. Preserve the dependency graph and existing worker inputs for this run.
Scroll sideways if needed. Open full-size diagram in a new tab
Text description
R, retrieving the source bundle, belongs to the evidence worker and is COMPLETE. R is a prerequisite for V, verifying the source bundle, which is PENDING with no assigned owner. V is a prerequisite for D, drafting the policy comparison. D is a prerequisite for C, checking citation links. D and C belong to the writing worker and are BLOCKED.
Available context:
- Evidence worker: Approved manifest, full retrieved source bundle, source IDs, and version metadata.
- Writing worker: Condensed excerpts and citation handles, with no manifest or version metadata.
Latest handoff messages:
Evidence worker: R is complete; the writing worker will handle V.
Writing worker: The evidence worker owns V; I am waiting for its verdict.
Which shared-task-state update best resolves the stalled dependency?
Options:
A. Assign
Vto the writing worker; require recorded confirmation that every draft citation resolves to a retrieved excerpt before markingVCOMPLETE.B. Assign
Vto the evidence worker; require recorded confirmation that a retrieval hit exists for every required source before markingVCOMPLETE.C. Assign
Vto the evidence worker; require recorded confirmation that all required source IDs and versions match the manifest before markingVCOMPLETE.D. Assign
Vto the writing worker; require recorded confirmation that all required source IDs and versions match the manifest before markingVCOMPLETE.
Best answer: C
Explanation: An unresolved handoff needs an explicitly accountable owner and evidence-based completion criteria. Here, V sits between retrieval and drafting, so verification must finish before the writing worker begins.
The evidence worker has the approved manifest and source identity/version metadata needed to verify that every required source is present at its listed version. Record that worker as V’s owner and retain the comparison verdict in shared state. Mark V COMPLETE only after the comparison passes, allowing D to start.
The later citation check serves a different purpose: confirming that the draft’s links resolve. Successful retrieval and citation linking do not independently establish compliance with the approved manifest.
Why each option fits or fails:
A. Citation checks require the draft, but drafting is blocked by verification; resolving citation links also does not establish approved source versions.
B. A retrieval hit confirms source availability, not that its version matches the approved manifest required for the policy comparison.
C. The evidence worker has the context needed for verification, and the recorded comparison provides completion evidence for the prerequisite blocking drafting.
D. The writing worker lacks the manifest and source-version metadata needed to produce this verdict under the unchanged input contract.
Question 35
Topic: Foundations of Context Engineering
An agent retrieves the following excerpts using Databricks AI Search. Source metadata contains only the filename and page number; section labels appear in the text.
User request:
When can rejected-applicant and former-employee records be deleted, including when a legal hold applies?
Retrieved context:
Source: records_standard_v4.pdf, page 6
Employee portal | Records standard
Home > Policies > Help
Rejected applicants
Delete records 90 days after the decision.
If a legal hold applies, retain records until it is released.
Source: records_standard_v4.pdf, page 7
Employee portal | Records standard
Home > Policies > Help
Former employees
Delete records 7 years after departure.
If a legal hold applies, retain records until it is released.
The team wants to reduce boilerplate without lowering answer correctness or source attribution on its existing policy evaluation set.
Which preprocessing change best supports this objective?
Options:
A. Remove later occurrences of identical lines across the batch; keep remaining text under its original section headings, with filename and page citations.
B. Remove the portal banner and navigation lines; keep policy text grouped by section heading, with filename and page citations on each group.
C. Extract sentences containing numeric retention periods; keep that text grouped by section heading, with filename and page citations on each group.
D. Remove the portal banner, navigation lines, and section headings; keep policy text grouped by page, with filename and page citations on each group.
Best answer: B
Explanation: Boilerplate removal should distinguish repeated presentation text from repeated substantive rules. The portal banner and navigation add no information needed to answer the request. The section headings do: they associate the 90-day rule with rejected applicants and the seven-year rule with former employees.
The legal-hold sentence is identical on both pages, but its placement establishes the exception within each section. Removing its later occurrence would leave that section incomplete. Retaining filename and page citations preserves access to the supporting source.
Evaluate the cleaned context against the same source snapshot and policy cases covering both record categories and legal holds. Compare correctness and source attribution as well as token count. A shorter context is useful only if it preserves the information needed for supported answers.
Why each option fits or fails:
A. Batch-wide deduplication removes the legal-hold clause from the former-employee section, even though that repeated clause establishes an exception for that category.
B. The banner and navigation do not contribute to the policy answer, while section headings, retention rules, legal-hold exceptions, and citations preserve meaning and provenance.
C. Numeric-period extraction retains the ordinary deadlines but drops the legal-hold exception, which has no numeric duration and directly affects the requested answer.
D. Removing section headings loses the explicit association between each retention rule and its record category because source metadata does not contain those labels.
Question 36
Topic: Memory Architecture with Lakebase and MLflow
An agent has resolved a user’s request to preparing a handoff for an ongoing incident investigation. It retrieves prior scope decisions and unresolved dependencies through an application-owned Lakebase task-memory retriever.
The retriever’s candidate depth selects a prefix of a fixed ranking, and every returned record is inserted unchanged into context. MLflow replays use the same query and source snapshot. Reviewers identify five memory records required for a complete handoff; none is available elsewhere in context.
- Memory allowance: 4,500 tokens, after reserving all non-memory context and output headroom.
- Objective: Retain all required memory records while minimizing memory tokens.
| Candidates returned | Required records found | Memory tokens |
|---|---|---|
| 8 | 3 of 5 | 950 |
| 16 | 4 of 5 | 1,700 |
| 32 | 5 of 5 | 2,900 |
| 64 | 5 of 5 | 4,300 |
Which candidate depth is best supported for this resolved intent?
Options:
A. Use a candidate depth of 32 for this resumed task.
B. Use a candidate depth of 16 for this resumed task.
C. Use a candidate depth of 8 for this resumed task.
D. Use a candidate depth of 64 for this resumed task.
Best answer: A
Explanation: Retrieval depth should follow task-specific coverage and context cost, rather than the largest depth that fits the window. For the incident handoff, all five required records must reach context. Depth 32 achieves that coverage using 2,900 memory tokens within the 4,500-token allowance. Deeper retrieval adds context without adding required evidence in the measured replay.
The relevant-hit count establishes memory coverage, not final-answer correctness; the agent must still use those records accurately. This depth is supported for the tested intent, ranking, and source snapshot, not as a universal setting. Changes to memory contents or task requirements warrant reevaluation.
Why each option fits or fails:
A. Depth 32 retrieves all five required records using 2,900 tokens, the lowest measured token consumption that achieves complete coverage.
B. Depth 16 reduces context consumption but omits one required memory record, leaving the handoff without complete supporting evidence.
C. Depth 8 has the smallest payload but excludes two required memory records, so its token savings sacrifice necessary evidence.
D. Depth 64 fits the allowance but consumes 1,400 additional tokens without improving required-record coverage, contrary to the minimization objective.
Question 37
Topic: Memory Architecture with Lakebase and MLflow
An assistant uses the current workspace-scoped managed memory service for durable knowledge and a separate managed session service for conversation continuity.
Application contract:
- Durable memories belong to authenticated users and follow those users across sessions.
- Anonymous requests may continue within isolated managed sessions.
- The backend service principal has store-level access. The memory service’s
actor_idpartitions memories; it is not an independent authorization boundary.
Application trace, before any memory operation:
verified_caller = null
session_cookie = valid
session_id = anon_73
body.actor_id = user_204
fallback_actor = guest
planned_memory_actor = guest
planned_operations = retrieve, retain
The server-issued cookie resolves only to anon_73. The request body was supplied by the unauthenticated client.
Which conclusions are supported by the contract and trace? Select TWO.
Options:
A. The handler must suspend durable-memory operations until a verified caller identity is available.
B. The handler can retrieve and update durable memory under the shared
guestactor.C. The handler can maintain conversation continuity in the isolated
anon_73session.D. The handler can retrieve and update durable memory under the server-issued
session_id.E. The handler can retrieve and update durable memory under the supplied
body.actor_id.
Correct answers: A and C
Explanation: Per-user durable memory requires an actor identity derived in trusted application code from the verified caller. Here, verified_caller is absent, so durable-memory retrieval and retention must not proceed. Backend store authorization permits service access; it does not authenticate the end user or make an actor identifier trustworthy.
Using guest would silently merge unknown users’ memories. Accepting the body’s actor value would trust an unauthenticated identity claim, while using the session ID would confuse conversation continuity with cross-session user identity.
The isolated managed session can still support the anonymous conversation. Once authentication establishes the caller, the application can derive the appropriate memory actor and perform authorized durable-memory operations.
Why each option fits or fails:
A. With no verified caller, the application cannot associate memory retrieval or retention with an authenticated user as its strategy requires.
B. A shared actor merges anonymous users’ retained knowledge, so the backend’s store access does not make this fallback compatible with per-user memory.
C. The valid cookie identifies an isolated session, and the application explicitly permits anonymous session continuity independently of durable user memory.
D. A session identifier identifies a conversation, not an authenticated user across sessions, so it does not satisfy the declared durable-memory strategy.
E. The client-supplied actor value is not a verified caller identity and cannot establish whose durable memories the request may access.
Question 38
Topic: Knowledge Retrieval and Genie Configuration
An AI support agent uses a SQL retrieval tool that executes as sp_case_reader on a Databricks SQL warehouse. Unity Catalog checks data access using that service principal.
Policy: The retrieval identity may read case_id, product, and resolution_text, but must not be able to read customer_email or tax_id.
Observed state:
- All five columns are in
support.prod.case_records. - The prompt and tool response schema list only the three approved columns.
sp_case_readerhas directSELECTon the source table, with no other privileges providing source access.- A steward-owned Unity Catalog view,
support.curated.case_lookup, exposes only the approved columns. Its owner retains source-table access, and the service principal already has required catalog and schema usage privileges.
Which change should the engineer make to meet the policy while preserving approved-field retrieval?
Options:
A. Move retrieval to the approved view; grant the service principal SELECT on it and revoke its SELECT on the source table.
B. Move retrieval to the approved view; grant the service principal SELECT on it and retain its SELECT on the source table.
C. Keep retrieval on the source table; restrict the tool’s response fields to the approved columns while retaining the service principal’s SELECT grant.
D. Keep retrieval on the source table; replace the service principal’s table-level SELECT with SELECT grants limited to the three approved columns.
Best answer: A
Explanation: Unity Catalog authorizes retrieval using sp_case_reader, so the policy must be enforced through that identity’s data permissions. Removing columns from a prompt or response schema does not reduce its source privileges.
The steward-owned view exposes only approved fields. Grant the principal SELECT on that view, revoke its source-table SELECT, and point retrieval at the view. The view owner retains the underlying-source privileges, so the retrieval principal does not need direct source-table access.
Unity Catalog does not provide column-scoped SELECT grants. A projection view, combined with removal of underlying-table access, provides the narrower authorized read surface required here.
Why each option fits or fails:
A. View access preserves retrieval of approved fields, while revoking source-table access removes the principal’s authorized path to restricted columns.
B. Changing the configured retrieval source does not remove the principal’s existing authorization to read restricted columns directly from the source table.
C. Response-field filtering limits what reaches the model, but the service principal remains authorized to read restricted columns from the source table.
D. Unity Catalog grants SELECT at the table or view level, not on individual columns, so this proposed privilege restriction is unsupported.
Question 39
Topic: Memory Architecture with Lakebase and MLflow
An agent recalls approved supplier payment terms through an application-defined MemoryRepository backed by Lakebase. The source row shown below is current and validated.
Adapter contract: fact.payment_term_days contains the approved payment term. metadata.review_interval_days specifies how often the memory record should be reviewed.
Simplified MLflow 3 trace:
request: Approved payment term for supplier SUP-17?
lakebase.read: memory_id=mem-52 supplier_id=SUP-17
fact.payment_term_days=45 metadata.review_interval_days=30
memory.retrieve: memory_id=mem-52 supplier_id=SUP-17
fact.payment_term_days=45 metadata.review_interval_days=30
context.assemble: Approved payment term for SUP-17: 30 days.
agent.respond: The approved payment term for SUP-17 is 30 days.
The context.assemble statement is the only payment-term information passed to the model.
Application assembly adapter:
def assemble(hit):
days = hit["metadata"]["review_interval_days"]
return f"Approved payment term for {hit['supplier_id']}: {days} days."
Which conclusions about the memory failure are supported by this evidence? Select TWO.
Options:
A. The retrieval stage selects another supplier’s memory record instead of the requested supplier’s record.
B. The assembly adapter injects the record-review interval as the requested supplier’s approved payment term.
C. The stored record contains an incorrect approved payment-term fact for the requested supplier.
D. The retrieval stage returns the requested supplier’s approved payment-term fact with its correct value.
E. The response stage disregards the correct payment term included in the assembled model context.
Correct answers: B and D
Explanation: Locate a memory-injection defect by comparing the source record, retrieval output, assembled context, and response in order. Here, retrieval preserves the approved 45-day fact for the correct supplier. The first divergence occurs during assembly: the adapter reads the 30-day review interval and labels it as the approved payment term. The response follows that incorrect context rather than disregarding correctly injected memory.
Repair the adapter to read hit["fact"]["payment_term_days"]. Keep the validated source record and correct retrieval behavior unchanged. Replay the saved retrieval result and verify that the assembled statement says 45 days, then evaluate the final response separately. This distinguishes an assembly defect from incorrect stored facts, retrieval-selection errors, and generation-stage failures.
Why each option fits or fails:
A. Both the requested supplier and the retrieved record identify SUP-17; the trace shows no supplier-selection mismatch.
B. The adapter reads metadata.review_interval_days, so context assembly converts a 30-day review cadence into a payment-term assertion.
C. The validated source row stores 45 days in fact.payment_term_days; its 30-day value belongs to the review schedule.
D. The retrieval span preserves mem-52, supplier SUP-17, and fact.payment_term_days=45, matching the current approved source row.
E. The model receives a 30-day payment-term statement, not the correct 45-day term, so the trace does not demonstrate generation-stage disregard.
Question 40
Topic: Memory Architecture with Lakebase and MLflow
An engineer is redesigning cross-session memory for tenants North and South using the current workspace-scoped managed-memory service. Session records are stored separately.
Current design:
- Both tenant workers use a shared memory store and service principal.
- The application derives
actor_idfrom verified tenant and user claims, creating partitions such asnorth:user17andsouth:user42. - An isolation test using North’s worker credentials calls the service directly with
south:user42and successfully retrieves South’s memory.
Access contract: The service authorizes the calling principal at the store level. Store grants cover all actor partitions within that store; there are no actor-level access grants.
Each tenant’s worker must retain direct service access. Credentials available to one tenant’s worker must not authorize reads or writes to the other tenant’s memory, even when application routing is bypassed.
Which design meets this requirement?
Options:
A. Create a store per tenant; retain a shared service principal with read/write access to every store and route by verified tenant.
B. Keep one shared store; assign a service principal per tenant and grant each principal read/write access to the shared store.
C. Keep one shared store and principal; enforce tenant-specific actor allowlists in each worker before any managed-memory call.
D. Create a store and service principal per tenant; grant each principal read/write access to only that tenant’s store.
Best answer: D
Explanation: The managed-memory store is the authorization boundary; an actor ID selects a logical partition within that boundary. Deriving actor IDs from verified callers helps the application select the intended user’s memories, but it does not restrict what an authorized principal can request directly. North’s successful retrieval from South’s partition demonstrates this distinction.
Use separate tenant stores and tenant-specific service principals whose grants are limited to their own stores. Continue deriving actor IDs in trusted application code for user-level partitioning. Separate stores alone are insufficient if credentials available to either tenant authorize access to both stores. Independently stored session records do not change long-term memory permissions.
Why each option fits or fails:
A. Separate stores do not isolate tenants when the principal available to both workers is authorized to read and write every store.
B. Distinct principals do not isolate actor partitions when every principal has read/write access to the same store.
C. Worker-side allowlists can be bypassed by direct service calls made with credentials authorized for the entire shared store.
D. Tenant-scoped store grants prevent access to the other tenant’s memory independently of the actor ID supplied by the worker.
Question 41
Topic: Tool Design, MCP, and Agent Context
An invoice-adjustment agent uses a custom MCP server and has submitted a user-approved adjustment. At 09:24, it is compacting history into a task-state note. Its default policy removes tool messages more than 15 minutes old.
Remaining task: Verify that the submitted operation completes with the invoice identity, revision, total, and currency specified by the approved preview.
Application tool contract: The read-only inspect_operation(operation_id) tool returns execution state and, once completed, final invoice values. These application-level records use simplified tool names.
09:00 preview_adjustment result
invoice_id=INV-204, expected_revision=17
expected_total=4800.00, currency=USD
09:02 user approval: Approve the 09:00 preview.
09:18 submit_adjustment result
transport_status=ok, operation_id=op-81, state=queued
09:24 task_state=verification_pending
Which retention decisions preserve information required for the remaining task? Select TWO.
Options:
A. Retain the older preview’s invoice identity and expected values as the baseline for verification.
B. Retain the approval note’s preview reference as the expected-value baseline for verification.
C. Retain the submission’s successful transport status as evidence that outcome verification is complete.
D. Retain the submission’s operation ID and queued status as the last observed operation state.
E. Retain the recent submission result as the complete expected-value baseline for verification.
Correct answers: A and D
Explanation: Compaction should preserve evidence according to remaining task dependencies, not age alone. The 09:00 preview is 24 minutes old, but it defines the approved outcome: invoice INV-204, revision 17, total $4,800.00, and currency USD. Removing those values would remove the reference needed to judge the eventual result.
The submission result contributes different information: operation op-81 and its last observed queued state. The agent needs that identifier to inspect the existing operation, and it must distinguish submission acknowledgement from completed execution. The compact task-state note should preserve both the approved target and the unresolved operation record. Verification can finish only after inspection reports completion and the returned invoice values match the retained target.
Why each option fits or fails:
A. The approved preview supplies the comparison target, so its age does not eliminate its role in the pending verification.
B. The approval references a preview but contains none of the expected invoice values needed to compare the operation’s final result.
C. transport_status=ok confirms a successful response, while the queued operation and pending verification do not establish a verified final outcome.
D. inspect_operation requires op-81, and the last observed queued status shows that submission did not establish execution completion.
E. The submission result supplies execution metadata, not the expected invoice identity, revision, total, and currency required for comparison.
Question 42
Topic: Foundations of Context Engineering
An agent is planning a supplier-review report on October 8, 2026. Tomorrow it may need to verify an order-level exception. Its next model request has room for at most 4,000 tokens of additional tool content.
Tool result:
{
"summary": "12 orders missed contractual delivery deadlines.",
"evidence_ref": "evidence://supplier-review/418/revision/3",
"full_evidence_tokens": 26000,
"snapshot_mode": "immutable",
"retained_until": "2026-11-08T00:00:00Z"
}
Application-defined reference contract:
- The reference resolves to the indicated snapshot through the retention deadline.
read_evidence(evidence_ref, section)returns only the requested section, containing at most 900 tokens.- Every read executes as the verified end user and checks that user’s current Unity Catalog read authorization on
procurement.audit.review_evidence. The initial tool call was authorized.
Which conclusions are supported? Select TWO.
Options:
A. The summary and reference can support planning while the full evidence remains outside the model’s current context.
B. The model’s current input budget must include the full snapshot’s token count once its reference is present.
C. The reference resolves to the latest evidence revision available at the time of the later read.
D. The current verified caller must pass a new Unity Catalog access check when a referenced section is read.
E. The initial tool call’s successful access check remains sufficient for source reads before the retention deadline.
Correct answers: A and D
Explanation: A source reference is a pointer to evidence, not evidence inserted into the model input or an access grant. Returning a compact summary with a durable, version-specific reference keeps the planning context small. The 26,000-token snapshot remains in storage rather than consuming the available 4,000-token allowance.
When verification is needed on October 9, the application can resolve revision 3 and return only the relevant section, containing at most 900 tokens. The November 8 retention deadline covers that later read and preserves the summarized revision. Availability and authorization remain separate: the tool must enforce the verified caller’s current Unity Catalog permissions. A valid reference does not preserve access if those permissions have changed.
Why each option fits or fails:
A. The compact result fits the 4,000-token allowance, and the retained reference lets the agent fetch specific evidence only when verification requires it.
B. A reference does not insert the stored snapshot into the model input; only content actually supplied to the model consumes its input budget.
C. The reference is pinned to immutable revision 3, so later resolution returns that snapshot rather than switching to a newer revision.
D. The tool evaluates the verified end user’s current permissions on every read, so an earlier successful call cannot authorize a later one.
E. The retention deadline guarantees source availability, not cached permission; each later read checks the caller’s current Unity Catalog authorization.
Question 43
Topic: Knowledge Retrieval and Genie Configuration
An engineer investigates why an agent using Databricks AI Search gives a shipping-insurance deadline when asked for the equipment warranty claim deadline after delivery.
Index configuration: Only chunk_text is embedded. doc_title and section are stored and returned as metadata. Retrieval is vector-only, with no filters or reranker.
Application replay: Both runs use the same ranked hits and generation settings. Only the fields shown to the model for each passage change.
rank=1 score=0.89
doc_title=Shipping Insurance Guide
section=Transit damage claims
chunk_text=Submit the claim within 7 days of delivery.
rank=2 score=0.85
doc_title=Equipment Warranty Guide
section=Warranty claims
chunk_text=Submit the claim within 30 days of delivery.
run=A passage_fields=chunk_text answer=7 days
run=B passage_fields=doc_title,section,chunk_text answer=30 days
Which conclusions are supported by the configuration and replay? Select TWO.
Options:
A. Rendering each passage with its title and section addressed the observed deadline-attribution error in the model’s generation context.
B. The stored title and section fields do not influence vector similarity under the current retrieval configuration.
C. The higher-scoring passage establishes the applicable seven-day deadline for the equipment warranty claim.
D. The replay establishes that adding document headings to embedding inputs improves candidate recall for warranty queries.
E. The replay establishes that returning document headings as metadata causes AI Search to rerank the retrieved passages.
Correct answers: A and B
Explanation: Locally ambiguous chunks can lose the subject established by a document title or section. Both passages mention “the claim” and delivery, but their headings assign them to different policies. Attaching each passage’s title and section supplies the scope needed to associate the 30-day deadline with equipment warranty claims.
The replay changes context assembly, not retrieval: both runs receive the same hits and scores. Since the index embeds only chunk_text and uses no metadata filters or reranker, stored headings do not alter vector similarity. Metadata helps only through a mechanism that consumes it. Testing heading-enriched embeddings would require changing the embedding input, recomputing the affected vectors, and evaluating retrieval separately. The observed correction does not establish a candidate-recall improvement.
Why each option fits or fails:
A. Both runs contain the warranty passage, but only the labeled format exposes its document scope and produces the matching 30-day deadline.
B. Only chunk text is embedded, and no metadata-dependent retrieval stage is used, so stored headings do not affect the vector comparison.
C. Vector similarity does not establish policy applicability; the source headings identify the seven-day passage as shipping insurance, not equipment warranty.
D. The embedding input and retrieved hits never change, so the replay provides no comparison of heading-enriched embeddings or candidate recall.
E. Both runs retain the same ranks and scores, and the configuration has no reranker; returning metadata does not create a ranking operation.
Question 44
Topic: Memory Architecture with Lakebase and MLflow
An application searches the current managed memory service for a saved hotel spending limit. It derives actor_id from the verified caller, employee-17. Both searches below use this actor and the same memory store.
Supplied retrieval contract:
- Candidates are scoped to the supplied
actor_idbefore lexical BM25 ranking of memory text. - Search performs no embedding comparison or synonym expansion. Records with no query-token overlap are not returned.
- The service provides memory APIs, not a SQL query interface.
Confirmed persisted records:
| ID | Actor | Memory text |
|---|---|---|
| m1 | employee-17 | Hotel reimbursement cap: $180 per night. |
| m2 | employee-29 | Hotel reimbursement cap: $220 per night. |
Observed searches:
| Query | Returned IDs |
|---|---|
| lodging allowance ceiling | None |
| hotel reimbursement cap | m1 |
Which conclusions are supported by the contract and observed results? Select TWO.
Options:
A. Retrieval of
m1is explained by lexical overlap after the query is rewritten using the stored vocabulary.B. Retrieval of
m1is explained by exact equality between the entire query and the entire stored memory text.C. Retrieval of
m1is explained by vector similarity between the query and memory text within the selected actor.D. Persistence of these records implies that the memory service also provides a SQL interface for retrieving them.
E. Exclusion of
m2is explained by its different actor identifier being filtered out before lexical candidate ranking.
Correct answers: A and E
Explanation: Lexical BM25 retrieval uses matching terms, not embedding-based similarity or equality of complete strings. Here, “lodging allowance ceiling” expresses a related intent but shares no indexed words with the saved hotel limit. Its empty result therefore does not show that the memory is missing. Rewriting the query as “hotel reimbursement cap” supplies matching words and retrieves m1, even though the query omits the amount and per-night wording.
Actor scoping is a separate step performed before ranking. Although m2 contains the same matching words, it belongs to employee-29 and cannot participate in employee-17’s search. Finally, persistent storage does not imply a SQL interface. Retrieval must follow the interfaces and semantics explicitly provided by the memory service contract.
Why each option fits or fails:
A. The first query shares no words with the memory text, while the rewritten query contains hotel, reimbursement, and cap.
B. The stored text includes an amount and per-night wording absent from the successful query, so the complete strings are not equal.
C. The supplied contract explicitly excludes embedding comparisons; the successful query instead shares indexed words with the stored memory.
D. Durable persistence does not establish SQL availability, and the supplied service contract explicitly provides no SQL query interface.
E. The contract scopes candidates to employee-17 before ranking, making employee-29’s record ineligible despite its matching vocabulary.
Question 45
Topic: Memory Architecture with Lakebase and MLflow
An agent is preparing a rollout checklist for project Cedar. Its memory interface is an application-managed Lakebase decision table with an application-side cache.
Retrieval contract:
- A cache hit returns a copied record and checks only its expiration time.
- A cache miss reads the latest committed Lakebase row and caches a copy for 30 minutes.
- Prompt rendering copies the retrieved decision into a string for later submission.
Application state record, all times on the same day:
10:00 cache loaded: project=cedar, version=7, expires_at=10:30
Version 7 decision: launch to all branches
10:11 prepared prompt rendered from version 7; not submitted
10:12 verified project owner correction committed as version 8
Version 8 decision: launch to pilot branches
10:14 no cache invalidation, version check, or prompt rebuild has run
Which conclusions about using the corrected decision in the next model call are supported? Select TWO.
Options:
A. Invalidating the cache does not change the prepared prompt, which still contains the version 7 rollout decision until its content is rebuilt.
B. Invalidating the cache now will still require waiting until 10:30 before a cache miss can return the committed version 8.
C. The next cache hit will return version 8 because committing the corrected Lakebase row updates the cached copy.
D. The next retrieval can still return version 7 before 10:30 unless the entry is invalidated or refreshed after a source-version check.
E. Re-rendering the prepared prompt from the unexpired cache will use version 8 because the source correction is already committed.
Correct answers: A and D
Explanation: A trusted correction changes the authoritative decision, but cached records and rendered prompts are separate snapshots. The cache’s 30-minute lifetime controls expiration, not how long the old decision remains authoritative. At 10:14, the unexpired entry still supplies the superseded instruction to launch to all branches.
To use the corrected pilot rollout decision, invalidate the project-specific entry or perform a source-version check that refreshes it when the version differs. Then rebuild the unsubmitted prompt using version 8. These address two distinct stale-context layers: refreshing retrieval does not alter text already assembled for a model call, while rebuilding through an unchanged cache simply reproduces the old decision.
Why each option fits or fails:
A. The prepared prompt contains copied text rather than a live reference, so cache invalidation cannot replace its existing rollout decision.
B. A cache miss reads the latest committed row immediately; the removed entry’s expiration time does not delay access to version 8.
C. The cache holds a copied record and checks only expiration, so the Lakebase commit does not update its version 7 entry.
D. At 10:14, the version 7 entry remains unexpired, and retrieval does not check whether Lakebase contains a newer version.
E. Re-rendering through the unchanged retrieval path still receives cached version 7, so rebuilding alone does not obtain the corrected decision.
Review your attempt
| Domain | Correct | Missed or guessed question numbers |
|---|---|---|
| Foundations of Context Engineering | ___ / 7 | ___ |
| System Prompt and Instruction Design | ___ / 4 | ___ |
| Knowledge Retrieval and Genie Configuration | ___ / 9 | ___ |
| Memory Architecture with Lakebase and MLflow | ___ / 8 | ___ |
| Tool Design, MCP, and Agent Context | ___ / 6 | ___ |
| Context Compression and Compaction | ___ / 5 | ___ |
| Multi-Agent and Long-Horizon Task Design | ___ / 6 | ___ |
For each missed question, name the decisive boundary: authority, access, retrieval, persistence, evidence or task dependency. Work through fresh questions in IT Mastery before repeating this fixed set. Use the cheat sheet for recall and the study plan to organize the next attempt.
If something seems incorrect or unclear, send private feedback with the question number and the evidence you want reviewed.
Continue in the web app
Use IT Mastery for interactive Databricks Context Engineer Associate practice with mixed sets, timed mocks, topic drills, explanations, and progress tracking.