University of Alberta crest University of Alberta · Incoming M.Sc. in Computing Science

Tanzim Hossain Romel

I study how AI agents behave in real software systems—and how to make their decisions inspectable, reliable, and trustworthy.

About three years of professional software-engineering experience, formerly at IQVIA.

AI4SE · LLM4Coding · Trustworthy AI · long-horizon coding agents · AI for SRE

Tanzim Hossain Romel seated outdoors

Selected evidence

Questions, systems, and proof

Four records connect the research questions I ask to the software and evidence I produce.

  1. 01 · Research

    How do we know an SRE agent recovered a live system instead of applying a shallow fix?

    University of Illinois Block I mark representing the UIUC systems reliability research context
    Research context · UIUC++ SRSE 2026
    Contribution
    As a UIUC++ SRSE research intern, I work on failure scenarios with realistic mitigation paths, business-level oracles, and evidence that rejects shallow recovery.
    Evidence
    Active research · UIUC++ SRSE 2026
    Methods
    • SREGym
    • Kubernetes
    • incident replay
    • business-level oracles
    Verified
  2. 02 · Engineering

    How can coding agents inspect the right repository evidence before they edit?

    ctxhelm architecture connecting coding agents to its compiler, repository intelligence, and local source-free storage
    ctxhelm system architecture
    Contribution
    I designed and built ctxhelm, a local read-only context compiler, and HelmBench, its source-free evaluation harness for real coding-agent runs.
    Evidence
    Open-source tool · ctxhelm v2.4.0 released through GitHub archives and Homebrew
    Methods
    • Rust
    • MCP
    • retrieval
    • source-free evaluation
    • release engineering
    Verified
  3. 03 · Research

    What security risks emerge when model-hosting platforms load community code and artifacts at ecosystem scale?

    Methodology overview for the remote-code-execution study across machine-learning model hosting ecosystems
    Study methodology across five hosting ecosystems
    Contribution
    I co-authored a cross-platform empirical study spanning five ML hosting ecosystems, static security analysis, malware signatures, and developer discussions.
    Evidence
    Under review at ICSE 2027
    Methods
    • static analysis
    • CodeQL
    • Semgrep
    • YARA
    • qualitative analysis
    Verified
  4. 04 · Engineering

    How can coding agents inspect structural software changes instead of relying only on raw diffs?

    Contribution
    I redesigned RefactoringMiner's MCP diff surface around source-parameterized analysis tools for coding agents.
    Evidence
    Merged contribution · RefactoringMiner PR #1087
    Methods
    • Java
    • MCP
    • AST differencing
    • refactoring detection
    Verified

Research agenda

Reliable agents in evolving systems

Three connected questions organize my work across software engineering, reliability, and agent security.

  1. AI4SE · LLM4Coding

    How can long-horizon coding agents preserve context, repository memory, and architectural intent as software evolves?

    ContextLedger turns long sessions into typed, provenance-carrying events so context compaction and recovery can be inspected instead of trusted by intuition.

    Status Research system

  2. AI4SE · AI for SRE

    How should agents be evaluated when diagnosis and recovery happen inside live reliability conditions?

    SREGym studies realistic failure lifecycles, safe mitigation, business-level recovery checks, and the evidence needed to distinguish durable recovery from a shallow fix.

    Status Active research

  3. Trustworthy AI · AI4SE

    How can we inspect agent choices, security boundaries, and evidence when the final answer still looks acceptable?

    SHIFT compares matched tasks with and without a known trigger to test whether a changed valid choice systematically favors an attacker's target.

    Status Submitted to TACL 2026 · manuscript not public

View the full research agenda

Engineering experience

About three years in software engineering

Selected roles show the systems I owned, the changes I made, and the evidence behind them.

  1. Software Development Engineer 1

    IQVIA

    Backend engineer on KPI Library, the dynamic reporting layer within Orchestrated Analytics, across configuration, dashboard, analysis, and export workflows.

    • Built and maintained C#/.NET services and shared components across four KPI Library workflow areas.
    • Consolidated six duplicated filter-resolution implementations into two shared query components with centralized validation and regression scenarios.
    • Extended CSV and presentation export paths with regional formatting, commentary support, and broader automated coverage.

    Methods C#/.NET · EF Core · SQL Server · MongoDB · AWS S3

  2. Full Stack Engineer · part-time

    Mindshare Bangladesh

    Built an e-commerce aggregation system for brands working across major online retail platforms in Bangladesh.

    • Built a Scrapy collection system covering about 20 e-commerce platforms.
    • Deployed the scraper as a ScrapyRT service on DigitalOcean for API-driven extraction.
    • Built Express.js and MongoDB APIs for product aggregation and brand promotion workflows.

    Methods Python · Scrapy · Express.js · MongoDB · DigitalOcean

View complete experienceBrowse engineering projects

Publications and open source

Outputs with direct evidence

Research status and merged engineering work remain separate so each claim keeps its proper boundary.

Research outputs

  1. An Empirical Study on Remote Code Execution in Machine Learning Model Hosting Ecosystems

    Mohammed Latif Siddiq, Tanzim Hossain Romel, Natalie Sekerak, Beatrice Casey, and Joanna C. S. Santos.

    Under review at ICSE 2027

    Co-authored the empirical study and its cross-platform security analysis; the manuscript is publicly available on arXiv.

  2. The Choice Can Be the Attack: Auditing Aligned Backdoors in LLM Agents

    Tanzim Hossain Romel with Chowdhury Rakin Haider.

    Submitted to TACL 2026 · manuscript not public

    Co-developed SHIFT, a matched-trigger audit for whether an agent's changed valid choice favors an attacker's target after accounting for ordinary option quality.

View the publication record

Open-source proof

  1. SREGym/SREGym

    Added a Calico route-reflector label-drift problem for reliability-agent evaluation.

    Merged contribution

  2. dotnet/efcore

    Fixed OriginalValues materialization for added entities with complex collections.

    Merged contribution

Review curated contribution evidence

News and education

Recent trajectory

Three dated milestones connect current research work to the next academic step.

View all newsReview education

Collaboration

Discuss reliable agent systems

If you are studying dependable AI agents or building reliable software systems, I would be glad to compare research questions, evaluation evidence, or engineering tradeoffs.